Back
Artificial Analysis launches the Cyber Index Alliance with Collinear, IBM, NVIDIA and Vercel
SiTech AI Team3 წთ. საკითხავი

Artificial Analysis launches the Cyber Index Alliance with Collinear, IBM, NVIDIA and Vercel

Artificial Analysis announced the Cyber Index Alliance on September 25, 2026, alongside the Artificial Analysis Cyber Index, which measures how well AI models handle enterprise cyber defense tasks.

Artificial Analysis announced the Cyber Index Alliance on September 25, 2026, together with the Artificial Analysis Cyber Index, a benchmark suite that measures how well AI models handle enterprise cyber defense tasks. The launch partners are Collinear AI, IBM, NVIDIA and Vercel.

The partners

Collinear AI, the developer of CWE-bench, contributed its benchmark as a private held-out evaluation, and Vercel did the same with DeepsecBench. IBM and NVIDIA joined with expert input on the scope and methodology rather than new datasets. According to Artificial Analysis, the partners help set a standard for evaluating models on cyber defense and may contribute datasets and external research directly. Organizations that want to join can contact the company by email.

How the index works

The index combines three evaluations that cover the defensive loop: finding vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing behavior. CWE-Bench-AA uses 120 held-out tasks covering all ten OWASP Top 10 (2025) categories across six languages, from C/C++ and Go to Python and Rust. DeepsecBench-AA isolates discovery, scoring an agent's findings in open-source application code against a golden set verified by human experts. CyberGym-E2E-AA draws 131 tasks from the 920-instance Berkeley RDI dataset and targets memory-safety bugs in C/C++ projects such as FFmpeg and CPython.

All three run on Stirrup, the company's open-source agent harness. Models work from source code, the way a security engineer auditing an application would, and none is asked to build a working exploit, because exploit realization sits outside a defense-focused index. Refusals on safety grounds are reported separately from scores.

First results

The best model solves 56% of index tasks at launch, with Grok 4.7 (xhigh) leading the combined ranking. The CWE-Bench-AA leader reaches 68% and the CyberGym-E2E-AA leader 79%. On DeepsecBench-AA the top F2 score is 46%, and the strongest model finds only 41% of the expert-verified issues.

Failures fall into a few patterns. Excluding refusals and timeouts, 55% of failed CWE-Bench-AA attempts fixed the main problem but left a related one open, while over-corrections that break legitimate behavior make up about 24% of failures. On CyberGym-E2E-AA, 42% of attempts reached the 90-minute limit without producing an input that crashes the program, and six frontier models, among them GPT-6 Sol, GPT-6 Astra and Claude Opus 5.5, refused at least 98% of tasks.

Roadmap

The company launched the index with three evaluations rather than waiting for full coverage, and plans to add incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.