A new benchmark arrived on September 28. The Artificial Analysis Cyber Index. Built with Collinear AI, IBM, NVIDIA, and Vercel, it evaluates AI for cyber defense.

The four-layer security post covered "review before merge." This post asks the next question. If AI does the review, what grades that AI? This index is the answer.

What this exam measures: a three-step defense loop

[1. Discover] find the flaw in code
  → [2. Reproduce] prove it with a crash or exploit condition
    → [3. Patch] fix it without breaking existing features

The scope is firm. Only defensive work with source code counts. Building live exploits is out of scope. It measures defense, not offense. Evals run on the open Stirrup harness in sandboxed, internet-blocked environments.

Read each of the three tests separately

Test Source What it checks Scale
CWE-Bench-AA Collinear AI Audit real repos and patch; verifier checks exploit blocked plus normal behavior 120 private tasks, all OWASP Top 10 areas
DeepsecBench-AA Vercel Report scanner-flagged files as confirmed flaws, scored F2 vs expert answers Recall-first scoring, median of 3 runs
CyberGym-E2E-AA Berkeley RDI Find, crash-reproduce, and patch C/C++ memory-safety bugs in 3 stages 131 tasks, 90-minute limit each

Each has a different flavor. CWE is "fix it end to end." Deepsec is "find it accurately." CyberGym is "do the whole chain." Like the SWE-bench, Terminal-Bench, and ProgramBench comparison showed, security tests measure different skills too.

Failure modes: where agents break

The published failure analysis is directly useful in production.

CWE-Bench:
- Partial patches 55%: main hole fixed, side entry still open
- Over-edits 24%: fix breaks normal features (more common in strong models, 40% in top 4)
- 38% of turns spent finding bugs, 62% patching and verifying

DeepsecBench:
- Best model finds only 41% of expert-confirmed issues
- Direct input-to-result bugs are easy; state-tracking bugs are missed
- Sequence-tracking hits about 30% on GPT-6 Sol and Astra, an early next-gen signal

CyberGym:
- 42% cannot even build a crashing input in 90 minutes (discovery is the bottleneck)
- Of passes, 31% fix a different real bug, not the target (stop-at-first-crash habit)

One line sums it up. Agents struggle less with finding flaws than with finding all of them and fixing them completely. Partial patches as the top failure hurt. "Fixed" does not mean done.

What the 98% refusal rate means: ship with a fallback

This is the most practical finding. In CyberGym, GPT-6 Astra and Sol, Fable 5.1, Opus 5.5, and the Qwen3.8 family refuse over 98% of tasks on safety grounds. Some exams draw lines by refusal rate, not score.

Product implications:
- Treat security agents as "calls that may refuse"
- Route refusal → fallback model or human queue
- Track refusal rate as its own metric (separate from success)
- Never merge refusal and failure into one number (AA keeps them apart)

The principle from the hallucination vs refusal post repeats in security. "I do not know / I will not do it" belongs in your metrics. Hide refusals and ops will break.

CodeBridge Mini Lab: run one defense loop on your own repo

1. Build 3 CWE-style mini tasks:
   - Prepare 3 small functions with known flaw patterns
   - Ask the agent to "audit and patch"

2. Judge in 3 stages:
   [ ] Discover: did it locate the flaw?
   [ ] Reproduce: did it prove it with a test or input?
   [ ] Patch: do all existing tests still pass?

3. Classify the failure:
   - Partial patch? (check the flank: add a second-entry test)
   - Over-edit? (catch with full tests)
   - Refusal? (check prompt and model policy, then inspect fallback path)

This applies the verification loop from the Grok self-check post and the five-question repo method from the Terminal-Bench 4.0 update to security tasks.

Conclusion: grade security agents with four numbers

The Cyber Index teaches a scorecard format.

Track discovery rate, patch success, refusal rate, and cost per task together.

Success rate alone gets fooled by partial patches. Skipping refusals blocks you in ops. Skipping cost blows your budget. Argon's cyber defense, Mythos Glasswing, and Fable science evals landed in the same month for a reason. Security is now its own eval axis, not a model side feature. Add that axis to your team scorecard first.

Further reading

References

Go deeper with a course

If you want to run find, fix, and verify loops in a real repo, this course builds harnesses and verification gates in Claude Code exactly like the defense loop here.