Beating AI News: Artificial Analysis has launched an independent Cyber Index specifically designed to evaluate AI Agents' capabilities in enterprise cyber defense. Rather than pulling a few tests from its original general Intelligence Index, it separately adopts three cybersecurity benchmarks—CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA—to test vulnerability auditing and patching, vulnerability discovery, and the full pipeline from discovering a vulnerability to reproducing and patching it, with each of the three accounting for one-third.
In the first batch of rankings, Grok 4.7 (xhigh) and Xiaomi MiMo-V2.6-Pro are tied at 56 points, GPT-6 Luna (max) at 53 points, and GLM-5.3-Flash at 50 points. MiMo costs only $0.18 per task on average, while Grok 4.7 costs $11.67; Luna is even lower, at just $0.12.
However, this leaderboard cannot simply be understood as a ranking of the models' own cybersecurity capabilities. GPT-6 Sol, GPT-6 Astra, and Claude Opus 5.5 all participated in the tests, but on CyberGym-E2E-AA they almost entirely refused to answer, with Sol and Astra refusing all tasks and Opus 5.5 refusing 98%. Refusals are directly scored as zero, while Luna was instead willing to continue completing these tasks, resulting in a higher overall score. Artificial Analysis also stated that the large number of refusals makes it difficult to judge the actual capabilities of these frontier models.
GLM-5.3 also showed a clear anomaly. The full version ranks higher than Flash on Artificial Analysis's general capability leaderboard, but on this Cyber Index it scored only 36 points, while Flash reached 50 points. The main gap comes from CyberGym-E2E-AA: Flash scored 74%, while the full version scored only 29%. Artificial Analysis has not yet explained the reason for this gap.

