Article

AI Cybersecurity ranking upset:GPT-6, Claude flagships fall behind due to security refusals, while Xiaomi MiMo ties Grok for first place

Dongcha Beating AI news: Artificial Analysis has launched an independent Cyber Index, specifically evaluating AI Agent' ability to conduct enterprise cyber defense. Rather than selecting several tests from its original general Intelligence Index, it instead uses CWE-Bench-AA、DeepsecBench-AA and CyberGym-E2E-AA three cybersecurity evaluations, testing vulnerability auditing and remediation, vulnerability discovery, and the full workflow from discovering a vulnerability to reproducing and remediating it, with each accounting for one-third.

In the initial ranking, Grok 4.7(xhigh) and Xiaomi MiMo-V2.6-Pro both scored 56 points,GPT-6 Luna(max)53 points, and GLM-5.3-Flash 50 points. MiMo's average cost per task was only 0.18 USD,Grok 4.7 compared with 11.67 USD; Luna was even lower, at only 0.12 USD.

However, this ranking should not be simply interpreted as a ranking of the models' cybersecurity capabilities.GPT-6 Sol、GPT-6 Astra and Claude Opus 5.5 both participated in the tests, but refused to answer almost all tasks on CyberGym-E2E-AA. Sol and Astra refused every task, while Opus 5.5 refused 98%. Refusals are directly scored as 0 points. Luna, by contrast, was willing to continue completing these tasks, resulting in a higher overall score. Artificial Analysis also stated that the large number of refusals makes the actual capabilities of these frontier models difficult to assess.GLM-5.3 also showed a clear anomaly. The full version ranks above Flash on Artificial Analysis's general-capability leaderboard, but this time Cyber Index scored only 36 points, while Flash reached 50 points. The main gap came from CyberGym-E2E-AA: Flash received 74%, while the full version received only 29%. Artificial Analysis has not yet explained the reason for this gap.

Original link https://m.theblockbeats.info/flash/369529