Booz Allen’s AI Threat Ranking: A Cheaper Software Overturns Results
Booz Allen scored 18 of the world’s most advanced AI models against a live corporate network, aiming to see how far each could penetrate unaided.
Only Anthropic’s Claude Mythos reached the end. Then, when given a model ranked 15th an attack harness, it matched the leader.
September 3, 2026 – 7:58 am
Credit: © Tiero via Canva.com
Booz Allen’s Cyber Weapon Index, published on Wednesday, scored 9 American and 9 Chinese models. The test measured their ability to infiltrate a real network with no human intervention.
What the Test Involved
- Each model had its own attacker machine.
- They executed commands sequentially, without any tool menu or supporting software.
- Scoring combined two parts: vulnerability discovery in compiled software and intrusions into an Active Directory network.
- Booz Allen relied on network logs, security records, and intrusion detection systems to score the models’ actions rather than their claims.
Key Findings
Claude Mythos scored 80, achieving administrator control with stolen credentials and autonomously navigating to higher privileges. It also successfully broke in from outside without any credentials.
Grok-4.5 and GPT-5.6 Sol followed closely behind.
A surprising discovery: Claude Sonnet 5, initially ranked 15th, improved significantly when paired with an attack harness, challenging the leader, Claude Mythos.
Booz Allen concludes that the "model is no longer the unit of risk; the system is." They also acknowledge unmeasured variables like optimized harnesses and Chinese or open-weight models.
A second notable finding: one model refused a task due to lack of credentials, while its sibling, cyber-tuned, completed it.