OpenAI Pauses and Ships a Model: Cyber Risk and GPT-5.6-Cyber
OpenAI paused a model over cyber risk on Friday. On Monday, it shipped one trained to refuse less, GPT-5.6-Cyber. This release comes three days after the company delayed another model, Astra, due to concerns about critical cyber capabilities.
The Preparedness Framework
Both incidents fall under the same Preparedness Framework, which governs a model’s capabilities rather than who gets to use it. This framework has sparked discussions about the industry’s stance on cyber risk.
GPT-5.6-Cyber: A Deep Dive
OpenAI released GPT-5.6-Cyber on Monday, a model built on GPT-5.6 Sol and trained for zero-day discovery and exploit-chain development. It is also designed to refuse fewer higher-risk dual-use cyber requests.
Access to this model is restricted through Daybreak, OpenAI’s vetted cybersecurity program. Axios‘ Sam Sabin was the first to report on this release.
A Reversal or a Shift in Perspective?
Three days prior, OpenAI had delayed Astra due to unresolved cyber risks. This new development might seem like a reversal, but it clarifies where the industry draws its lines regarding cyber capabilities.
Daybreak now consists of two tracks:
- Daybreak Blue: General-purpose frontier models, including GPT-5.6 Sol, with system-level cyber guardrails removed, suitable for defensive tasks like vulnerability discovery and incident response.
- Daybreak Red: Purpose-trained cyber models like GPT-5.6-Cyber for authorised vulnerability research and security testing.
Performance Metrics
OpenAI tracks an internal measure called the Advanced Cybersecurity Completion Rate, which measures how often a model completes tasks in categories such as exploit-chain development and privilege escalation.
- GPT-5.6-Cyber scores 95.0%.
- Its predecessor, GPT-5.5-Cyber, had a completion rate of 57.3%.
- Models through Daybreak Blue have much lower rates: 2.0% without guardrails and 1.5% with standard safeguards.
Discoveries and Vulnerabilities
GPT-5.6-Cyber was tested against V8, the JavaScript engine inside Chrome, uncovering two previously unknown vulnerabilities that could be chained to corrupt memory and escape the sandbox. Google received these vulnerabilities through coordinated disclosure and issued a fix, assigning them a high-severity identifier, CVE-2026-15903.
The model also identified at least five vulnerabilities in a mobile operating system, three critical ones in a popular database, and over 400 privilege escalation flaws in an OS kernel.
Benchmarks and Limitations
While GPT-5.6-Cyber excels at specific offensive tasks, it doesn’t consistently outperform its predecessor across the board. For instance, on vulnerability discovery and report writing, it performs worse, and on ExploitBench, it uses fewer tokens at 300 turns but narrows the gap at 600 turns.