OpenAI’s New Model, GPT-6 Astra: Benchmarks and Capabilities
OpenAI has released GPT-6 Astra, succeeding GPT-5.6 Sol as their top-tier model. The company highlights its advancements in pre-training, reinforcement learning, and alignment research, making it the most intelligent and aligned system to date.
Key Features and Benchmarks:
- Astounding Performance: Astra demonstrates significant improvements in various benchmarks. It achieves a 99.9% score on ARC-AGI-3 (up from 7.8% for its predecessor), 97.6% on FrontierMath Tier 4, and 64.6% on a science workflow benchmark.
- Speed: On OSWorld 2.0, Astra completes tasks 47% faster than Sol, with a completion time of approximately 40 minutes compared to 75 minutes.
- Practical Applications: Astra can create documents, presentations, and spreadsheets using business templates, enhance Codex’s note-taking capabilities, and perform asynchronous question-asking.
- Cybersecurity: Astra excels in cybersecurity, scoring 100% on ExploitBench (compared to 78.5% for Sol) and discovering two previously unknown zero-day vulnerabilities during testing. This performance triggers OpenAI’s own safety framework’s "Critical" threshold.
Safety Measures:
OpenAI has implemented several safety restrictions with Astra:
- Refusing to write proof-of-concept exploits.
- Pausing or stopping tasks that may involve offensive security.
- Enterprise access is disabled by default until enabled by an administrator.
Security researchers will gain access to a specialized version of the model through OpenAI’s Daybreak program, designed to handle legitimate defensive requests.