OpenAI’s new model aces the benchmarks and admits it is better at hiding

OpenAI’s New Model, GPT-6 Astra: Benchmarks and Capabilities

OpenAI (164news.com's guide to openai) (164news.com's guide to openai) has released GPT-6 Astra, succeeding GPT-5.6 Sol as their top-tier model. The company highlights its advancements in pre-training, reinforcement learning, and alignment research, making it the most intelligent and aligned system to date.

Key Features and Benchmarks:

  • Astounding Performance: Astra demonstrates significant improvements in various benchmarks. It achieves a 99.9% score on ARC-AGI-3 (up from 7.8% for its predecessor), 97.6% on FrontierMath Tier 4, and 64.6% on a science workflow benchmark.
  • Speed: On OSWorld 2.0, Astra completes tasks 47% faster than Sol, with a completion time of approximately 40 minutes compared to 75 minutes.
  • Practical Applications: Astra can create documents, presentations, and spreadsheets using business templates, enhance Codex's note-taking capabilities, and perform asynchronous question-asking.
  • Cybersecurity: Astra excels in cybersecurity, scoring 100% on ExploitBench (compared to 78.5% for Sol) and discovering two previously unknown zero-day vulnerabilities during testing. This performance triggers OpenAI’s own safety framework's "Critical" threshold.

Safety Measures:

OpenAI has implemented several safety restrictions with Astra:

  • Refusing to write proof-of-concept exploits.
  • Pausing or stopping tasks that may involve offensive security.
  • Enterprise access is disabled by default until enabled by an administrator.

Security researchers will gain access to a specialized version of the model through OpenAI's Daybreak program, designed to handle legitimate defensive requests.