OpenAI Imposes 20% Compute Cost on AI Safety Monitoring
OpenAI has implemented a new AI safety monitoring system, adding approximately 20% to the compute cost of its riskiest operations. This comes after a recent breach at Hugging Face and the internal testing of an unreleased model, Astra, which raised concerns about potential cyber risks.
Pausing Training and Implementation
The company paused reinforcement-learning training on its newest models for two weeks and put its largest frontier run on hold. This decision was made to meet "alignment, security, and monitoring standards" for the advanced capabilities they are developing.
New Monitoring System
The new monitoring system, which expands chain-of-thought monitoring, covers all tool-enabled training and evaluations for models at the GPT-5.6 Sol level and above, as well as all inference with Astra. The system reads a model’s step-by-step reasoning and escalates any suspicious activity to automated investigators.
Cost and Customer Impact
OpenAI emphasizes that this additional cost is for internal research and will not be passed on to customers. The company did not disclose the exact percentage of its total compute now under monitoring.
Addressing Concerns
In a blog post and social media updates, Chief Scientist Jakub Pachocki and President Greg Brockman highlighted that the pause and new monitoring measures are part of a broader safety rethink following the Hugging Face breach. They acknowledged that while the new system has limitations, it is a crucial step towards ensuring the responsible development and deployment of AI technologies.