OpenAI paused some frontier RL training, and says the next launches are unaffected
OpenAI says the controls will govern higher-capability models, while separate reporting links other safety changes and a two-week training pause to a Hugging Face breach.
OpenAI said it was pausing internal Astra activities that did not meet strengthened security-control requirements after its internal evaluations and expert assessments indicated that it could not rule out critical cyber capabilities under its Preparedness Framework. OpenAI described Astra as an upcoming model. On August 7, 2026, OpenAI said its latest internal evaluations showed significant advances in agentic coding and cybersecurity.
On August 18, Sam Altman said OpenAI had paused some frontier reinforcement-learning training to ensure it could meet alignment, security and monitoring standards for what he called the new level of capabilities in front of the company. In a follow-up post he said OpenAI still expects to ship new models soon and that the pause affects further-out releases.
OpenAI said the strengthened controls were for higher-capability models. The company said the measures include isolated testing, restricted network and tool access, protections for model weights, monitoring and detection, and sandboxed execution. OpenAI said internal Astra work would be paused when it did not meet those requirements.
Techmeme reported that OpenAI said it had made several changes to its safety practices following the Hugging Face breach. Techmeme also reported that OpenAI had paused two weeks of deployment-focused reinforcement-learning training. That account does not say that the two-week training pause was the same action as the pause of Astra activities.
WIRED published a headline saying that OpenAI had overhauled safety protocols after its AI agents went rogue. The available reporting does not independently verify either the alleged rogue-agent incident or the extent of the claimed overhaul. OpenAI's statements instead describe the stronger controls and the Astra pause alongside its evaluation findings.
OpenAI described Astra as upcoming and said internal activities would stop if they failed the stronger requirements. The company reported advances in agentic coding and cybersecurity, but said critical cyber capabilities could not be ruled out under its framework. That wording does not establish that Astra demonstrated those capabilities or was publicly released.
Sources
This article was written from these pages. Read them.
- discoveryOpenAI says it has made several changes to its safety practices following the Hugging Face breach and has paused two weeks of deployment-focused RL training (Intechmeme.com
- secondaryOpenAI Overhauls Safety Protocols After Its AI Agents Went Roguewired.com
- discoveryOpenAI Overhauls Safety Protocols After Its AI Agents Went Roguereddit.com
Written from verified primary sources by Epoch's editorial pipeline and checked by a human before publication.