OpenAI announced a temporary slowdown in the development of its most advanced AI models, instituting a two-week pause on reinforcement learning training for its latest models as it strengthens security and monitoring measures. The company said the move is intended to better protect against risks arising from increasingly capable AI systems.
OpenAI noted that its next model, Astra, could meet a critical cybersecurity capability threshold under its Preparedness Framework. That potential classification, combined with a recent operational incident involving OpenAI and Hugging Face, led the company to intensify its safeguards throughout the training lifecycle.
The organization's single largest planned frontier reinforcement learning run remains paused. In the interim, OpenAI is running smaller-scale training and evaluation experiments to observe model behavior and to validate the effectiveness of newly implemented protections.
New security mandates have been added for frontier research workloads. Those requirements include stronger isolation for workloads that execute model-generated code and extra controls to segregate higher-risk workloads from Internet access. Workloads that involve Astra or other cyber models are now subject to the strictest level of security safeguards.
A substantial portion of higher-risk workloads remains on hold while teams migrate them to comply with the new security standards. OpenAI said it will prioritize migrating safety and alignment workloads ahead of other work.
To improve detection and response, OpenAI expanded its monitoring architecture to deploy activation classifiers that run at every sampled token. When these classifiers flag concerning activity, the system escalates to automated investigators. The monitoring design aims to generate an alert within 30 minutes of detecting worrying behavior. If the system determines a likely breach of a critical security boundary, it will emit a highest-priority alert that pages safety, security, and research personnel.
OpenAI reported that the additional monitoring imposes overhead: roughly 20% of the inference compute being monitored is required for the monitoring processes, with the precise cost impact varying between training and evaluation workloads.
Finally, the company said it plans to update and evolve its Preparedness Framework to align with the capabilities of future models and the environments in which those models operate.
Summary
OpenAI has paused certain reinforcement learning activities for two weeks and delayed a major frontier run while it implements stricter isolation, enhanced monitoring, and migration of high-risk workloads to meet new security requirements. The pause follows the assessment that Astra may hit a critical cybersecurity threshold and after a recent operational incident prompted broader safeguard upgrades.