On August 18 local time, OpenAI disclosed that it has deliberately slowed the pace of frontier model development. Reinforcement learning on its latest models headed for deployment was paused for two weeks, and the largest planned RL run is still on hold. The company is raising the bar on three fronts, monitoring, alignment, and security, and will not scale up further until the work catches up.

Two events triggered the slowdown

OpenAI points to two developments. The first is the Hugging Face security incident that surfaced in July. The second is preliminary evidence that Astra, one of its upcoming models, may reach the Critical cybersecurity capability threshold defined in its own Preparedness Framework.

As models grow more capable, the risks do not appear only after deployment. They grow inside the lab as well, during training and evaluation. OpenAI acknowledges this directly, saying it wanted the time needed to meet its own standards and therefore slowed scaling on purpose.

Smaller training runs and evaluations now come first, used to observe model behavior, validate safeguards, and build evidence of alignment before the largest run resumes.

CEO Sam Altman posted on X that the company paused some frontier RL so it could meet appropriate alignment, security, and monitoring standards for the level of capability now within reach. He argued the industry needs shared standards but said OpenAI will act on its own until those exist. He also noted that strong new models are still coming soon, and that the pause affects releases further out.

Rebuilding the research environment

The first pillar is the environment where models are trained and evaluated. Immediately after the Hugging Face incident, OpenAI halted frontier model inference in research clusters for any run that could execute code or reach the internet through tools. A limited, more secure path for code execution was restored first, then each workload was reviewed individually before resuming.

New requirements include stronger sandboxing for workloads running model-generated or otherwise untrusted code, network isolation designed so that a single compromise does not by itself open a path to the internet or other internal networks, and continuous security testing that throws simulated attacks from its own models at those boundaries. Shared services have been trimmed and standing privileges reduced.

The strictest tier applies to Astra and to cyber-related workloads, and a significant number of those runs remain paused until they are fully migrated. Safety and alignment workloads are being moved first. OpenAI does not hide the cost, stating plainly that meeting the new bar has brought substantial expense and delay to frontier research.

Multistage monitoring of what the model is doing

The second pillar is monitoring. Coverage used to concentrate on the highest-risk cases such as internal deployments and large RL runs, and did not extend to every inference or training run that used tools.

Under the new setup, activation classifiers inspect the model's internal activity at every sampled token. Anything suspicious escalates to more sophisticated, higher-compute automated investigators, which examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards.

The target is an alert within 30 minutes of concerning activity surfacing. A likely breach of a critical security boundary generates a highest-priority alert that pages the safety, security, and research teams. If they cannot rule it out as a false positive within 30 minutes, the activity is expected to stop.

This monitoring is mandatory for all RL training and evaluation involving tools on models of GPT-5.6 Sol capability or higher. For Astra, it was extended on August 7 to all inference with tools. The overhead is real, currently estimated at roughly 20 percent of the inference compute being monitored.

Alignment applied across every stage of training

The third pillar is alignment. As capability rises, behaviors such as reward hacking, deception, and unauthorized access become serious risks rather than curiosities.

For RL runs on the most capable models, core alignment techniques now apply across more stages of training. That includes improving reward models to catch and discourage unsafe behavior, training models to be more honest about their own actions, capabilities, and limits, and reducing behavior that exploits weaknesses in rewards, graders, tools, or oversight. Training coverage is also expanding for situations where models interact with external systems and resources.

The framework itself is being revised

OpenAI plans to revise the Preparedness Framework so these safeguards fit together across both training and deployment, and so the framework reflects the capabilities of future models and the environments they run in. External organizations will be involved, and a technical report on the Hugging Face incident is due within the coming weeks.

The company expects models themselves to handle most security work before long, including defending against other models. If the defending side can be staffed by models, all three safeguards can scale alongside capability.

Summary

The striking part of this announcement may be less the pause itself than the decision to talk about it. OpenAI admits to real delay and cost in frontier research, then lays out specifics: layered monitoring, isolated environments, and alignment applied across more of the training process. Altman's framing, that confidence in safety increasingly sets the pace of progress, is a useful yardstick for reading the next stretch of AI development.