OpenAI has decided to temporarily slow down the scaling of its new AI models and has paused reinforcement learning (RL) training for its latest systems for two weeks, which are intended for future deployment.

As models become more capable, the risks associated with developing and testing them internally also grow.

We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research…

— OpenAI (@OpenAI) August 18, 2026

This decision was influenced by two incidents: a situation involving Hugging Face and a preliminary assessment of Astra, which led OpenAI to conclude that the model could reach the Critical level — the highest tier of cyber capabilities in the company's Preparedness Framework.

OpenAI acknowledged that these signals, coupled with the rapid progress of its internal research, necessitated urgent enhancements to its monitoring, model behavior alignment methods, and technical limitations at all stages of development.

Part of the Training Remains on Hold

Following the Hugging Face incident, OpenAI halted the operation of advanced models in research clusters for tasks where the systems might execute code or utilize tools with potential internet access.

Later, the company reinstated a more limited and secure code execution mechanism and began to evaluate research tasks individually. Some processes resumed under new restrictions, while others required further modifications.

Concurrently, OpenAI paused RL training for its latest deployment models for two weeks, using this time to strengthen research environments, conduct red team assessments, and enhance monitoring.

The largest planned RL launch has not yet been resumed. Instead, OpenAI is conducting smaller-scale experiments and additional evaluations to assess model behavior and protective mechanisms before further scaling.

A significant portion of the training and testing processes for Astra also remains on hold until the transition to updated infrastructure is completed, prioritizing security-related tasks.

On August 7, OpenAI released preliminary results from Astra testing, showing considerable improvements in agent programming and cybersecurity tasks. Based on the test outcomes and expert evaluations, the company stated that it cannot yet rule out the model reaching Critical capability.

Despite the pause in some processes, prediction market participants continue to anticipate the imminent release of Astra. According to Decrypt, as of August 18, traders on Polymarket estimated a 59% probability of the model's release by September 15 and a 72% chance by the end of the month.

Research Infrastructure Overhauled

In the weeks following the incident, OpenAI established stricter requirements for working with advanced models. The company enhanced the isolation of environments where generated or untrusted code is executed, and added further network access restrictions for high-risk tasks.

The infrastructure was reorganized to ensure that the compromise of a single workload does not automatically allow the model unauthorized internet access or connections to other internal networks.

OpenAI also eliminated potentially vulnerable shared services, reduced persistent privileges, and strengthened trust boundaries between elements of the research infrastructure.

The most stringent security protocols are now mandatory for tasks related to Astra and specialized cyber models.

Model Actions Now Monitored in Stages

A separate set of changes pertains to monitoring. Specialized classifiers analyze the model's internal activity during generation. Potentially dangerous signals are sent to more complex and resource-intensive automated investigation systems.

These systems verify the model's actions with tools, access recorded reasoning, and scrutinize the complete sequence of activities to detect signs of unauthorized access, data theft, destructive behavior, or circumvention of protective measures.

The goal of this system is to issue a warning no later than 30 minutes after suspicious activity is detected.

If monitoring identifies a likely breach of critical security thresholds, a high-priority alert is generated, which is sent to security and research teams.

If experts cannot confidently determine within the next 30 minutes that the alert was false, the corresponding activity must be suspended. Monitoring is mandatory for all RL training and assessments using tools for models of Sol level and above.

Following the preliminary assessment of Astra on August 7, OpenAI extended this requirement to any inference involving Astra with tools, not just training and testing.

This level of oversight requires additional resources. The company estimates that overhead costs currently account for about 20% of the computations needed for monitored inference, with the exact figure varying based on specific workloads.

OpenAI Strengthens Model Behavior Management Techniques

Simultaneously, the company is enhancing the methods for aligning the behavior of its most powerful models. For RL launches, OpenAI has begun to apply core approaches across a greater number of training stages.

Specifically, the company is improving models designed to identify and suppress unsafe behavior and is training systems to describe their own actions, capabilities, and limitations more accurately and honestly.

Another focus is on reducing the tendency of models to seek ways to achieve high rewards without completing tasks as intended by developers. OpenAI is also working to lower the likelihood of deception and unauthorized access to systems and data.

The next step for the company is to revise the Preparedness Framework, with the new version expected to integrate training and deployment requirements and better account for the capabilities of future models.

Brockman: OpenAI Underestimated Model Capabilities

On August 17, OpenAI President Greg Brockman acknowledged that the Hugging Face incident revealed the company underestimated the actual cyber capabilities of its own models.

He believes the incident illustrates how significantly the capabilities of ordinary attackers could evolve in the coming months. Brockman also emphasized that the same advancements could benefit defenders.

OpenAI is already utilizing models to protect its own infrastructure. According to him, nearly all initial alerts about potential incidents are first analyzed by AI before the most critical cases are escalated to humans.

Systems are also employed to identify potential attack vectors, configuration errors, excessive access privileges, and weak trust boundaries.

Brockman urged companies not to wait for even more powerful cyber models but to start using AI agents for controlled analysis of codebases, infrastructure, and technical documentation. He warned against an immediate shift to fully autonomous defense, advocating for starting with limited scenarios and read-only access, preserving critical decisions for humans.

Brockman predicts that organizations will need to significantly automate their cybersecurity programs in the coming months to keep pace with the increasing capabilities of attackers.

Additionally, it was noted that Astra from OpenAI achieved 10 results in mathematics and theoretical computer science in August.