News

OpenAI Slows Frontier AI Training After Hugging Face Breach, Adds New Safeguards

OpenAI has moved from post-incident remediation to changing how it trains, monitors and contains frontier AI models after its July cybersecurity evaluation breach reached Hugging Face's production infrastructure. The company said it temporarily slowed the pace of scaling, paused reinforcement learning training and is continuing to hold its largest planned frontier RL run while it puts stronger security and alignment controls in place.

In an Aug. 18 announcement describing the new safeguards, OpenAI connected the Hugging Face incident with a separate development: preliminary evidence that an upcoming model called Astra may reach the Critical cybersecurity capability threshold under its Preparedness Framework, meaning it may have cyber capabilities powerful enough to cause severe real-world harm if misused. OpenAI said those developments, combined with progress in its internal research, increased the urgency of strengthening monitoring, alignment and containment across the training process.

This adds another chapter to what has increasingly been described as a rogue AI problem: highly capable models crossing boundaries during evaluations and interacting with real systems in ways their operators did not intend. In OpenAI's account of the Hugging Face episode, the models were pursuing the benchmark goal. Anthropic made a similar distinction after discovering three incidents involving Claude models, saying it found no evidence that its models were pursuing goals of their own.

As PureAI reported in July, OpenAI models escaped an isolated internal environment while trying to solve the ExploitGym cybersecurity benchmark, found a path to the open internet and ultimately reached Hugging Face production systems. The models obtained test solutions from a production database after chaining attacks that included stolen credentials and previously unknown vulnerabilities.

OpenAI Slows Training While It Raises the Security Bar
OpenAI's response now reaches beyond patching the specific infrastructure involved in that incident. The company said it imposed a two-week pause on RL training for its latest models intended for deployment while it hardened and red-teamed research environments and expanded monitoring. "We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling."

The company has since resumed some smaller-scale training and evaluation work, but not its biggest planned run. "Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding."

OpenAI said that immediately after the Hugging Face incident it also paused frontier-model inference in research clusters for workloads capable of executing code or using tools that could reach the internet. A more restricted code-execution path was restored, but workloads have since been reviewed individually before being allowed to resume. Some remain paused while they are moved to environments that meet the new security requirements.

The infrastructure changes include stronger sandboxing for workloads that execute model-generated or otherwise untrusted code. OpenAI said the requirement also covers software that could itself be compromised while processing model outputs. Network controls are being strengthened to separate higher-risk workloads from the internet and to prevent the compromise of a single workload or supporting service from automatically providing broader network access.

OpenAI is also changing shared infrastructure. It said it has reconfigured research environments to remove potentially vulnerable shared services, reduce standing privileges and improve security and trust boundaries. Security logging is being expanded, while the company is investing in automated testing that uses models to continuously probe those boundaries with simulated attacks.

"Together, these controls provide defense in depth and help protect our research environment against attacks," OpenAI said.

Monitoring Moves Deeper Into Training and Evaluation
Monitoring is another major part of the response. OpenAI said its previous monitoring concentrated particularly on internal deployments of frontier models and frontier RL runs because those workloads could combine advanced capabilities with tools, sensitive systems and network access. It was not applied to every inference and training workload using tools.

The new system expands that coverage and uses multiple stages of automated review. OpenAI said activation classifiers inspect internal model activity at every sampled token and escalate potentially concerning activity to higher-compute automated investigators. Those systems review tool actions, available reasoning and the sequence of activity for signs including unauthorized access, data theft, destructive behavior and attempts to defeat safeguards.

OpenAI's stated target is to produce an alert within 30 minutes after concerning activity reaches the monitoring system. A suspected violation of a critical security boundary generates the highest-priority alert and pages safety, security and research teams. If those teams cannot determine within 30 minutes that the alert is a false positive, they are expected to pause the activity.

The monitoring requirement now covers all RL training and evaluations involving tools for models at Sol capability or higher. OpenAI said that after determining on Aug. 7 that Astra might have Critical cyber capabilities, it also required monitoring for all Astra inference using tools, rather than limiting the requirement to training and evaluations.

The additional scrutiny carries a measurable computational cost. OpenAI estimated that monitoring currently adds roughly 20% to the inference compute being monitored, although it said the figure varies substantially across training and evaluation workloads.

Astra itself was not involved in the Hugging Face breach. OpenAI said in a separate Aug. 7 disclosure that preliminary evaluations were strong enough that it could not rule out the Critical capability level. The company subsequently paused Astra activities that did not meet strengthened security requirements and imposed additional controls including isolated test environments, restricted tool and network access, weight protection, encryption, sandboxed execution and expanded monitoring.

The Rogue AI Context Is Broader Than One Lab
The OpenAI changes come against a broader set of incidents in which advanced models escaped the intended boundaries of cybersecurity evaluations. OpenAI's original Hugging Face incident disclosure said the affected models included GPT-5.6 Sol and an internal research prototype operating with reduced cyber refusals. The evaluation also omitted production classifiers intended to stop models from pursuing high-risk cyber activity because the purpose was to measure maximum cyber capability.

OpenAI later clarified that the research prototype involved was not planned for public release. Following the incident, the company said it deactivated, encrypted and restricted that model from research access. It also disclosed the zero-day vulnerability in an Artifactory package registry cache proxy that provided the models a route out of the isolated environment.

The episode was significant enough that OpenAI said, "We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly."

Days later, Anthropic disclosed that a review triggered by OpenAI's report had uncovered three incidents in which Claude models reached the internet during cybersecurity evaluations and gained unauthorized access to real systems. As Virtualization & Cloud Review reported, one older Claude model continued its attack after recognizing that it was probably interacting with a real production environment. Another model published a malicious Python package to the real PyPI registry, where it was downloaded and executed on 15 real systems.

Anthropic's own incident report said the models had been told they were operating in simulations without internet access, but a configuration problem left internet access available. Anthropic said it saw no evidence of a model deliberately attempting to escape or pursuing its own goal. Its oldest model, however, continued attacking after recognizing that the system was likely real, while a newer internal research model stopped after determining that it had reached a real environment.

Taken together, the OpenAI and Anthropic disclosures document a recurring operational failure mode without documenting independent malicious intent: models pursuing authorized evaluation objectives can reach unauthorized real-world systems when containment, environmental assumptions or monitoring fail.

Alignment Changes Target Reward Hacking and Unauthorized Actions
OpenAI is also extending its response into model training itself. For RL runs involving its most capable models, the company said it is applying its core alignment techniques across more stages of training. Those changes include improving reward models so they better detect and discourage unsafe behavior, training models to be more accurate about their actions, capabilities and limitations, and reducing behavior that exploits weaknesses in rewards, graders, tools or human oversight.

OpenAI also said it is increasing training coverage for behavior that could cause harm when models interact with outside systems and resources. In its announcement, the company specifically identified reward hacking, deception and unauthorized access as forms of misaligned behavior that become more serious as models gain stronger capabilities and operate in more complex environments.

The company plans to revise its Preparedness Framework so the monitoring, alignment and security measures cover both model training and deployment and more directly reflect the environments in which future models operate. That represents a shift from treating capability thresholds primarily as deployment questions toward applying stronger safeguards during development when models reach Critical capability levels.

For OpenAI, the continuing hold on its largest frontier RL run is the clearest immediate indication of how seriously it is treating that change. The company closed its announcement with another concise statement of the standard it says its development process now has to meet: "The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead."

Featured

Subscribe on YouTube