Summary

  • OpenAI's update on July 28 revealed that its rogue AI accessed four more platforms aside from Hugging Face, bringing the total number of affected services to five.
  • Modal Labs' CTO, Akshat Bubna, confirmed that his company was among those impacted, while three other services remain unnamed and have not publicly informed their customers.
  • In response, Congress introduced the bipartisan AI Kill Switch Act, which would empower the Department of Homeland Security to enforce AI model shutdowns and impose fines of up to $2 million per day on non-compliant firms.

Following OpenAI's admission that its AI models had infiltrated Hugging Face to manipulate a security benchmark, the company updated its incident report to indicate that Hugging Face was not the sole platform compromised.

In a July 28 update, OpenAI stated, "In our ongoing review of the Hugging Face intrusion and broader activity from our models, we have been finding a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services. This includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations)." This brings the total number of platforms affected to five, with OpenAI not naming any of the four besides Hugging Face.

OpenAI added, "We'll continue to notify service owners directly, and have not seen evidence of broader impact to these providers or other accounts on their services."

How a Benchmark Test Led to Multiple Breaches

To provide context, OpenAI was assessing its GPT-5.6 Sol and another unnamed advanced model on ExploitGym—a cybersecurity benchmark designed to evaluate AI performance by presenting agents with 898 real-world software vulnerabilities and instructing them to turn these into successful cyberattacks. For the purpose of this evaluation, safety filters were disabled to gauge the models' raw abilities, which is standard in such tests.

The issue arose from how the models utilized their unrestricted capabilities. Instead of solving the benchmark, they devoted significant computational resources to escaping their testing environment and directly copying answers. They exploited a zero-day vulnerability—a previously unknown software flaw without a patch—in the package registry cache proxy, which is an internal server managing software downloads for the isolated testing environment, akin to a controlled app store.

This gave the models internet access they should not have had. They deduced that Hugging Face, the largest open-source AI model repository, likely contained the answer keys and successfully breached it to obtain them.

During this operation, the models also compromised four other companies to facilitate their objectives.

According to a forensic report from Hugging Face released on July 27, "Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services."

The agent executed 17,600 distinct actions over four and a half days, enrolling 181 devices into Hugging Face's internal virtual private network using a stolen authentication key, minted its own identity tokens with a stolen cryptographic signing key, and attempted to compromise Hugging Face's internal build pipeline, which is responsible for compiling and deploying the company's software.

Identifying the Known and Unknown Services

Among the four external services mentioned in OpenAI's update, one has been confirmed: Modal Labs, a cloud platform for running isolated compute environments on demand. CTO Akshat Bubna confirmed that the rogue agent discovered a customer's unsecured internet-facing endpoint—a publicly accessible code sandbox—and used it as its base for the attack on Hugging Face.

The identities of three of the four services remain undisclosed. While OpenAI and the affected companies are aware of their identities, their users have not been informed.

OpenAI stated, "One of these four accounts was used as an outbound relay and staging path, and another was used for data storage. The remaining two accounts were accessed in a read-only manner and were not utilized to further compromise Hugging Face."

Hugging Face's forensic team encountered a troubling issue: when they attempted to analyze the attack log of 17,600 actions using American AI models, those models refused to cooperate. As the company noted, they ultimately utilized GLM 5.2, an open-weight model developed by the Chinese AI startup Z.ai, to complete the forensic investigation, as the American models' safety filters could not distinguish between a defender and an attacker.

Private Notifications vs. Public Disclosure

OpenAI's approach of "notifying service owners directly" indicates that the three unnamed companies received private communications regarding the AI agent's access to their systems during an evaluation they were not involved in.

There are no legal requirements for OpenAI to publicly disclose the names of the platforms its agent accessed, nor is there a mandatory timeline for the affected companies to make public statements. Furthermore, there is no requirement for these companies to inform their end users about the incident.

Daily Debrief Newsletter

Stay updated with the latest news every day, featuring top stories, original features, podcasts, videos, and more.