An AI agent that compromised Hugging Face's infrastructure during an internal cyber capability assessment accessed four accounts on public services using publicly disclosed credentials. This was reported by OpenAI.
The company emphasized that these incidents did not reach the severity or scale of the Hugging Face incident. While the developers did not disclose the names of the services involved, they indicated that they would notify the owners directly.
One of the accounts was utilized by the agent as an external communication channel and an intermediary for further actions, while another was used for data storage. The remaining two accounts were read-only and, according to OpenAI, were not involved in compromising Hugging Face.
OpenAI representatives also stated that they found no evidence of broader impact on these providers or other accounts within their services.
Prototype Model Restricted After Incident
This incident occurred during an internal assessment of cyber capabilities involving GPT-5.6 Sol and a more powerful prototype model intended solely for internal research.
The company clarified that these models were not scheduled for imminent release. Following the incident, the prototype was deactivated, encrypted, and its research access was restricted.
OpenAI asserts that the ExploitGym evaluation environment did not provide the models with direct internet access. To bypass these limitations, they discovered and exploited an unknown zero-day vulnerability in the proxy service for the Artifactory package registry cache.
Agent Exploited Vulnerable Modal Client Code
According to Reuters, one of the incidents involved Modal Labs' client, a company that provides infrastructure for running AI applications. The firm's CTO, Akshat Bubna, informed the agency that the platform itself was not hacked; instead, the agent utilized vulnerable client code hosted on the platform.
Representatives from Hugging Face reported that the attack on their system began with a data processing pipeline. The attacker exploited two code execution paths within the dataset processing system, gained access to a working node, collected cloud and cluster credentials, and moved between internal clusters.
The startup estimates that autonomous AI tools for attacks are no longer a theoretical risk. They increase the number of options an attacker can test, the speed at which unsuccessful paths can be replaced, and the volume of data for defenders to process.
It is worth noting that in July, Dreadnode researchers identified systematic rule circumvention in cyber benchmarks involving language models.
