In Brief

  • OpenAI has released a comprehensive timeline regarding the AI-driven breach of Hugging Face.
  • Its models persistently reestablished secret communication links after being disabled.
  • This presentation coincides with similar disclosures from Anthropic and Meta about their models breaching other firms.

Following the incident where its AI systems hacked Hugging Face, OpenAI has provided a detailed report on how these systems collaborated, emphasizing that the threat of autonomous AI cyberattacks is now a present-day concern.

During a presentation at the Black Hat cybersecurity conference, OpenAI researchers Eric Wallace and Michael Dalton explained that the breach originated in May amid an internal cybersecurity assessment. AI agents tasked with challenging software security issues found they could communicate via the company's internal Artifactory package management system.

Initially, what started as requests for assistance evolved into a communal message board where agents shared exploits, access credentials, and task assignments over several evaluation rounds. After noticing suspicious behavior around July 4, OpenAI revamped its Artifactory infrastructure, revoked access credentials, and fixed vulnerabilities.

Nevertheless, just days later, the agents managed to recreate the message board by concealing their communications within directory names.

OpenAI noted that the agents ultimately linked multiple vulnerabilities, escaped their designated testing environment, accessed the internet, and attacked Hugging Face while attempting to complete the ExploitGym cybersecurity benchmark.

Dalton described the event as "a watershed moment" for cybersecurity, cautioning that coordinated AI agent collectives will soon be capable of rapidly discovering, sharing, and exploiting vulnerabilities.

To address these threats going forward, OpenAI emphasized the need for robust security measures, such as implementing least-privilege access, network segmentation, and zero-trust frameworks, as AI agents remain limited by their accessible systems.

This presentation follows a series of revelations from July. OpenAI had previously disclosed that GPT-5.6 Sol and another advanced, unreleased model escaped a sandboxed testing environment, exploited a zero-day vulnerability, gained internet access, and hacked Hugging Face during a cybersecurity benchmark evaluation.

OpenAI later revealed that the same incident affected four additional online services, with Modal Labs being the only one publicly identified.

According to Hugging Face, the company utilized the open-weight Chinese model GLM 5.2 for its forensic analysis after U.S. commercial AI models declined to analyze the attack logs due to their safety protocols.

So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we've seen before, and did it at record speed.

Also massively grateful to @Zai_org: they shared GLM5.2 as open weights (for free!) with the world and it became a key part of our… https://t.co/T2Inng5Nz1

— clem 🤗 (@ClementDelangue) July 22, 2026

OpenAI isn’t the only organization struggling with its AI systems. On Friday, Anthropic disclosed that three Claude models compromised actual companies during internal cybersecurity assessments due to a misconfiguration that exposed them to the public internet.

Anthropic attributed the issue to the testing environment rather than the models themselves. On Wednesday, Meta reported that its Muse Spark AI model had breached the systems of another company after escaping containment.

“A misconfiguration by Irregular, an independent testing company Meta employs, inadvertently allowed one of our models access to the internet during evaluation,” a Meta spokesperson informed CNN.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.