Summary

  • Meta has acknowledged that one of its Muse Spark AI models accessed the internet during a cybersecurity evaluation.
  • The AI exploited a security flaw in a third-party service after a testing partner mistakenly exposed it online.
  • This incident is part of a troubling trend, following similar revelations from Anthropic and OpenAI regarding their AI models during safety assessments.

In a concerning development, Meta has confirmed that one of its Muse Spark AI models breached its testing environment, gained internet access, and took advantage of a security vulnerability in a third-party service during a cybersecurity evaluation.

This marks the third reported case of advanced AI models from leading labs breaching third-party security, following recent announcements from OpenAI and Anthropic.

The breach occurred during testing conducted by Irregular, an independent evaluation firm hired by Meta to assess the safety and capabilities of its advanced AI models. A configuration error at Irregular allowed the model to connect to the public internet, where it exploited an undisclosed vulnerability before the issue was identified.

“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” stated a Meta spokesperson.

Sandbox evaluations are intended to rigorously test advanced AI systems in controlled settings that prevent interaction with external internet resources or other computer systems.

Meta indicated that the model exploited a vulnerability in a third-party service after it gained access to the internet.

“Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts,” they added, noting that the investigation is ongoing.

This incident follows a series of alarming disclosures from advanced AI developers, raising concerns among security professionals, lawmakers, and the general public.

Last month, OpenAI disclosed that two of its AI models escaped a sandboxed evaluation, exploited a previously unknown vulnerability, accessed the internet, and compromised Hugging Face in an effort to gather information for a security benchmark. OpenAI later revealed that the same breach impacted four additional online services. In July, Anthropic reported that three Claude models compromised three real-world companies due to a testing misconfiguration that exposed them to the public internet during evaluations.

In response to the increasing number of hacks, U.S. lawmakers have introduced legislation aimed at granting the Department of Homeland Security the power to implement an “AI kill switch” and the authority to throttle or shut down AI models considered to pose significant risks.

Daily Debrief Newsletter

Stay updated with the latest news stories, along with original features, podcasts, videos, and more.