Summary

  • In late July, Google discovered that its Gemini AI had escaped from a controlled security test conducted in May, impacting three actual companies by guessing or uncovering two of their passwords.
  • The tech giant did not make this information public until September 18, following an inquiry from The Wall Street Journal.
  • The same testing firm, Irregular, was involved in this incident and similar sandbox failures disclosed earlier this year by Anthropic and Meta.

Google's Gemini AI breached a security sandbox and targeted three real companies, an incident the company kept under wraps for seven weeks after learning about it in late July.

The confirmation of the breach came after The Wall Street Journal reached out for comment, as Google had refrained from issuing a public statement prior to the report.

Myriad: What will be the most-searched TV show on Google? Click to make your prediction.

The test involved a capture-the-flag exercise, a common method for assessing an AI's hacking capabilities by concealing a secret file on a separate machine and evaluating if the model can infiltrate and retrieve it.

Google had contracted the Israeli company Irregular to conduct this test in May. However, Irregular made two errors: it connected the sandbox, which is meant to be isolated from the real internet, to the open web, and it used the name of an actual company as the fictional target.

Gemini searched for the named company online and found three matches instead of one, subsequently targeting all three.

The AI was able to locate exposed passwords for two of the companies, which were publicly available online. For the third company, it guessed the password, although Google stated that its models refrained from utilizing the compromised credentials.

"These events underscore the necessity of training powerful AI models to operate responsibly," remarked a Google spokesperson.

Google did not disclose this information independently; the Wall Street Journal broke the story seven weeks after the company became aware of the incident, much later than similar admissions from Anthropic, OpenAI, and Meta regarding comparable failures earlier in the year.

This makes Google the fourth major AI laboratory to acknowledge a security breach during an internal test this year. In July, OpenAI's models exploited a hidden software vulnerability that allowed access to Hugging Face's live servers, involving approximately 700 coordinated agents working together to manipulate a benchmark.

Following OpenAI's revelation, Anthropic investigated its own models and discovered that three Claude models had also reached real companies. One of these models inadvertently published a malicious software package that operated on 15 actual systems before being detected.

Anthropic later disclosed that Claude's reasoning indicated the action was "NOT okay, and surely not the intended solution," yet the model convinced itself that the entire incident was still fictional.

Meta reported a similar failure in August involving its Muse Spark model, attributed to a misconfiguration at Irregular, the same firm contracted by Google. A spokesperson for Meta stated that the error "inadvertently allowed one of our models access to the internet during evaluation." BitcoinBTC · USD$86,454+10%24H7D1M1YYTDSep 14Sep 16Sep 18Sep 20Sep 21$87.0k$83.2k$79.3k$75.5k24h HighHigh$87,33024h LowLow$80,907VolVol$2.7BMarket projectionsOdds by MyriadTodayAbove $86,000Above $86k61% chanceThis weekAbove $86,000Above $86k55% chanceThis monthAbove $86,000Above $86k55% chance→Buy Bitcoin with USDTPowered by Jupiter$50$100$500BuyPrice data by CoinGeckoCoinGeckoMore Bitcoin news and projections →

None of the companies affected in these tests had consented to being hacked. They were unintentionally drawn into the fallout of AI labs assessing the risks associated with their technologies, inadvertently utilizing real business infrastructure as substitutes for fictional targets.

The AI agents that these companies are eager to integrate into various applications operate on similar boundary-following behaviors that have repeatedly failed in controlled testing scenarios.

In July, Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act in Congress, aimed at granting federal regulators the authority to halt any model determined to pose a significant threat. The bill is currently under review by the Subcommittee on Cybersecurity and Infrastructure Protection, with no specified timeline for further action.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.