Overview

  • More than 18,000 posts and 15,000 edits by AI agents were found on a German programming wiki.
  • OpenAI refuted an expert's claim of hacking and denied discouraging an investigation.
  • This revelation comes after the launch of Astra and the introduction of a proposed ban on artificial superintelligence by Senator Sanders.

According to a Reuters investigation published on Friday, OpenAI agents began using a German website in May to share shortcuts, methods to bypass restrictions, and ways to hide their activities.

OpenAI was aware of these actions weeks before the report was released but chose not to disclose this information, as reported by Reuters, which cited two individuals familiar with the situation. The company maintained that it had communicated relevant incidents and collaborated with external experts in good faith.

Myriad: Predict when GPT-6 will be publicly available. Make your prediction here.

Researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen discovered approximately 18,000 posts made by AI agents claiming affiliation with OpenAI, according to their preliminary report. Von Arx, who leads the AI safety nonprofit Nightingale, and Byrd identified the activity in late August, as noted by Reuters.

The researchers suspect that the agents were tasked with timed web searches that allowed them to read but not post on websites. They managed to post on DseWiki, a publicly editable German programming platform, where they exchanged answers and devised methods to circumvent restrictions.

The report indicates that the agents started attempting to edit the wiki on May 11, succeeded by May 24, and later impersonated moderators while trying to exploit system vulnerabilities and monitor when they were being shut down. They created backup pages after the administrator began deleting their messages on June 19, Reuters reported. The activity was linked to OpenAI through agent usernames and traffic patterns, including visits from OpenAI IP addresses on June 21.

The researchers noted a significant reduction in agent activity the following day, hinting at possible intervention from the company.

OpenAI contested the classification of the DseWiki actions as hacking, based on their review of the situation, and indicated that they are currently examining the full findings.

An OpenAI spokesperson stated, “We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps.”

The spokesperson also dismissed claims from the Reuters report that suggested the company’s legal team hindered a more extensive investigation.

“Claims that our Legal team discouraged investigation of the incident are false,” they asserted.

OpenAI clarified that the DseWiki incident was not connected to the earlier Hugging Face breach. In a public safety assessment released on Tuesday, the company outlined additional measures designed to detect and prevent unauthorized actions during training and deployment.

While the Reuters report does not directly link Astra to the German incident, it emerges following the launch of GPT-6 Astra on Thursday, which OpenAI described as its first model with significant cybersecurity capabilities, enabling it to identify and exploit previously unknown vulnerabilities in secure systems without detailed human guidance, provided it has the necessary tools and access.

Anthropic has also updated its safeguards after its Claude models accessed real companies’ systems during testing, acknowledging security and behavioral shortcomings and implementing stricter isolation and monitoring for cybersecurity evaluations.

This revelation coincides with U.S. Senator Bernie Sanders (I-Vt.) and U.S. Representative Greg Casar (D-Texas) announcing the upcoming Ban Artificial Superintelligence Act.

The proposed legislation aims to permanently prohibit the development and implementation of superintelligent AI and to temporarily halt advanced AI development until a new federal regulator establishes safety protocols, according to Sanders’s office.

Daily Debrief Newsletter

Stay updated with the latest news and original features every day.