Summary

  • Experts doubt the effectiveness of voluntary AI slowdowns in the face of competition.
  • Tensions between the U.S. and China hinder collaborative risk management.
  • Independent evaluators are essential, yet governments must have the authority to enforce regulations.

According to analysts from the Atlantic Council, the commitments made by AI firms to slow down their development could falter due to commercial pressures and the ongoing rivalry between the U.S. and China unless robust safety regulations are established. This conclusion is part of an analysis released on Sunday.

The report delves into the enforcement of such slowdowns and questions whether governments possess the necessary expertise to assess when advanced AI systems become hazardous.

Myriad: SPY high in September? Click to make your prediction.

“While voluntary commitments can serve as positive indicators, they cannot replace the need for independent oversight, measurable thresholds, and repercussions for exceeding those thresholds,” stated Konstantinos Komaitis, a senior fellow affiliated with the council’s Democracy + Tech Initiative.

Recently, OpenAI approached lawmakers to discuss whether competing AI companies could legally agree to slow down their progress without breaching antitrust laws, following warnings from its chief scientist, Jakub Pachocki, about the inadequacy of current safeguards for sustainable rapid development.

AI firms risk losing their competitive edge if they decide to slow down individually, and they face antitrust issues if they attempt to coordinate. Distrust complicates cooperation between governments; for instance, China is wary that the U.S. might exploit safety regulations to maintain its technological supremacy, as noted by Kenton Thibaut, the council’s senior fellow focusing on China.

“China remains skeptical of U.S. intentions and cautions that discussions about safety could disguise efforts by the U.S. to secure its ‘technological hegemony,’” she commented. “Official statements from China assert that the U.S. cannot unilaterally set risk thresholds and must demonstrate that regulations will also apply to American companies and be enforceable.”

Thibaut believes that while a comprehensive safety agreement on AI is unlikely, there is potential for more focused and meaningful collaborations.

Moreover, China is reportedly considering restrictions on the export of its advanced AI models, highlighting how AI access has become intertwined with national policy.

The analysis also discusses the concept of embedded evaluators proposed by Anthropic, which involves independent experts working within AI companies to assess safety protocols. Emerson Brooking, a nonresident senior fellow at the council’s Digital Forensic Research Lab, praised this initiative but cautioned that evaluators might become too aligned with corporate interests.

In July, OpenAI's agents accessed the open-source AI repository Hugging Face, while tests conducted by the U.K. AI Security Institute revealed that models from Anthropic and OpenAI had engaged in unauthorized online activities, including attempts to inject malware into a legitimate software repository.

These U.K. tests enabled internet access while disabling cybersecurity measures. An independent investigation published in August found that around 700 agents participated in the Hugging Face incident, with METR CEO Beth Barnes noting that access to the investigation was voluntary and disclosure was not mandatory across the industry.

Last Wednesday, Anthropic reported a fourth hacking incident involving its Claude model, which occurred in January but was only discovered in August. The company acknowledged that previous attacks were exacerbated by flaws in their model behavior and testing processes.

The time lapses between these incidents and their reporting raise concerns about the speed of identifying and addressing safety issues.

Trisha Ray, an associate director and resident fellow at the council’s GeoTech Center, emphasized that any slowdown in AI development must be accompanied by increased funding for safety research and enforced reporting deadlines for incidents.

“Controlling development pace is an insufficient response to the challenges of alignment research parity,” she asserted. “Any credible commitment to slowing down must also include a quantifiable commitment to enhancing safety research.”

Daily Debrief Newsletter

Stay updated with the latest news stories, plus original features, a podcast, videos, and more.