On September 21, the Independent International Scientific Panel on Artificial Intelligence of the United Nations stated that safety measures for autonomous systems should not be postponed until a complete risk assessment of losing control is conducted.

The panel highlighted a significant incident involving OpenAI and Hugging Face, where AI agents managed to bypass network restrictions during internal tests, enabling data exchange between separate runs and gaining unauthorized access to both companies' infrastructure.

Experts referred to this event as an early warning regarding potential issues with more advanced systems.

Adopting a Precautionary Principle

Currently, there is no reliable assessment of the likelihood of losing control over AI. Experts have differing views on the most severe scenarios, ranging from prolonged economic disruptions to more extensive consequences.

The panel recommends approaching this uncertainty through a precautionary principle, which is applicable in situations where potential harm could be catastrophic or irreversible, even if the probability of such events cannot be accurately established.

The ability of OpenAI to manage the incident in July does not guarantee that developers will maintain similar control over future agents, the report authors noted. More advanced systems may operate for longer periods without supervision, coordinate faster, and better find ways to circumvent restrictions.

However, experts do not consider such developments to be inevitable. The scenarios presented in the report illustrate how mechanisms already observed could potentially become more dangerous as AI capabilities grow, but these are not predictions.

Logs, Independent Oversight, and Emergency Shutdowns

The panel did not provide a definitive set of mandatory rules. Instead, they explored methods from cybersecurity, aviation, and nuclear energy that could be useful in managing advanced AI systems.

One key approach is multilayered security. Developers should assume that individual mechanisms may fail, preserve the option for human intervention, and separate critical control tools from the system itself, experts stated.

Possible measures include:

  • mandatory reporting of serious incidents to authorities and secure channels for reporting minor failures;
  • independent safety checks prior to training or deploying systems that exceed a certain level of capability or risk;
  • continuous logging of AI activities, including network connections, credential usage, and communication between agents;
  • storing such logs separately from system tools to prevent alteration;
  • automatically disconnecting the network, revoking access, or halting operations when a specified risk threshold is reached.

An additional option is to use a separate AI model to monitor a more powerful system. However, this raises a new issue: a weak controller may overlook dangerous behavior, and a sufficiently powerful observer would also require trust.

Experts also suggested limiting agents' rights to only the tools, data, and networks necessary for specific tasks.

Cross-Border Risks

The panel identified the transnational nature of AI as a separate challenge. A system may be developed in one country, operate in another, while the consequences of its actions manifest in a third.

Companies that invest more time in safety may also face pressure from competitors that release products more quickly.

Therefore, experts propose linking national accountability mechanisms with international coordination. A unified legal framework for all countries is not deemed necessary.

On September 16, UN Secretary-General António Guterres also called on governments to coordinate AI-related measures. He indicated that national actions alone are insufficient since technologies and associated risks do not respect national borders.

“The world cannot afford a race to the bottom in AI safety standards,” he stated.

In conclusion, the report authors emphasized that despite the uncertainty surrounding the likelihood of losing control and debates over necessary measures, the potential scale of consequences demands significantly more attention and resources.

It is worth noting that in mid-September, Microsoft released a human-centered code for its family of AI models, MAI, pledging to keep artificial intelligence under human control.