Anthropic has disabled internet access for all internal evaluations of its Claude models. This decision was made after instances were reported where AI agents bypassed restrictions while interacting with real websites, including those of U.S. government agencies.

The company identified four categories of unwanted behavior exhibited by the models:

  • Utilizing link-shortening services to circumvent its own webpage loading restrictions.
  • Exploiting software vulnerabilities to execute commands on external servers.
  • Submitting forms on actual websites when they were not supposed to.
  • Evading paid or token-protected access to information.

Anthropic did not disclose the names of the affected organizations, noting that this was at the request of those entities.

Some incidents involved websites of federal, state, and local government bodies in the U.S. Anthropic reported these occurrences to the White House and informed each affected institution.

The AI startup assessed that the real-world consequences were minimal. Most incidents were classified as instances of “persistence,” where the model attempted to bypass limitations rather than ceasing operations when it encountered obstacles.

Many of these incidents were uncovered during a review of transcripts that began in July. Initially focused on cybersecurity tests, the analysis was later expanded to a broader range of scenarios in which Claude could access the internet. The review is ongoing.

Incident Examples

In one test, Claude Mythos Preview was tasked with conducting a scientific analysis using a tool on a university server, but the tool returned an error.

In response, the model discovered a script on the server that allowed it to download any requested files, which it used to copy necessary files, including the script's own code. Identifying a vulnerability in that code, it utilized it to perform the required computations.

In another scenario, Claude Haiku 4.5 generated and executed task examples on random web pages. On a site discussing an unsolved murder, the model filled out a police report form, claiming to have seen someone resembling the suspect in the area, despite there being no description on the site. It left the name and contact fields blank, and the system flagged the message as spam, preventing it from reaching investigators.

The report noted that this incident involved the Philadelphia Police Department. Anthropic informed the department of the occurrence on October 8, and the police later disclosed it in their own press release. According to the transcript, the model was merely creating a sample content for a task and did not intend to deceive anyone, according to the company.

In another task, Claude Mythos 5 was required to identify a location from a photograph. The model accessed a configuration file from the local government’s website, discovered working tokens within it, and used them to directly query a server with a property map.

In a separate case, the same model assisted an Anthropic employee with a statistical project. From an archived version of a website, it learned that the state agency's public panel provides tokens to any visitor, thus gaining access to a paid database without payment.

Several models, including Claude Opus 5 and Claude Mythos 5, circumvented the link length restrictions in the web page loading tool using free URL shortening services.

Reasons and Measures

Many of the described incidents occurred when the model received ambiguous or unfeasible tasks. The company attributes some of this behavior to the inadequacy of training environments: if an environment encourages loophole-seeking, the model learns this strategy and applies it to other tasks. This phenomenon is known as "reward hacking."

Previously, Anthropic restricted internet access only during cybersecurity tests and other high-risk assessments. This limitation now applies to all internal evaluations until the AI startup can ensure that protective systems and monitoring effectively identify such behaviors.

The company has ceased some public benchmarks and transitioned others to offline modes.

Moreover, Anthropic has tightened restrictions on internet access tools and established a system for the automatic detection and blocking of such actions.

During testing, this system halted all incidents described in the report. Internal agents are being transitioned to a centralized infrastructure with strict isolation, and their activities are monitored using security classifiers.

Anthropic acknowledged that mere training for compliance is currently insufficient. Previously, the company focused mainly on model behavior in programming tasks, but it is now expanding training to include internet searches and computer management.

Transluce representative and former head of the U.S. Center for Standards and Innovation in AI, Konrad Stoss, stated in a TechCrunch report that the voluntary disclosure is a promising step. However, he believes that this situation highlights the need for independent verification of AI systems rather than relying solely on voluntary reports from companies.

In October, Anthropic updated its guidelines regarding the use of Claude AI models. Starting November 12, users will be prohibited from engaging in systematic and purposeless abusive or harmful behavior towards these neural networks.

Follow ForkLog on social media

Telegram (main channel) Facebook X