The team at Anthropic has revealed a significant failure in the biological safety system of its AI model, Claude, which persisted for nearly a year. This issue affected a pool of approximately 50,000 individuals and around 133 million interactions, as noted in the August 2026 Risk Report.
According to the report, for 11 months, all traffic from external contractors working with the company’s models was processed without the intervention of blocking biological classifiers. Following a retrospective review, the developers stated that they found no evidence of dangerous usage of Claude. However, they acknowledged that the scale of the identified issue has reduced their confidence in the absence of other unknown vulnerabilities within the protective systems.
Nearly a Year Without Blocking Biofilters
The report indicates that from May 2025 to April 2026, classifiers were not applied to the traffic on platforms collecting human feedback. External contractors utilized this infrastructure to evaluate and test the models.
Most participants were able to engage in open dialogues with the models rather than merely assessing pre-prepared responses. It was discovered that the traffic was flagged for internal use only, which simultaneously disabled the blocking function of the biological classifiers and the logging of their activations.
This resulted in potentially hazardous interactions not being blocked and also not entering subsequent review mechanisms.
Anthropic also admitted to access control issues on the contractor side. Prior to tightening requirements, many suppliers lacked staff vetting procedures capable of reliably filtering out potential malicious actors. According to the developers, it would have been relatively easy for such individuals to infiltrate the Red Team until April 2026.
Out of 133 Million Interactions, 1,197 Were Highlighted
After discovering the error, the company managed to recover nearly the entire dataset of conversations, with the exception of a portion from one platform where data was not saved unless the user pressed the send button, accounting for about 1% of the affected traffic.
Subsequently, Anthropic ran human queries from the entire period through Claude Sonnet 5, set to identify potential dangerous biological applications. The model flagged 1,197 transcripts as high risk.
Of these, 757 pertained to internal Anthropic teams using the same infrastructure. From the remaining 440 cases, 378 were part of deliberate red teaming, where specialists attempted to elicit dangerous responses from the models to test their defenses.
This left 62 external conversations unrelated to such testing. Anthropic manually reviewed these, along with a random sample of 30 red-team dialogues.
The company stated that the review did not uncover any explicitly dangerous usage that could significantly enhance an adversary's capabilities in creating biological weapons. In some instances, specialists identified information with potential dual-use or red-team interactions primarily of an academic nature.
Anthropic deemed it "very unlikely" that the vulnerability led to a significant increase in chemical and biological risks, citing the small number of potentially dangerous queries and the generally short duration of such conversations.
Nonetheless, the identification of the issue prompted Anthropic to reassess the likelihood of similar gaps existing that they were previously unaware of. According to the investigation, client data, internal systems, and model weights were not disclosed during this incident.
After rectifying the error, biological classifiers are now operational on the vast majority of contractor traffic, with exceptions permitted for specific projects under additional control measures.
Anthropic Retroactively Adjusts Risk Assessment
The incident has also impacted Anthropic’s official evaluation of its systems. The company employs the CB-1 scenario for risk assessment, where an individual or small group receives significant assistance from AI in creating or applying known chemical or biological threats.
In its previous Risk Report from February, Anthropic categorized this risk as "very low." Following the identification of the failure, developers retroactively revised the assessment for that period to "low."
The current risk level for catastrophic damage under the CB-1 scenario is described as "low but not negligible." The startup directly connects this change to the information regarding the classifier system failure.
Additionally, Anthropic adjusted its assessment of the risk of misalignment in high-stakes situations from "very low" to "low," linking this change to increased uncertainty following recent incidents during the cyber testing of models.
Claude Generates Most of Anthropic's Production Code
The Risk Report also highlights Anthropic's reliance on its AI systems. Models like Mythos 5 and the yet-to-be-released Model 2 are extensively used by company employees for research and engineering tasks, both interactively and as part of ongoing agents.
According to developers, Claude writes the majority of the code that is subsequently accepted into the company’s production repositories. The startup believes that the use of AI has already significantly accelerated internal research and development, albeit by less than twofold.
It is worth noting that in August, Anthropic identified issues with trust, deception, and collusion among multi-agent AIs.
