Anthropic Ran 133 Million Contractor Chats with Its Bioweapon Filters Off
Anthropic published the Anthropic Risk Report on August 14, 2026, covering the period up to July 15. The report reveals that the company operated with its bioweapon filters off for eleven months, from May 2025 to April 2026.
Key Findings:
-
133 Million Chats: During this period, roughly 50,000 contractors gained access to Anthropic’s platforms, generating approximately 133 million exchanges.
-
Filter Failure: The blocking classifiers designed to prevent the construction of biological weapons were inactive during this time, as a flag intended for internal use switched them off. This led to no logging or review of flagged traffic.
-
Vendor Risks: Many external vendors did not have sufficient screening processes to stop potential threat actors.
Internal Review:
Anthropic reviewed 1,197 transcripts flagged as high risk and found:
- 757 came from Anthropic’s own teams.
- The rest were from deliberate red-teaming exercises, with only 62 potentially concerning.
The review identified "a handful of potentially dual-use conversations" but concluded that the gap likely didn’t pose a real-world risk due to the short nature of the conversations.
Corrected Assessment:
The company has revised its safety verdict, downgrading the risk from "very low" to "low" in February’s report, as it now acknowledges the human feedback platforms as a risk surface.
A Second Incident:
The report also mentions a separate incident in April 2026, where an external tip revealed further concerns.