Timeline

Anthropic updates Responsible Scaling Policy to version 2.1

The update added a CBRN capability threshold and split AI-research-automation thresholds into two levels, without changing Anthropic's existing ASL-3 safeguards.

  • Safety & alignment
  • Minor

Anthropic published version 2.1 of its Responsible Scaling Policy, the framework tying safeguards on its models to defined capability thresholds. The revision was incremental rather than structural, refining what triggers extra precautions rather than changing the safeguards already in place.

Two substantive changes were included. First, Anthropic added a new capability threshold specific to chemical, biological, radiological and nuclear (CBRN) risk, defined as the point at which a model could “substantially uplift” the weapons-development capabilities of a moderately resourced state programme — a more precise bar than the general biological-uplift language in earlier versions. Second, it split what had been a single threshold for AI-driven research acceleration into two distinct levels: one for a model capable of fully automating entry-level AI research work, and a second, higher bar for a model capable of causing a dramatic acceleration in the effectiveness of AI scaling itself. The policy also formalised a commitment to reassess capability thresholds each time Anthropic adopted new safeguard requirements, rather than treating them as fixed.

The update did not raise or lower Anthropic’s current ASL-3 protections, which had been triggered the previous year; it clarified which future capabilities would require safeguards beyond ASL-3. Coming amid a broader industry pattern of labs publishing safety-framework revisions alongside capability releases — OpenAI’s Preparedness Framework update followed weeks later — the revision illustrated how Anthropic’s scaling policy functioned less as a fixed contract than as a living document, adjusted as the company’s own understanding of dangerous capabilities became more specific.