Timeline

Anthropic's Responsible Scaling Policy v3.1 takes effect

The update clarifies that Anthropic's automated-AI-R&D threshold means doubling aggregate capability rather than researcher productivity, and reaffirms it can pause unilaterally at any time.

  • Safety & alignment
  • Minor

Anthropic published version 3.1 of its Responsible Scaling Policy, a narrower update than the structural rewrite it made to the policy as version 3.0 five weeks earlier. Anthropic described the new version as clarificatory rather than substantive, making two specific changes rather than revising the policy’s underlying commitments.

The first change tightened the definition of the policy’s automated-AI-R&D capability threshold. Earlier language had described the trigger as being able to “compress two years of 2018–2024 AI progress into a single year,” phrasing that could be read either as doubling aggregate model capability or as doubling the productivity of individual researchers using AI tools. Version 3.1 specifies that it means the former — a threshold about how much the technology itself has advanced, not how much faster people work with it.

The second change made explicit something the policy had left implicit: that Anthropic “remain[s] free to take measures such as pausing the development of our AI systems in any circumstances in which we deem them appropriate,” independent of whether any formally defined safeguard commitment has been triggered. Anthropic paired the release with an updated Frontier Safety Roadmap marking two previously planned safety goals as complete.

The update illustrates a pattern in how frontier labs maintain safety frameworks under scrutiny: rather than a fixed document, the RSP has gone through repeated point revisions since its 2023 original, each narrowing ambiguity that could otherwise read as stricter or looser than intended once put into practice.