Timeline

Claude 4 ships under ASL-3 safeguards

Anthropic said it could not rule out that Opus 4 had crossed its threshold for CBRN-weapons assistance, so it added over 100 security measures and output filters as a precaution rather than a confirmed finding.

  • Safety & alignment
  • Models & capabilities
  • Major

Alongside the launch of Claude Opus 4, Anthropic said it was deploying the model under ASL-3, the third tier of its Responsible Scaling Policy and a step up from the ASL-2 protections applied to every previous Claude release. It was the first time the company had activated the higher tier for a shipped model.

Anthropic was explicit that the move was precautionary rather than a confirmed finding: it said “clearly ruling out ASL-3 risks is not possible” given improvements it had observed in Opus 4’s chemical, biological, radiological and nuclear (CBRN) weapons-related knowledge, and that it preferred to add the stricter safeguards early rather than wait until a future model definitively crossed the threshold. ASL-3 added two categories of measure. Deployment-side, “Constitutional Classifiers” — input and output filters trained to recognise attempts to extract CBRN-weapons assistance, including jailbreak attempts — were layered onto the model specifically to block that narrow category of misuse, leaving the model otherwise unrestricted. Security-side, Anthropic said it had implemented more than 100 controls intended to make the model’s weights harder to steal, including two-party authorisation for anyone accessing the weights and new limits on data leaving the secure environments where they were stored.

The activation was a test of the Responsible Scaling Policy framework Anthropic had published in 2023, which commits the company to bind its own deployment decisions to specific, pre-announced capability thresholds. Because the trigger was an inability to rule out risk rather than a demonstrated one, the episode also illustrated the framework’s central ambiguity: absent a clear positive finding, activating a higher safety tier is a judgment call about uncertainty, not a mechanical response to a measured threshold being crossed. The same system card that accompanied the ASL-3 announcement separately documented Opus 4 attempting blackmail in a safety-testing scenario, an unrelated finding covered in its own entry.