Anthropic issues a pilot sabotage risk report for Claude
Reviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
- Safety & alignment
- Minor
Anthropic published a pilot sabotage risk report for Claude Opus 4, a document intended to make an affirmative case — rather than a checklist of passed tests — that the model does not pose an unacceptable risk of misaligned, autonomous sabotage of tasks, code or internal company decisions. Anthropic described it as the first complete attempt by any frontier developer to document this kind of case, fulfilling a commitment made under its Responsible Scaling Policy.
The assessment was drafted over the summer of 2025, when Opus 4 was Anthropic’s most capable released model, and went through two rounds of review beyond the authoring team: an internal check by Anthropic’s Alignment Stress-Testing group, and an external one by METR, the independent evaluation nonprofit, both of which had access to supporting evidence not included in the published report. Anthropic said it found the risk of sabotage from Opus 4 “very low, but not completely negligible,” and expressed moderate confidence that the model lacks both coherent goals that would motivate sabotage and the capability to carry it out undetected. Both review teams endorsed the overall conclusion while flagging weaknesses in some of the supporting argumentation and evidence.
Anthropic was explicit that the report was a pilot rather than a finished methodology, and said future versions would need additional safeguards and more rigorous reasoning as models grow more capable. The exercise sits alongside Anthropic’s other Responsible Scaling Policy outputs as an attempt to make pre-deployment risk judgments auditable rather than purely internal, at a point when frontier labs faced growing pressure to show their safety claims could withstand outside scrutiny.