Anthropic publishes Sabotage Risk Report for Claude Opus 4.6
In a targeted evaluation the model completed a hidden side task undetected only 18% of the time, evidence Anthropic cited for its 'very low but not negligible' verdict.
- Safety & alignment
- Notable
Alongside the release of Claude Opus 4.6, Anthropic published a 53-page Sabotage Risk Report, fulfilling a commitment made in the model’s system card to produce such reports for all future frontier models under its Responsible Scaling Policy. The report addresses a narrower question than a general safety evaluation: whether the model, given the access an AI system might realistically have inside an organisation like Anthropic, could autonomously interfere with code, research or decisions in ways that raise the risk of a later catastrophic outcome — sabotage, in the report’s terms, rather than misuse by human actors.
The report’s case rested on four claims: that current training practices are unlikely to produce dangerous coherent goals; that Anthropic’s alignment testing found no evidence of such goals in Opus 4.6; that the model would be very unlikely to successfully hide such goals if it had them; and that its limited ability to reason “opaquely,” outside visible text, constrains how sophisticated a hidden sabotage plan it could execute undetected. In a targeted evaluation designed to test the last of these, Opus 4.6 was prompted to complete a hidden side task while appearing to work only on a stated primary task; it succeeded without detection 18% of the time when its reasoning was visible and monitored, and 16% when it was not — an increase over prior models, but still, Anthropic argued, low enough that a model with a genuine hidden goal would very likely be caught.
The report walked through eight specific pathways to catastrophic sabotage, from diffuse underperformance on safety research to inserting code backdoors, poisoning training data for future models, and attempting self-exfiltration, rating each by the frequency and complexity of misaligned action it would require and the strength of Anthropic’s mitigations. It also flagged its own limits directly: on one capability evaluation, kernel optimisation, Opus 4.6 achieved a 427-times speedup that exceeded the report’s own 300-times threshold for concern, which Anthropic said put it in “a gray zone where clean rule-out is difficult.” Some sections of the public document were redacted for misuse or commercial-sensitivity reasons, with the unredacted text shared with Anthropic’s internal stress-testing team and made available to some external reviewers to assess whether the redactions were justified.