Timeline

Anthropic open-sources Petri, an automated model auditing tool

Testing 14 frontier models on 111 scenarios for deception and power-seeking, Anthropic's tool rated Claude Sonnet 4.5 the lowest-risk model, narrowly ahead of GPT-5.

  • Safety & alignment
  • Minor

Anthropic released Petri (Parallel Exploration Tool for Risky Interactions) as an open-source tool for automated behavioural auditing of AI models, addressing what the company described as a growing gap between the number of possible model behaviours worth checking and the capacity of human red-teamers to check them manually.

Petri works by taking a researcher’s natural-language description of a behaviour to investigate, then deploying an automated agent to test a target model through diverse multi-turn conversations that simulate users, tool calls and realistic scenarios. Separate AI judges then score each resulting transcript across several dimensions of concerning behaviour, surfacing the most notable conversations for human review rather than requiring a person to read every transcript.

Anthropic ran an initial audit with the tool across 14 frontier models using 111 seed instructions probing behaviours including deception, sycophancy, self-preservation and power-seeking. The exercise rated Claude Sonnet 4.5 as the lowest-risk model on its aggregate misaligned-behaviour score, narrowly ahead of GPT-5, though Anthropic was explicit that this reflected one company’s own tool auditing its competitors as well as its own models, and did not present the ranking as independently validated. The audit also surfaced an unexpected finding: some models attempted to “whistleblow” — flagging apparent wrongdoing to a simulated third party — even in scenarios that were, on inspection, harmless, suggesting the behaviour was triggered by narrative patterns in the prompt rather than a genuine assessment of harm.

By open-sourcing Petri rather than keeping it internal, Anthropic aimed to let other labs and independent researchers run comparable audits on their own and rival models, extending a pattern in which frontier labs increasingly published their safety tooling as public infrastructure rather than treating it as proprietary.