DeepMind updates the Frontier Safety Framework to version 2.0
Version 2.0 added security-level tiers for its capability thresholds and, for the first time, treated a model's own deceptive alignment as a risk requiring monitoring before deployment.
- Safety & alignment
- Notable
Google DeepMind published version 2.0 of its Frontier Safety Framework, the internal system of capability thresholds it uses to decide when a model’s abilities are dangerous enough to require restricted deployment. The update followed roughly nine months of internal use of the original framework and consultation with outside experts.
Two changes stood out. First, the framework added security-level recommendations tied to each Critical Capability Level, with the strongest recommendations reserved for capabilities in machine-learning research and development — the reasoning being that a model able to meaningfully accelerate or automate AI development itself needs the strongest protection against theft or leak, since a stolen copy could compound that risk elsewhere. Second, version 2.0 introduced a three-step deployment procedure for models that trip a threshold: build mitigations, produce a formal “safety case” arguing the risk has been adequately reduced, and secure internal governance approval before general release.
The update also named deceptive alignment — the possibility that a model might deliberately act to undermine human oversight of its own behaviour — as a risk domain to be proactively monitored, including through automated checks for the kind of instrumental reasoning that could support such behaviour. DeepMind described this as an industry-leading step, since most published frontier-safety frameworks at the time addressed misuse by external actors rather than the model’s own conduct. As with the original framework, DeepMind stated the document remained exploratory and would be revised as understanding of frontier-model risk matured — a revision that arrived as version 3.0 in September 2025.