Timeline

Microsoft publishes Frontier Governance Framework

The framework tracks CBRN, offensive cyberoperations and advanced-autonomy capabilities using benchmarks that best-performing models score below 70% on, plus a 10^26 FLOP compute trigger.

  • Safety & alignment
  • Minor

Microsoft published version 1 of its Frontier Governance Framework, a document setting out how it monitors its most capable AI models for high-risk capabilities and mitigates them before deployment. The framework grew out of the voluntary Frontier AI Safety Commitments that Microsoft and fifteen other labs made in May 2024, and tracks three categories: models providing significant uplift toward chemical, biological, radiological or nuclear weapons; significant uplift for highly disruptive offensive cyberoperations, including against critical infrastructure; and “advanced autonomy,” meaning a model’s ability to complete expert-level tasks, including AI research and development, without human oversight.

The framework describes a two-stage process. A “leading indicator” assessment, run on any model Microsoft is optimising for frontier capabilities or that is pre-trained with more than 10²⁶ floating-point operations, screens for advanced general-purpose capabilities — reasoning, scientific and mathematical ability, long-context handling, spatial understanding, autonomous tool use and software engineering — using benchmarks Microsoft says must be unsaturated, meaning the best current models typically score below 70% on them. A model that trips this screen undergoes a deeper capability assessment, including adversarial testing, classified against low, medium, high or critical risk levels, which then determines what security and safety mitigations apply before release.

The framework explicitly covers only these tracked capability risks and states it sits alongside Microsoft’s broader AI governance programme, which handles the wider set of contextual, use-case-dependent risks such as bias and misuse that vary by deployment and jurisdiction. It followed similar published frameworks from Anthropic, Google DeepMind, OpenAI and Meta, and Microsoft described it as the first version of a document it expected to revise substantially as evaluation practice matured.