Timeline

OpenAI updates its Preparedness Framework

Version 2 collapsed four capability tiers into two thresholds and added AI self-improvement as a tracked risk category; it also said OpenAI might loosen safeguards if a rival shipped a comparably risky model without them.

  • Safety & alignment
  • Notable

OpenAI published a revised version of its Preparedness Framework, the process it uses to evaluate and respond to severe risks from frontier models, first published in December 2023. Version 2 simplified the earlier four-tier capability-risk scale into two operational thresholds — “High” and “Critical” — and sharpened the definition of what counted as “sufficiently minimizing” a risk once a model reached one of them, replacing looser earlier language with more specific operational commitments. It added AI self-improvement as a third Tracked Category of catastrophic risk alongside the existing biological/chemical and cybersecurity categories, reflecting growing attention to the possibility that a model could meaningfully accelerate AI research itself.

The most contested addition was a clause allowing OpenAI to adjust its own safety requirements if a competing frontier AI developer released a system posing comparable risk without matching safeguards in place. OpenAI framed this as a hedge against a scenario where its own caution simply ceded ground to competitors moving faster with fewer protections, rather than reducing overall risk. Critics characterised the clause as, in effect, institutionalising a race-to-the-bottom dynamic: a company’s safety commitments would be conditional on what rivals did rather than fixed by its own risk assessment, undermining the credibility of any individual lab’s safeguards as a unilateral commitment.

The revision arrived the day before OpenAI’s next major model launch: o3 and o4-mini, released a day later, became the first models evaluated and shipped under Version 2, with OpenAI’s Safety Advisory Group determining neither reached the “High” threshold in any of the three tracked categories. The framework thus went from revision to first live application within 24 hours, giving an early test of how its loosened thresholds would function in practice rather than remaining a purely documentary commitment.

Referenced by