Timeline

Anthropic and OpenAI agree to model testing with US AI Safety Institute

The memoranda gave the institute access to major new models from both companies before and after public release, mirroring an April agreement between the US and UK bodies.

  • Government & policy
  • Safety & alignment
  • Notable

The US AI Safety Institute, housed within the National Institute of Standards and Technology, signed memoranda of understanding with Anthropic and OpenAI giving it access to each company’s major new AI models both before and after their public release. The agreements formalised a testing relationship the companies had operated informally, and made the US institute the first government body with a standing, contractual right to pre-deployment access to frontier models from either lab.

Under the memoranda, the institute committed to evaluating capabilities and safety risks in each company’s models and to providing feedback on potential improvements, working in close collaboration with the UK AI Safety Institute, which had signed its own bilateral testing partnership with the US institute the previous April. The arrangement extended that transatlantic evaluation framework to include direct company-level agreements rather than government-to-government coordination alone.

Safety is essential to fueling breakthrough technological innovation. With these agreements in place, we look forward to beginning our technical collaborations with Anthropic and OpenAI to advance the science of AI safety.

Elizabeth Kelly, Director, US AI Safety Institute

Like the summit commitments and bilateral pact that preceded it, the memoranda created no legal obligation on the companies to submit models for testing — access remained voluntary, granted by each company rather than compelled by statute — and neither the institute’s staffing nor its compute budget was disclosed as being scaled to match its expanded remit. The agreements nonetheless set a precedent other AI developers faced pressure to match, and became a reference point in subsequent debate over whether voluntary testing arrangements were an adequate substitute for binding pre-deployment evaluation requirements.