Timeline

OpenAI partners with Cerebras for 750MW of compute

The multi-year, reportedly $10bn-plus deal covers wafer-scale chips for fast inference — not model training — deployed in phases through 2028.

  • Compute & infrastructure
  • Notable

OpenAI and Cerebras Systems announced a multi-year agreement to deploy 750 megawatts of Cerebras’s wafer-scale AI chips for OpenAI’s inference workloads — the compute used to serve responses to users, rather than to train new models. Reporting put the value of the deal above $10 billion, with capacity coming online in phases through 2028.

Cerebras chips integrate compute, memory and interconnect on a single, dinner-plate-sized wafer rather than the smaller dies used in GPUs, which the companies said let them return chatbot and coding-agent responses up to fifteen times faster than GPU-based inference systems. OpenAI framed the deal as aimed at low-latency use cases — coding agents and voice interaction in particular — where response speed affects usability more directly than in offline or batch workloads, rather than as a substitute for the Nvidia and AMD GPU capacity underpinning its training runs.

For Cerebras, a company that had built its business selling wafer-scale hardware as a niche alternative to Nvidia, a multi-year commitment of this size from OpenAI — one of the largest buyers of AI compute in the world — represented a significant validation of the wafer-scale approach and diversified OpenAI’s compute supply chain beyond its existing large commitments to Nvidia, AMD, Microsoft, Oracle and others, part of a broader pattern in which OpenAI spent 2025 and early 2026 assembling compute capacity from as many suppliers as it could sign.