OpenAI releases GPT-5.3-Codex-Spark, an ultra-low-latency coding model
Served on Cerebras' Wafer Scale Engine 3 rather than OpenAI's usual infrastructure, the smaller model hit over 1,000 tokens per second — about 15 times the standard Codex model's speed.
Models & capabilities