OpenAI releases GPT-5.3-Codex
OpenAI reported the model roughly doubled its predecessor's OSWorld-Verified computer-use score, from 38.2% to 64.7%, and was the first Codex model rated 'High capability' for cybersecurity tasks.
- Models & capabilities
- Minor
OpenAI released GPT-5.3-Codex, a coding model built for agent-style development workflows in which the model uses tools, operates a computer and completes longer tasks with less supervision, following GPT-5.2-Codex by around three months and arriving three days after the standalone Codex desktop app. It launched across the Codex app, CLI, IDE extension and web interface for paid ChatGPT plans, with API access said to be coming later; OpenAI described it as 25% faster than its predecessor.
OpenAI reported GPT-5.3-Codex scoring 77.3% on Terminal-Bench 2.0, up from 64.0% for GPT-5.2-Codex, and 64.7% on OSWorld-Verified, a computer-use benchmark, against 38.2% for the prior model — roughly doubling that score. On SWE-Bench Pro the gain was marginal, 56.8% against 56.4%. On a cybersecurity capture-the-flag benchmark it scored 77.6%, against 67.4% previously, which OpenAI said made it the first Codex model it classified as having “High capability” for cybersecurity tasks under its preparedness framework.
The release continued a pattern of frequent, incremental Codex updates from OpenAI through early 2026, with gains concentrated in computer-use and terminal-operation tasks rather than in the code-correctness benchmarks, such as SWE-Bench, that had shown diminishing returns from one release to the next.