OpenAI ships GPT-5.2-Codex
OpenAI reported an 'unmatched' 56.4% on the SWE-Bench Pro benchmark and 64% on Terminal-Bench 2.0, alongside new defensive-cybersecurity capabilities.
- Models & capabilities
- Minor
OpenAI released GPT-5.2-Codex, a coding-specialised variant of GPT-5.2 built for long-horizon, agentic software-engineering work rather than general chat. OpenAI said the model held context better across extended sessions through a technique it called context compaction, letting it continue large refactors, code migrations and multi-file feature builds without losing track of earlier steps, and reported markedly improved reliability working in native Windows environments — building on ground covered by the earlier GPT-5.1-Codex-Max.
The company reported what it called an unmatched score of 56.4% on SWE-Bench Pro, a harder successor to the widely used SWE-bench Verified benchmark, and 64% on Terminal-Bench 2.0, which tests agentic performance in realistic terminal environments. It also added defensive-cybersecurity capabilities, framed around AI-assisted vulnerability discovery, testing and mitigation, and improved handling of visual inputs such as screenshots and technical diagrams.
The release continued OpenAI’s practice of splitting its flagship line into a general-purpose model and a coding-optimised sibling, following the same pattern set by GPT-5.1 and GPT-5.1-Codex-Max. As with other self-reported frontier benchmark figures, SWE-Bench Pro and Terminal-Bench 2.0 scores were published by OpenAI without independent verification, and the model shipped alongside a system-card addendum covering its cybersecurity capabilities specifically.