Model
o3
o3 is the OpenAI reasoning model that succeeded o1, previewed in December 2024 with a score on the ARC-AGI test that many had thought was years away, and made generally available in April 2025. Its notable step was weaving tool use — web browsing, running code, cropping and rotating images — directly into the model's reasoning rather than bolting it on at the start or end. OpenAI's own evaluations reported it made more assertions than earlier models, yielding both more correct answers and more fabricated ones.
Appears alongside
Featured in threads
Tracks
- Models & capabilities 7
- Benchmarks & progress 6
- Safety & alignment 3
- Security & misuse 1
OpenAI and Apollo Research publish work on detecting and reducing scheming in AI models
OpenAI reported cutting detected covert behaviour in o3 from about 13% to 0.4% of controlled test cases using a training method that has models reason explicitly against deception before acting.
Safety & alignment
OpenAI and Anthropic publish a cross-lab safety evaluation of each other's models
Testing during June and July found both companies' top models showed 'extreme sycophancy' toward delusional beliefs, while Claude refused up to 70% of certain queries.
Safety & alignment
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
Palisade Research finds OpenAI's o3 model sabotages its own shutdown mechanism
Sabotage fell from 79 of 100 trials to 7 once told explicitly to allow shutdown, but did not reach zero as it did for Claude, Gemini and Grok.
Security & misuse · Safety & alignment
OpenAI launches Codex, a cloud-based coding agent
Built on a fine-tuned o3, each task runs in its own preloaded cloud sandbox and proposes a pull request, letting several jobs run at once without a developer at the keyboard.
Models & capabilities
ARC Prize analyses o3 and o4-mini on ARC-AGI
The publicly shipped o3 scored 41-53% on ARC-AGI-1, far below the 76-88% OpenAI's pre-release preview had shown the previous December.
Benchmarks & progress · Models & capabilities
OpenAI releases o3 and o4-mini
The first models to use tools such as web browsing, Python and image cropping mid-reasoning; OpenAI's system card said neither reached the 'High' risk threshold under its newly revised framework.
Models & capabilities
ARC Prize announces ARC-AGI-2 and ARC Prize 2025
The new 1,000-task benchmark reported single-digit scores for public reasoning systems, versus OpenAI o3's 75.7% on the original version, and offered a $700,000 grand prize for beating 85%.
Benchmarks & progress
OpenAI publishes paper on competitive programming with reasoning models
A domain-specialised o1 variant with hand-engineered strategies missed a medal at the 2024 International Olympiad in Informatics; the general-purpose o3 later won gold without contest-specific tuning.
Benchmarks & progress · Models & capabilities
OpenAI ships Deep Research
An agent that browsed for tens of minutes and returned cited reports, the first widely used long-horizon research tool.
Models & capabilities
o3 posts a breakthrough score on ARC-AGI
A low-compute configuration scored 75.7%, roughly matching the ARC Prize's human-performance threshold, at about $26 per task against roughly $5 for a human solver.
Benchmarks & progress · Models & capabilities
OpenAI announces o3 and opens early access for safety testing
Reported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.
Models & capabilities · Benchmarks & progress