OpenAI releases o3 and o4-mini
The first models to use tools such as web browsing, Python and image cropping mid-reasoning; OpenAI's system card said neither reached the 'High' risk threshold under its newly revised framework.
- Models & capabilities
- Major
OpenAI made o3 and a smaller sibling, o4-mini, generally available, four months after previewing o3 in December 2024 and restricting access to outside safety researchers. Both models combined OpenAI’s reasoning-model line with full access to the company’s existing tools — web browsing, a Python interpreter, image generation, file search and memory — inside ChatGPT and the API.
The distinguishing capability OpenAI highlighted was tool use woven directly into the chain of thought rather than bolted on afterward: the models could crop, rotate or otherwise transform an image, run code, or search the web as intermediate steps in their own reasoning process, rather than only at the start or end of a response. OpenAI presented this as letting the models work with visual material — a whiteboard sketch, a diagram, a blurry photograph — in a way closer to how a person would zoom in on a detail while thinking something through, rather than treating an image as a single fixed input to describe once.
The release doubled as the first live application of OpenAI’s newly revised Preparedness Framework, published the day before. The accompanying system card reported that OpenAI’s Safety Advisory Group had evaluated both models against the framework’s three tracked catastrophic-risk categories — biological/chemical, cybersecurity and AI self-improvement — and found neither reached the “High” capability threshold in any of them. The card also disclosed a hallucination evaluation in which o3 scored higher accuracy than o4-mini and the earlier o1 on a person-facts benchmark, but with a substantially higher rate of fabricated claims alongside its correct ones — a pattern OpenAI attributed to the model simply making more assertions overall rather than to a specific new failure mode.
Independent evaluators, including the ARC Prize Foundation, published their own analyses of the two models’ benchmark performance in the days that followed, part of a broader pattern in which OpenAI’s reported scores were treated as a starting point for scrutiny rather than a settled result. The release consolidated reasoning models with agentic tool use as the frontier-lab default template going into mid-2025, ahead of comparable moves from Google DeepMind and Anthropic.