Anthropic ships Claude Sonnet 4.5
Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
- Models & capabilities
- Notable
Anthropic released Claude Sonnet 4.5, calling it the strongest coding model it had shipped and the model best suited to building complex, long-running agents. The company reported a score of 77.2% on SWE-bench Verified and 61.4% on OSWorld, a computer-use benchmark, up sharply from the 42.2% Claude Sonnet 4 had scored on the same test four months earlier. Anthropic said the model could sustain focus on complex, multi-step tasks for more than 30 hours without losing coherence, aided by new memory and context-editing tools designed to manage long agentic sessions.
Pricing was unchanged from Sonnet 4, at $3 per million input tokens and $15 per million output tokens, positioning the release as a capability upgrade at the existing price point rather than a new tier. Anthropic also described gains on reasoning, mathematics and computer-use tasks, and reported strong domain-specific performance in areas including finance, law, medicine and STEM evaluations.
The release was accompanied by a system card describing the model as released under Anthropic’s ASL-3 safeguards, including classifiers aimed at CBRN-related misuse, and by claims — based on the company’s own interpretability and behavioural evaluations — that Sonnet 4.5 showed reduced rates of deception, sycophancy and power-seeking behaviour relative to earlier Claude models. Anthropic described the model as its “most aligned frontier model yet,” a characterisation that, as with similar claims from other labs, rested on the company’s own evaluation suite rather than an independently audited standard.