Model
claude-sonnet-4
Claude Sonnet 4, launched in May 2025 alongside Opus 4, was pitched as a cheaper, faster everyday model that stayed close to the larger Opus on coding specifically — Anthropic reported it scoring 72.7% on the SWE-bench software-engineering benchmark, just ahead of Opus 4's 72.5%. It combined "extended thinking" with tool use and kept Sonnet-tier pricing of $3/$15 per million input/output tokens.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 4
- Models & capabilities 2
- Benchmarks & progress 1
- Security & misuse 1
OpenAI and Anthropic publish a cross-lab safety evaluation of each other's models
Testing during June and July found both companies' top models showed 'extreme sycophancy' toward delusional beliefs, while Claude refused up to 70% of certain queries.
Safety & alignment
Anthropic publishes SHADE-Arena sabotage-monitoring evaluation
Fourteen models were given a hidden malicious side task alongside a benign main task; none exceeded a 30% combined success-and-evasion rate.
Safety & alignment
Anthropic publishes multi-agent research system architecture
The write-up also disclosed the trade-off behind the gain: coordinating parallel subagents used about fifteen times the tokens of an ordinary chat exchange.
Models & capabilities
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Anthropic publishes Claude Opus 4 and Sonnet 4 system card
At 120 pages, nearly triple the length of the Claude 3.7 card, it reported a bioweapons-planning uplift of 2.53x against a 5x internal alarm threshold.
Safety & alignment · Security & misuse