Model
claude-3.7-sonnet
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 1
- Money & business 1
- Models & capabilities 1
ETH Zurich's 'Proof or Bluff?' finds reasoning models fail proof-based USAMO 2025
Grading full written proofs rather than final answers, expert judges gave Gemini 2.5 Pro 24% and every other tested model under 5%, out of a possible 100%.
Benchmarks & progress
Anthropic raises $3.5 billion at a $61.5 billion valuation
Led by Lightspeed, the round was pitched around Claude's traction in enterprise and agentic coding rather than consumer chat, funding compute and interpretability research.
Money & business
Anthropic ships Claude 3.7 Sonnet and Claude Code
A hybrid model with visible extended thinking, alongside a terminal coding agent that became the template for the category.
Models & capabilities