Model
claude-2
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 1
- Models & capabilities 1
SWE-bench paper published
Built from 2,294 real GitHub issues across 12 Python repositories, the benchmark proved so hard that the best model of the day, Claude 2, solved under 2%.
Benchmarks & progress
Anthropic releases Claude 2
A 100,000-token context window and a jump to 71.2% on the Codex HumanEval coding test, alongside a consumer web app opened to the US and UK.
Models & capabilities