Benchmarks · Aggregate indices & arenas

Artificial Analysis Coding Agent Index

also: AA Coding Agent Index, Artificial Analysis Coding Index

How well a coding agent — a specific model paired with a specific harness, such as Claude Code or Codex — completes real software-engineering work end to end, not just whether the underlying model answers a coding question correctly.

Artificial AnalysisReleased May 2026Live

Artificial Analysis’s Coding Agent Index asks a narrower question than its general Intelligence Index: not how capable a model is in the abstract, but how well a specific agent setup — a model wired into a coding harness such as Claude Code, Codex or Cursor CLI — gets through real software work. Results are reported by that model-and-harness pairing rather than by model alone, because the same underlying model can score differently depending on which tool-calling scaffold it runs inside.

The index combines three benchmarks in equal measure: DeepSWE’s long-horizon engineering tasks, Terminal-Bench v2’s agentic terminal use, and Scale AI’s SWE-Atlas-QnA repository question-answering set, each scored on the first-attempt pass rate averaged over three tries. Alongside the composite score, Artificial Analysis publishes cost, token usage and wall-clock time per task, letting a reader weigh a faster or cheaper agent against a marginally more capable one rather than reading capability alone. When OpenAI released GPT-5.6 in July 2026, it cited a score of 80 for its Sol model running at maximum reasoning effort, ahead of Anthropic’s Claude Fable 5 on the same index at the time, while using fewer tokens and less compute per task.

Because it scores agent-and-harness combinations rather than static model outputs, and because its component benchmarks are comparatively new and still being iterated — the index stood at version 1.3 as of August 2026 — it functions less as a settled reference than as a running snapshot of which combination of model and tooling currently gets real coding work done fastest and cheapest.

The set

As of v1.3, a composite of three benchmarks weighted equally: DeepSWE (113 long-horizon software-engineering tasks), Terminal-Bench v2 (84 agentic terminal-use tasks) and SWE-Atlas-QnA (124 repository question-answering tasks by Scale AI), each scored as pass@1 averaged across three attempts. Results are reported per model-and-harness combination — for example 'Claude Code – Opus 5' versus 'Codex – GPT-5.6 Sol' — rather than per model alone, alongside cost, token usage and wall-clock time for each.

Where it stands

Launched in May 2026 and revised three times since — components and scoring methods have already changed twice — so it functions as a running snapshot rather than a settled reference. Scores agent-plus-harness combinations rather than raw models, distinguishing it from Artificial Analysis's broader Intelligence Index.

How the top score changed hands

  1. July 2026GPT-5.6 Sol (max reasoning)80Reported in OpenAI's own release announcement, ahead of Anthropic's Claude Fable 5 on this index at the time, using under half the output tokens and roughly a third less cost.

Current best: Claude Code – Claude Opus 5 (xhigh) / Codex – GPT-5.6 Sol (max) — 67 (tied) Live leaderboard snapshot; OpenAI's own launch announcement separately reported Sol scoring 80 on this index in its maximum-reasoning configuration, ahead of Anthropic's Claude Fable 5, using fewer tokens and lower cost — a different snapshot than the one cited here.

In the timeline · 1 entry

More aggregate indices & arenas benchmarks