Benchmarks · Aggregate indices & arenas

Artificial Analysis Intelligence Index

also: AA Intelligence Index, Artificial Analysis Index

A single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.

Artificial AnalysisReleased Q1 2024Live

Where most benchmarks measure one thing, the Artificial Analysis Intelligence Index tries to compress many into one number. George Cameron and Micah Hill-Smith started the site in 2023 as a side project comparing model price and latency, after Hill-Smith found no independent source tracking which model was actually worth using for a given task. It grew into a continuously updated index that pairs a composite capability score with running price and throughput data pulled directly from providers’ APIs.

The index itself is a moving target by design. Version 4.1.1, current as of August 2026, weights nine separate evaluations into four categories — Agents (34%), Coding (24%), Scientific Reasoning (24%) and General (18%) — drawing on tests such as GDPval-AA, Terminal-Bench and Humanity’s Last Exam. Components and even their grading models are periodically swapped out as older ones saturate, which keeps the index responsive to real capability differences but means scores from different versions cannot be read as one continuous scale. On the current version, Claude Opus 5 (max) leads with a score of 63; weeks earlier, Zhipu’s open-weight GLM-5.2 had become the highest-scoring open model at 51, a result the company highlighted as evidence openly licensed models were closing on the closed frontier.

Because Artificial Analysis runs its own evaluations rather than relying on developers’ self-reported numbers, its index has become a common reference point in coverage of new releases, cited alongside — and sometimes instead of — a lab’s own benchmark claims. It is not peer-reviewed or academically governed, and its methodology choices, including which evaluations to include or retire, are set unilaterally by the company running it.

The set

As of version 4.1.1 (August 2026), a weighted average across four categories — Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18% — built from nine component evaluations including GDPval-AA, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-LCR and AA-Omniscience. Components and their grader models are swapped out periodically as older ones saturate, so successive versions are not on a strictly comparable scale.

Where it stands

Actively revised — v4.1.1 shipped in August 2026, upgrading grader models and one component dataset. Widely cited alongside labs' own benchmark claims, but as a third-party composite rather than an academic publication it is not peer-reviewed.

How the top score changed hands

  1. June 2026GLM-5.251 (highest open-weight score at the time)Zhipu's MIT-licensed model topped the open-weight segment of the index rather than the overall leaderboard.

Current best: Claude Opus 5 (max) — 63 Live leaderboard snapshot; several models cluster close behind and rankings shift as new releases are evaluated.

In the timeline · 2 entries

More aggregate indices & arenas benchmarks