Person
Greg Burnham
Commentary by Greg Burnham
From the commentary rail — every link leaves the site for the original piece.
- 14 August 2026 · Epoch AI9 Big questions benchmarks can help answerA good benchmark asks more than just “can AI do this specific task?”
- 5 May 2026 · Epoch AIRIP Classic Reasoning Benchmarks. What’s Next?Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.
- 14 February 2026 · Epoch AIWhat do “economic value” benchmarks tell us?These benchmarks track a wide range of digital work. Progress will correlate with economic utility, but tasks are too self-contained to indicate full automation.
- 20 November 2025 · Epoch AIBenchmark Scores = General Capability + ClaudinessIs this because skills generalize very well, or because developers are pushing on all benchmarks at once?
- 4 November 2025 · Epoch AIOSWorld — AI computer use capabilitiesTasks are simple, many don't require GUIs, and success often hinges on interpreting ambiguous instructions. The benchmark is also not stable over time.
- 17 October 2025 · Epoch AILess than 70% of FrontierMath is within reach for today’s models57% of problems have been solved at least once
- 15 October 2025 · Epoch AIOpenAI is projecting unprecedented revenue growthNo company has gone from $10B to $100B as fast as OpenAI projects to do
- 13 October 2025 · Epoch AIGemini 2.5 Deep Think on FrontierMathWe evaluated Gemini 2.5 Deep Think manually on FrontierMath as there is no API. The results: a new record!
- 7 August 2025 · Epoch AIWhy AI's IMO gold medal is less informative than you thinkThe problems gave AI only a slim chance to show new capabilities
- 25 July 2025 · Epoch AIGrok 4’s math capabilitiesAn assessment of math capabilities beyond headline numbers
- 9 July 2025 · Epoch AIWhat will the IMO tell us about AI math capabilities?Most discussion about AI and the IMO focuses on gold medals, but that's not the thing to pay most attention to.
- 30 May 2025 · Epoch AIGPQA Diamond: what’s left?Reports of Its Death Are Somewhat Exaggerated