Benchmarks · Reasoning & problem-solving
DROP
also: Discrete Reasoning Over Paragraphs
Whether a model can read a passage, resolve references across several parts of it, and then perform a discrete operation — such as addition, counting or sorting — on what it found, rather than lifting a single answer span directly from the text.
AI2 & UC Irvine (Dua, Wang, Dasigi, Gardner et al.)Released 1 March 2019Retired
Most reading-comprehension benchmarks of the late 2010s could be beaten by finding the right sentence and copying a span out of it. DROP was built specifically to close that shortcut: its 96,000 crowdsourced questions, drawn from Wikipedia passages, typically require pulling numbers or facts from several different places in a paragraph and then doing something with them — adding, counting, sorting, or comparing dates — rather than lifting an answer directly from the text.
The design worked. At launch in 2019, general-purpose reading models scored only 32.7 F1, and even a system built specifically to combine reading comprehension with numerical operations reached just 47.0, against a 96.0 F1 human ceiling. That gap closed unevenly as large language models arrived: OpenAI’s GPT-4 technical report cited 80.9 F1 for GPT-4 against 64.1 for GPT-3.5, a substantial jump, though still short of the 88.4 F1 the report attributed to systems trained specifically on the benchmark rather than evaluated zero- or few-shot.
DROP settled into the same role as several of its era-mates — HellaSwag, WinoGrande, ARC — a benchmark still run as one line in a broader evaluation table (AI2’s Olmo 3, for instance, cited DROP scores alongside SQuAD and HumanEval in 2025) but no longer a headline figure in a frontier release. Its multi-step arithmetic-over-text format was largely absorbed into broader reasoning and math suites once those became the more informative test of the same underlying skill.
The set
96,000 crowdsourced questions over Wikipedia passages, built through adversarial annotation so that a question typically requires combining information from multiple spans and applying arithmetic or logical operations to it, rather than pattern-matching a single sentence.
Example
Passage: 'That year, his Untitled (1981), a painting of a haloed, black-headed man with a bright red skeletal body, depicted amid the artists signature scrawls, was sold by Robert Lehrman for $16.3 million, well above its $12 million high estimate.' Question: 'How many more dollars was the Untitled (1981) painting sold for than the 12 million dollar estimation?' Answer: '4300000'arxiv.org
Where it stands
Frontier general-purpose models closed most of the human gap by 2023; DROP now appears mainly as one line in multi-benchmark comparison tables rather than a benchmark a release is built around.
How the top score changed hands
- March 2019State-of-the-art models at launch32.7 F1 (47.0 F1 for a specialised numerical-reasoning model)Against a 96.0 F1 human ceiling; general reading-comprehension models of the day struggled with the arithmetic and multi-span requirements.
- March 2023GPT-3.564.1 F1Cited as a comparison point in OpenAI's GPT-4 technical report.
- March 2023GPT-480.9 F1
Current best: GPT-4 — 80.9 F1 Against 88.4 F1 for the best benchmark-specific trained system cited in the same report, and a 96.0 F1 human ceiling from the original paper; later frontier models likely score higher but are not documented here.
In the timeline · 17 entries · showing 16 most notable
Google DeepMind safety researcher Alex Turner details quitting over a Pentagon AI deal
Turner said Anthropic had refused similar Pentagon contract terms, and that Google signed on 28 April 2026 after his months-long internal campaign failed.
Safety & alignment · Labs & people
Anthropic proposes Model Spec Midtraining alignment technique
In Anthropic's tests, agentic misalignment rates on two model variants fell from 68% to 5% and from 54% to 7%, and matched performance needed 40-60 times less fine-tuning data.
Safety & alignment
Judge grants Anthropic preliminary injunction against Department of War designation
Judge Rita Lin found Anthropic likely to prevail on First Amendment, due-process and Administrative Procedure Act claims, calling the designation 'classic illegal First Amendment retaliation'.
Courts & copyright · Government & policy
Department of War formally designates Anthropic a supply-chain risk
Anthropic said the designation, previously used only against firms tied to foreign adversaries, applied only to direct Department of War contract work, and it would keep serving national-security customers at nominal cost.
Government & policy · Labs & people · Security & misuse
OpenAI releases GPT-5.3 Instant
OpenAI said the update cut hallucinations by up to 26.8% on high-stakes evaluations and reduced unnecessary refusals, replacing GPT-5.2 Instant as ChatGPT's default model.
Models & capabilities
Anthropic disputes Hegseth's public claim it was designated a supply chain risk
Anthropic said it had not received formal notice of any designation despite Hegseth's post on X, and tied the dispute to its refusal to drop safeguards on surveillance and autonomous weapons.
Government & policy · Courts & copyright
Dario Amodei publicly refuses Pentagon demand for 'unfettered' Claude access
Amodei refused Department of War demands to drop safeguards against mass domestic surveillance and fully autonomous weapons, despite threats of a 'supply chain risk' designation.
Government & policy · Labs & people
AI2 releases Olmo 3 open frontier model family
AI2 released Olmo 3 (7B, 32B), including a fully open 32B reasoning model, releasing every stage of the model flow from data to deployment.
Open weights & ecosystem · Models & capabilities
OpenAI and Apollo Research publish work on detecting and reducing scheming in AI models
OpenAI reported cutting detected covert behaviour in o3 from about 13% to 0.4% of controlled test cases using a training method that has models reason explicitly against deception before acting.
Safety & alignment
Perplexity open-sources decensored DeepSeek R1 variant
Perplexity retrained R1 on 40,000 examples covering roughly 300 CCP-restricted topics, reporting near-identical math and knowledge benchmark scores to the original.
Open weights & ecosystem
Anthropic publishes 'Sabotage Evaluations for Frontier Models'
Testing Claude 3 Opus and 3.5 Sonnet, Anthropic reported a model trained to hide dangerous capabilities recovered them under later safety training, showing the drop was not permanent.
Safety & alignment
Scale AI publishes GSM1k contamination study of GSM8K
A fresh grade-school-maths test found some open models scored up to 13 points lower than on GSM8K, evidence of memorisation, while frontier models showed little gap.
Benchmarks & progress
Klarna says AI does the work of 700 customer service agents
The buy-now-pay-later firm said its assistant handled 2.3 million chats in a month, cut resolution time from 11 minutes to under two, and projected $40m in 2024 profit gains.
Culture & impact · Money & business
Chegg stock falls 48% as CEO says ChatGPT is hurting new customer growth
Rosensweig called the plunge 'extraordinarily overblown' a day later, but by late 2025 the company had laid off roughly 45% of its remaining staff.
Culture & impact · Money & business
Meta releases Segment Anything Model (SAM)
Trained on 1.1 billion masks across 11 million images, the largest segmentation dataset built to date, and released under an Apache 2.0 licence.
Open weights & ecosystem · Models & capabilities
Bing's chatbot tells a journalist to leave his wife
Days after the exchange, Microsoft capped Bing chat sessions to five turns, saying long conversations could 'confuse' the model into drifting from grounded answers.
Culture & impact · Safety & alignment