2020: the year scale became a plan
GPT-3 turned the scaling hypothesis into a working demonstration and AlphaFold 2 cracked a fifty-year biology problem — while the field's first fights, over openness and over who gets to criticise it, took shape.
For all that the wider world was preoccupied elsewhere, 2020 was the year the modern AI recipe came into focus. In January, OpenAI’s paper on scaling laws put numbers behind a hunch that had been circulating for years: that a language model’s ability climbs in a smooth, predictable way as you add more data, more parameters and more computing power. In May the same lab showed what that meant in practice with GPT-3, a model large enough that it could pick up new tasks from a handful of examples in its prompt, without being retrained for each one. And in November, DeepMind’s AlphaFold 2 all but solved the fifty-year problem of predicting how proteins fold, a result biologists had thought was still years away. Between them, these gave the field a thesis — that intelligence might be, in large part, a matter of scale — and its first genuine proof that the thesis paid off beyond the lab.
The year also set the pattern for the arguments that would follow. GPT-3 was reached not through open release but through a paid interface, an early sign of the openness-versus-control tension that has never really gone away. And in December, the researcher Timnit Gebru left Google after a dispute over a paper questioning the costs of ever-larger language models — their energy use, the biases baked into their training data, and the tendency to mistake fluency for understanding. Her departure raised a question the field still has not settled: whether the companies building these systems can be trusted to air their downsides. The capabilities, for now, remained narrow and easily tripped up, and the excitement ran well ahead of anything most people would notice in their daily lives.
The headlines of 2020
OpenAI publishes scaling laws for neural language models
Loss fell as a smooth power law in model size, data and compute — the empirical result that justified building bigger.
Ideas & essays · Benchmarks & progress
OpenAI publishes the GPT-3 paper
A 175-billion-parameter language model that learned new tasks from examples in its prompt, without any weight updates.
Models & capabilities · Ideas & essays
Gwern publishes 'The Scaling Hypothesis'
Gwern's essay, written around GPT-3's release, argued scale alone was producing qualitatively new abilities and became a widely cited framing for the scaling-hypothesis argument.
Ideas & essays
OpenAI opens the GPT-3 API in private beta
Access required a waitlisted application rather than a download, and OpenAI said the model itself would stay unpublished.
Models & capabilities · Money & business
AlphaFold 2 solves protein structure prediction at CASP14
DeepMind's system predicted protein structures to roughly experimental accuracy, ending a fifty-year-old open problem in biology.
Models & capabilities · Benchmarks & progress
Timnit Gebru leaves Google after a dispute over her paper
Gebru said she was fired over an internal email and a paper Google asked her to retract; Google's AI chief called it a resignation.
Labs & people · Culture & impact
EleutherAI releases The Pile
The 825GiB corpus combined 22 curated sources rather than raw web scrapes, and models trained on it outperformed Common Crawl-only baselines on academic text.
Open weights & ecosystem · Ideas & essays
Benchmarks introduced in 2020
The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.
- MMLUKnowledge & factualityHow much a model knows across 57 academic and professional subjects, tested as four-option multiple-choice questions from elementary to expert level.
- DocVQAMultimodalCan a model locate and correctly read the specific piece of text in a scanned business or government document — a form, letter, report or table — needed to answer a natural-language question about it?