The scaling hypothesis and its critics
The bet that more data, parameters and compute keep buying capability — given an empirical spine by GPT-3, revised by Chinchilla, and by 2025 argued to have reached the end of its pretraining era.
The scaling hypothesis is the claim that a neural network’s capability rises predictably with the compute, data and parameters poured into it — and that intelligence might be, in large part, a matter of scale. OpenAI gave it an empirical spine with its 2020 scaling-laws paper, then a demonstration months later with the 175-billion-parameter GPT-3, whose few-shot abilities emerged from size rather than task-specific training.
The hypothesis drew critics as fast as adherents. “On the Dangers of Stochastic Parrots” argued that a larger language model was a better mimic, not a better reasoner, and weighed the costs of building one; Stanford’s foundation-models report named the category even as it questioned it. The recipe was revised from within, too: DeepMind’s Chinchilla paper showed most large models were badly under-trained on data for their size, redrawing the optimal trade-off between parameters and tokens.
GPT-4 was the apex of the pretraining-scale era. Then the axis moved: a test-time-compute paper and OpenAI’s o1 showed that spending more compute at inference — letting a model think for longer — could beat simply building it bigger. By late 2025 even Ilya Sutskever was declaring the scaling era over, arguing the field had to return to new research ideas rather than raw scale. Whether that is the hypothesis failing or merely changing its variable is the argument this thread leaves open.
OpenAI publishes scaling laws for neural language models
Loss fell as a smooth power law in model size, data and compute — the empirical result that justified building bigger.
Ideas & essays · Benchmarks & progress
Microsoft announces Turing-NLG at 17 billion parameters
Briefly the largest published language model, trained with Microsoft's new DeepSpeed library and beating Megatron-LM on WikiText-103 and LAMBADA.
Models & capabilities
OpenAI publishes the GPT-3 paper
A 175-billion-parameter language model that learned new tasks from examples in its prompt, without any weight updates.
Models & capabilities · Ideas & essays
Gwern publishes 'The Scaling Hypothesis'
Gwern's essay, written around GPT-3's release, argued scale alone was producing qualitatively new abilities and became a widely cited framing for the scaling-hypothesis argument.
Ideas & essays
OpenAI announces DALL·E and CLIP
Two multimodal models released on the same day: one that generated images from text captions, one that learned to match images to language.
Models & capabilities · Open weights & ecosystem
Google trains a 1.6-trillion-parameter Switch Transformer
Routing each input to a single expert rather than blending several simplified prior mixture-of-experts designs, and the paper reported up to 7x faster pre-training than a dense T5 baseline.
Models & capabilities
"On the Dangers of Stochastic Parrots" is published
The paper that named the case against ever-larger language models — and whose suppression cost two Google researchers their jobs.
Ideas & essays · Culture & impact
BAAI announces Wu Dao 2.0 at 1.75 trillion parameters
The Beijing Academy of Artificial Intelligence put the parameter count at ten times GPT-3's, a claim reported widely but never independently benchmarked.
Models & capabilities
Stanford's foundation models report names the category
Over 100 researchers at Stanford's newly formed Center for Research on Foundation Models coined the term for models like GPT-3 and BERT, adapted rather than retrained for each task.
Ideas & essays
Microsoft and NVIDIA train Megatron-Turing NLG at 530 billion parameters
The largest dense language model publicly described at the time, and a demonstration of multi-thousand-GPU training.
Models & capabilities · Compute & infrastructure
DeepMind publishes Gopher, RETRO and a risk taxonomy together
A 280-billion-parameter model, a smaller retrieval-augmented alternative that matched larger models, and a taxonomy of six categories of language-model harm.
Models & capabilities · Safety & alignment
DeepMind's Chinchilla paper rewrites the scaling laws
Existing large models were badly under-trained: for a fixed compute budget, parameters and training tokens should scale together.
Ideas & essays · Benchmarks & progress · Models & capabilities
Google announces PaLM at 540 billion parameters
Trained on the Pathways system across two TPU v4 pods, it posted large gains on reasoning benchmarks and explained its own jokes.
Models & capabilities · Compute & infrastructure
DeepMind's Gato does 600 tasks with one set of weights
A single 1.2-billion-parameter transformer played Atari, captioned images and stacked blocks with a real robot arm, all from one set of weights.
Models & capabilities · Ideas & essays
FlashAttention makes exact attention IO-aware
Reordering attention around GPU memory rather than approximating it cut training time and unlocked longer sequences — and became default infrastructure.
Ideas & essays · Compute & infrastructure
Emergent abilities of large language models are described
Jason Wei and co-authors catalogued tasks where accuracy jumped from near-chance to strong performance past a scale threshold, a pattern later disputed as a metric artefact.
Ideas & essays · Benchmarks & progress
Epoch AI publishes 'Will We Run Out of Data?'
Modelled dataset growth against the finite stock of public text and projected the two curves would cross between 2026 and 2032, sooner if models were overtrained.
Ideas & essays · Compute & infrastructure
OpenAI releases GPT-4
A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
Models & capabilities · Benchmarks & progress
'Are Emergent Abilities of Large Language Models a Mirage?' challenges emergence claims
Reanalysing the same benchmark results with linear metrics, the Stanford authors made the apparent phase transitions disappear; the paper won a NeurIPS 2023 outstanding paper award.
Ideas & essays · Benchmarks & progress
DeepSeek releases DeepSeek-V2
A 236-billion-parameter mixture-of-experts model with only 21 billion active per token, released open-weight; DeepSeek said training costs fell 42% versus its prior model.
Open weights & ecosystem · Models & capabilities
Epoch AI publishes 'Training compute of frontier AI models grows by 4-5x per year'
The estimate drew on 333 compute figures for notable models since 2010 — roughly triple Epoch's 2022 dataset — and flagged an unexplained slowdown around 2018.
Ideas & essays · Compute & infrastructure
Scaling test-time compute paper argues extra inference compute can beat bigger models
On some problems, extra inference-time computation matched the gains from a pretrained model roughly 14 times larger, the authors reported.
Ideas & essays · Benchmarks & progress
OpenAI releases o1, trading inference time for reasoning
A model trained to think before answering opened a second scaling axis: spend more compute at inference and accuracy rises.
Models & capabilities · Benchmarks & progress
Sam Altman publishes 'The Intelligence Age' essay
Altman wrote that superintelligence could arrive within 'a few thousand days,' framing scaling deep learning as an already-solved algorithmic problem needing only more compute.
Ideas & essays
Reports emerge that pre-training gains are slowing
Reuters cited a dozen AI scientists and investors, and quoted Ilya Sutskever saying results from scaling up pre-training had plateaued, pointing instead to inference-time reasoning techniques.
Ideas & essays · Benchmarks & progress
OpenAI releases GPT-4.5
Priced at $75/$150 per million tokens, about thirty times GPT-4o's rate, and retired from the API within five months in favour of the cheaper GPT-4.1.
Models & capabilities
Sam Altman publishes 'The Gentle Singularity'
Altman declared 'we are past the event horizon' of the singularity, predicting novel-insight-generating systems in 2026 and real-world robots by 2027.
Ideas & essays
Dwarkesh Patel's second interview with Ilya Sutskever declares the scaling era over
Sutskever said models 'generalize dramatically worse than people,' citing an example of an AI that fixes a bug, breaks it again, then reverts to the original error when corrected.
Ideas & essays
xAI confirms Grok 5 training on Colossus 2, targets 6T-parameter MoE model
xAI disclosed the training alongside its $20 billion Series E raise; reports described xAI training two Grok 5 variants in parallel, at 6 trillion and 10 trillion parameters.
Models & capabilities · Compute & infrastructure
Yann LeCun launches AMI Labs to build JEPA-based world models
The $1.03 billion seed round, at a $3.5 billion pre-money valuation, was co-led by Cathay Innovation, Greycroft, Hiro Capital and HV Capital, with Jeff Bezos, Nvidia and Temasek among the backers.
Labs & people · Models & capabilities
OpenAI releases GPT-5.5
Pitched as OpenAI's most agentic model yet, it shipped alongside a $50,000 bug-bounty for jailbreaks that could extract biological-weapons help.
Models & capabilities
Anthropic publishes 'When AI builds itself', calls for coordinated pause option
The essay says the length of tasks models complete unassisted has doubled roughly every four months since 2024, and proposes a verification scheme for a coordinated slowdown.
Safety & alignment · Ideas & essays
Anthropic launches Claude Fable 5 and Claude Mythos 5
Fable 5 and Mythos 5 share the same underlying model, but only Fable 5 carries safety classifiers that can refuse requests; Mythos 5 is restricted to vetted cyber-defence and biosecurity partners.
Models & capabilities · Safety & alignment
OpenAI releases GPT-5.6
Released in three tiers — Sol, Terra and Luna — after a delayed rollout attributed to US government review, with OpenAI billing the flagship as its strongest cybersecurity model yet.
Models & capabilities · Security & misuse