Epoch AI publishes 'Will We Run Out of Data?'
Modelled dataset growth against the finite stock of public text and projected the two curves would cross between 2026 and 2032, sooner if models were overtrained.
- Ideas & essays
- Compute & infrastructure
- Notable
Epoch AI published an analysis, “Will We Run Out of Data?”, projecting when the supply of publicly available human-generated text would be exhausted relative to the growing appetite of large language model training runs. Led by Pablo Villalobos and Jaime Sevilla, the paper modelled two trends — the rate at which frontier labs were scaling training datasets, and estimates of the total stock of public text of usable quality — and projected the curves would cross sometime between 2026 and 2032, somewhat earlier if models continued to be “overtrained” on more tokens per parameter than was compute-optimal.
The paper distinguished high-quality text, such as edited prose and code, from lower-quality but far larger sources such as social media and general web scrapes, estimating the higher-quality pool would be exhausted first. It proposed three ways scaling could continue past that point: synthetic data generated by models themselves, transferring capability from data-rich domains into data-poor ones, and training more efficiently on data that already existed — all three of which subsequently became active areas of lab research rather than remaining hypothetical.
The projection was explicitly a forecast under stated assumptions rather than a measured physical limit, and the paper was revised in 2024 as trends became clearer. Even so, it gave the field’s informal worry about a coming “data wall” a specific, citable estimate, and it became one of the standard references in later arguments — including reporting in 2024 that some frontier training runs were already showing diminishing returns from further pretraining scale — about whether scaling data and compute together could keep producing capability gains indefinitely.