Timeline

Reports emerge that pre-training gains are slowing

Reuters cited a dozen AI scientists and investors, and quoted Ilya Sutskever saying results from scaling up pre-training had plateaued, pointing instead to inference-time reasoning techniques.

  • Ideas & essays
  • Benchmarks & progress
  • Notable

Reuters reported that AI labs were increasingly turning to new training techniques after runs of ever-larger pre-training produced smaller gains than expected. The article cited about a dozen unnamed AI scientists, researchers and investors, and quoted Ilya Sutskever, the former OpenAI chief scientist who had since co-founded Safe Superintelligence, saying that results from scaling up pre-training had “plateaued.”

As a concrete example, the report said OpenAI had completed initial training of a new model, internally called Orion, in September 2024, and that early results showed a smaller improvement over GPT-4 than GPT-4 itself had shown over GPT-3.5 — including underperformance on coding tasks outside its training distribution. Reuters framed this alongside broader industry discussion that simply adding more data and compute to next-token prediction training was yielding diminishing returns, a view attributed to researchers at multiple labs including OpenAI, Google DeepMind and Anthropic.

The alternative the report highlighted was techniques behind OpenAI’s recently released o1 model: instead of, or in addition to, scaling pre-training, models could be given more computation at inference time to reason through a problem step by step before answering. Reuters said researchers viewed this shift as carrying implications well beyond model architecture, potentially reshaping the balance of resources — energy, chip types, data — that AI companies would need going forward.

The report crystallised an argument that had been building through 2024 among researchers sceptical that the scaling laws underpinning GPT-3 and GPT-4’s gains would continue indefinitely. It was disputed by lab leaders who continued to publicly defend pre-training scaling’s trajectory, and subsequent releases — including reasoning models from OpenAI, Google and Chinese labs through 2025 — showed the industry pursuing test-time compute and pre-training scale in parallel rather than treating either as exhausted.