Timeline

xAI confirms Grok 5 training on Colossus 2, targets 6T-parameter MoE model

xAI disclosed the training alongside its $20 billion Series E raise; reports described xAI training two Grok 5 variants in parallel, at 6 trillion and 10 trillion parameters.

  • Models & capabilities
  • Compute & infrastructure
  • Minor

xAI disclosed that its next flagship model, Grok 5, was training on its Colossus 2 cluster, alongside the announcement of its $20 billion Series E funding round. Reports based on that disclosure described xAI running two Grok 5 variants in parallel — one with roughly 6 trillion parameters, another at roughly 10 trillion — built as sparse mixture-of-experts models in which only a fraction of the total parameters activate for any given input, the same broad architectural approach used by DeepSeek’s largest releases.

Colossus 2, xAI’s Memphis, Tennessee compute cluster, crossed 1 gigawatt of capacity and was targeting roughly 2 gigawatts, with Elon Musk describing it as the world’s first gigawatt-scale AI training cluster. Musk said the site was running seven concurrent training jobs, of which the Grok 5 variants were among the largest, and that a further upgrade to 1.5 gigawatts was scheduled for April.

Neither Grok 5 configuration had been released as a product by the time of the disclosure. xAI had previously missed earlier informal release windows for the model, and precise parameter counts and architecture details for Grok 5 were not independently verified beyond what xAI and Musk stated publicly — a pattern common to frontier model announcements, where headline compute and scale figures typically precede any benchmark evidence of what the resulting model can do.

Referenced by