xAI releases Grok-3
xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
- Models & capabilities
- Benchmarks & progress
- Notable
xAI released Grok 3 in an evening livestream on X, its first model built for extended reasoning (“Grok 3 Think”) alongside a lighter “Grok 3 mini” variant. Elon Musk called it “an order of magnitude more capable than Grok 2,” trained on what xAI said was roughly ten times the compute used for its predecessor, run on the Colossus supercomputer in Memphis, by then expanded to around 200,000 Nvidia GPUs.
xAI published its own benchmark comparisons showing Grok 3 ahead of OpenAI’s GPT-4o and, in its reasoning configuration, ahead of o3-mini-high on AIME 2025, GPQA and other maths- and science-heavy tests, along with a strong placement on the crowdsourced Chatbot Arena leaderboard. As with earlier xAI releases, these figures came from the company’s own reporting rather than independent verification at launch. Musk again framed the model around free expression, describing it as “a maximally truth-seeking AI, even if that truth is sometimes at odds with what is politically correct” — a framing that had drawn scrutiny before, after researchers found earlier Grok versions leaned left on several social and political topics despite similar pledges.
Grok 3 launched alongside a new subscription tier, SuperGrok, priced at $30 a month or $300 a year, offering expanded access to reasoning mode and the model’s “DeepSearch” research-agent feature. The release followed a Grok 2 that had been criticised as behind the frontier, and it established xAI as a fourth serious entrant in the reasoning-model race alongside OpenAI, Google and DeepSeek — though its improvement leaned on compute scale more than a distinct training method.
Five months later, xAI released Grok 4, pushing further on the same benchmarks amid renewed scrutiny of the company’s other products.