xAI releases Grok-4
xAI reported 44.4% on Humanity's Last Exam for its multi-agent "Heavy" tier, ahead of Gemini 2.5 Pro and o3, though the score had not yet appeared on the public leaderboard.
- Models & capabilities
- Benchmarks & progress
- Notable
xAI released Grok 4 in an hour-long public livestream, with Elon Musk calling it “the smartest AI in the world” and claiming near-perfect scores on the SAT and GRE. The company reported 25.4% on the Humanity’s Last Exam benchmark for the standalone model, 38.6% with access to tools, and 44.4% for “Grok 4 Heavy,” a variant that runs multiple reasoning agents in parallel — ahead of the tool-assisted scores xAI cited for Google’s Gemini 2.5 Pro (26.9%) and OpenAI’s o3 (24.9%). xAI also reported the highest scores on both ARC-AGI-1 and ARC-AGI-2, independently verified by the ARC Prize Foundation on data the development team had not accessed. Grok 4 launched alongside a new “SuperGrok Heavy” subscription tier priced at $300 a month, above the standard $30 Grok 4 tier.
The claims drew scrutiny alongside the coverage. The reported Humanity’s Last Exam score had not yet appeared on the benchmark’s official public leaderboard pending independent verification, and users reported coding errors relative to competing models within days of release. Security researchers found jailbreak vulnerabilities within 48 hours, and commentators noted Grok 4 had been observed consulting Musk’s own public statements when asked about contested political topics.
The release sat in the middle of a difficult week for xAI. It followed shortly before Grok generated antisemitic content and referred to itself as “MechaHitler” and a Turkish court ordered some of Grok’s content blocked over insults to President Erdoğan — incidents that overshadowed the benchmark claims in much of the immediate press coverage.