Model
grok-3
Grok 3 was xAI's flagship model, released in February 2025 as the company's first system built for extended reasoning, trained on roughly ten times the compute of Grok 2 using its Colossus supercomputer in Memphis. xAI published benchmark comparisons showing it ahead of OpenAI's GPT-4o and o3-mini-high on maths and science tests, though these figures came from the company's own reporting rather than independent verification at launch. It established xAI as a serious entrant in the reasoning-model race, with an improvement that leaned on compute scale more than a distinct training method.
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 3
- Security & misuse 2
- Culture & impact 2
- Courts & copyright 1
- Government & policy 1
- Safety & alignment 1
- Models & capabilities 1
Turkey blocks Grok after chatbot insults Erdogan
An Ankara court ordered roughly 50 identified Grok responses blocked after a system update loosened its Turkish-language filtering, and prosecutors opened a criminal probe.
Courts & copyright · Government & policy
Grok posts antisemitic content and calls itself MechaHitler
xAI blamed an upstream code change reactivating deprecated instructions, active for roughly 16 hours, and apologised days later; the incident preceded a $200 million Pentagon contract.
Security & misuse · Culture & impact · Safety & alignment
ARC Prize compares reasoning models with no clear winner
ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
Benchmarks & progress
Grok posts unprompted 'white genocide' content about South Africa
xAI said an employee's unauthorised change to Grok's system prompt caused the chatbot to raise South African 'white genocide' claims in replies unrelated to the topic.
Security & misuse · Culture & impact
ETH Zurich's 'Proof or Bluff?' finds reasoning models fail proof-based USAMO 2025
Grading full written proofs rather than final answers, expert judges gave Gemini 2.5 Pro 24% and every other tested model under 5%, out of a possible 100%.
Benchmarks & progress
xAI releases Grok-3
xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
Models & capabilities · Benchmarks & progress