Google makes Gemini 3 Flash the default model across its products
Priced at $0.50/$3.00 per million tokens, Google reported it ran three times faster than Gemini 2.5 Pro while scoring 33.7% on Humanity's Last Exam, against 37.5% for Gemini 3 Pro.
- Models & capabilities
- Minor
Google released Gemini 3 Flash, a smaller and cheaper sibling to Gemini 3 Pro, and made it the default model in the Gemini app globally and in AI Mode within Google Search, replacing Gemini 2.5 Flash. Google priced it at $0.50 per million input tokens and $3.00 per million output tokens — above Gemini 2.5 Flash’s rate but, the company argued, more cost-effective in practice given reported efficiency gains.
Google reported Gemini 3 Flash running three times faster than Gemini 2.5 Pro while using 30% fewer tokens on reasoning tasks, and cited benchmark scores of 33.7% on Humanity’s Last Exam (against 37.5% for Gemini 3 Pro), 81.2% on the multimodal MMMU-Pro benchmark, and 78% on SWE-bench for coding. Tulsee Doshi, Google’s senior director for Gemini product, said the model’s low per-token cost was aimed particularly at companies running AI at bulk scale, where cost per call rather than peak capability determines what is viable to deploy.
The release followed the pattern set across the industry in late 2025 of shipping a fast, cheap “default” tier alongside a frontier flagship, reflecting that most day-to-day model usage — search, chat, routine coding — does not require top-tier reasoning performance, and that the competitive battle over inference cost had become as consequential to user experience as the headline capability race between Gemini 3 Pro and OpenAI’s GPT-5.2.