Google releases Gemini 3.1 Flash-Lite
Priced at $0.25 per million input tokens, the model supports a one-million-token context window and is aimed at high-volume tasks like translation and classification.
- Models & capabilities
- Colour
Google released Gemini 3.1 Flash-Lite, a cut-down variant of Gemini 3 Pro positioned as the fastest and cheapest model in the Gemini 3 line, aimed at high-volume, latency-sensitive tasks such as translation and classification rather than frontier reasoning. It accepts text, image, audio and video input, with up to one million tokens of context, and was priced at $0.25 per million input tokens and $1.50 per million output tokens.
On Google’s own disclosed benchmarks the model scored 86.9% on GPQA Diamond and 76.8% on MMMU-Pro, figures intended to show a low-cost model still clearing a reasonably high capability bar. The release continued Google’s practice of shipping smaller “Lite” and “Flash” variants alongside flagship Pro models, for developers where per-token cost, not peak capability, determines what is viable to deploy at scale.