Quick answer: Gemini 2.0 Flash-Lite (released February 5, 2025) is Google DeepMind's most cost-efficient model in the Gemini 2.0 family — faster than Gemini 2.0 Flash and priced at the same $0.075/$0.30 per million tokens as Gemini 1.5 Flash. It delivers Gemini 2.0-generation quality at maximum throughput for high-volume, latency-sensitive workloads.
Where Gemini 2.0 Flash-Lite leads
Where it lags
Best for: real-time chat, high-volume classification and routing, lightweight sub-agents, and latency-critical production deployments.
Gemini 2.0 Flash-Lite was released during Google's Gemini 2.0 rollout in early 2025. Where Gemini 2.0 Flash balanced speed and quality, Flash-Lite pushed the speed/cost tradeoff further — Google's stated goal was a model "faster and more cost-efficient than 1.5 Flash" while maintaining Gemini 2.0-era improvements in instruction following, reasoning, and code generation.
For the Benchgen platform, Flash-Lite is particularly useful as a fast, cheap evaluation harness: running many model calls in parallel to score other models on MMLU-Pro or custom benchmarks, where the evaluator itself doesn't need frontier reasoning capability.
| Field | Value |
|---|---|
| Organization | Google DeepMind |
| API identifier | gemini-2.0-flash-lite |
| Context window | 1,000,000 tokens |
| Max output | 8,192 tokens |
| License | Proprietary |
| Release date | February 5, 2025 |
| Modality | Multimodal (text, image, audio, video input; text output) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Google AI | $0.075 | $0.30 |
Source: Google AI pricing.
| Model | Context | Price (in/out per 1M) | Generation |
|---|---|---|---|
| Gemini 2.0 Flash-Lite | 1M | $0.075 / $0.30 | Gemini 2.0 |
| Gemini 1.5 Flash | 1M | $0.075 / $0.30 | Gemini 1.5 |
| Gemini 2.0 Flash | 1M | $0.10 / $0.40 | Gemini 2.0 |
| Gemini 3 Flash | 1M | TBD | Gemini 3 |
Flash-Lite matches Gemini 1.5 Flash on price but delivers Gemini 2.0 quality. Gemini 2.0 Flash offers more reasoning depth for 33% more cost. (Scores from respective announcements; not Benchgen measurements.)
Gemini 2.0 Flash-Lite is a large language model developed by Google.
This model isn’t on any benchmark leaderboard yet.