Benchgen
Models/google/

Gemini 2.0 Flash-Lite

DraftPublic

Model Details

Gemini 2.0 Flash-Lite

Organization Context Pricing License Modality Released

Quick answer: Gemini 2.0 Flash-Lite (released February 5, 2025) is Google DeepMind's most cost-efficient model in the Gemini 2.0 family — faster than Gemini 2.0 Flash and priced at the same $0.075/$0.30 per million tokens as Gemini 1.5 Flash. It delivers Gemini 2.0-generation quality at maximum throughput for high-volume, latency-sensitive workloads.

At a Glance

Where Gemini 2.0 Flash-Lite leads

  • Fastest model in the Gemini 2.0 family — optimised for minimum latency at scale
  • Gemini 2.0-class quality (better than 1.5 Flash) at the same price point
  • 1M-token context window with full multimodal support
  • Native tool use and function calling for agent integrations

Where it lags

  • Lower reasoning depth than Gemini 2.0 Flash for complex multi-step tasks
  • No image/video generation (output is text only)

Best for: real-time chat, high-volume classification and routing, lightweight sub-agents, and latency-critical production deployments.

What Gemini 2.0 Flash-Lite Is

Gemini 2.0 Flash-Lite was released during Google's Gemini 2.0 rollout in early 2025. Where Gemini 2.0 Flash balanced speed and quality, Flash-Lite pushed the speed/cost tradeoff further — Google's stated goal was a model "faster and more cost-efficient than 1.5 Flash" while maintaining Gemini 2.0-era improvements in instruction following, reasoning, and code generation.

For the Benchgen platform, Flash-Lite is particularly useful as a fast, cheap evaluation harness: running many model calls in parallel to score other models on MMLU-Pro or custom benchmarks, where the evaluator itself doesn't need frontier reasoning capability.

Specifications

FieldValue
OrganizationGoogle DeepMind
API identifiergemini-2.0-flash-lite
Context window1,000,000 tokens
Max output8,192 tokens
LicenseProprietary
Release dateFebruary 5, 2025
ModalityMultimodal (text, image, audio, video input; text output)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Google AI$0.075$0.30

Source: Google AI pricing.

Gemini 2.0 Flash-Lite vs Alternatives

ModelContextPrice (in/out per 1M)Generation
Gemini 2.0 Flash-Lite1M$0.075 / $0.30Gemini 2.0
Gemini 1.5 Flash1M$0.075 / $0.30Gemini 1.5
Gemini 2.0 Flash1M$0.10 / $0.40Gemini 2.0
Gemini 3 Flash1MTBDGemini 3

Flash-Lite matches Gemini 1.5 Flash on price but delivers Gemini 2.0 quality. Gemini 2.0 Flash offers more reasoning depth for 33% more cost. (Scores from respective announcements; not Benchgen measurements.)


Specs sourced from Google — Gemini 2.0 release. Last updated 2026-06-19.

Gemini 2.0 Flash-Lite

Gemini 2.0 Flash-Lite is a large language model developed by Google.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.