Quick answer: Gemini 1.5 Flash (released May 14, 2024) is Google DeepMind's fast production model with a 1-million-token context window. At $0.075/$0.30 per million tokens, it is designed for high-throughput tasks requiring broad multimodal capabilities — text, images, audio, and video — at a fraction of Pro pricing.
Where Gemini 1.5 Flash leads
Where it lags
Best for: large-document analysis, audio/video understanding, high-volume classification, and cost-sensitive multimodal pipelines.
Gemini 1.5 Flash was released at Google I/O 2024 alongside Gemini 1.5 Pro. Where Pro targeted maximum quality, Flash targeted speed and scalability — built on the same architecture (Mixture of Experts) but with a narrower, faster-activating parameter budget. The result was a model that could process hour-long videos, thousands of pages of text, or mixed audio/text inputs at a price accessible for production deployments.
Flash's defining feature at launch was the full 1M-token context window (expandable to 2M in limited preview) — a 10× increase over GPT-4's 128K that opened entirely new use cases in document understanding, codebase analysis, and long-session agents.
| Field | Value |
|---|---|
| Organization | Google DeepMind |
| API identifier | gemini-1.5-flash |
| Context window | 1,000,000 tokens |
| Max output | 8,192 tokens |
| License | Proprietary |
| Release date | May 14, 2024 |
| Knowledge cutoff | May 2024 |
| Modality | Multimodal (text, image, audio, video, documents) |
| Architecture | Mixture of Experts |
| Tier | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Up to 128K tokens/request | $0.075 | $0.30 |
| Over 128K tokens/request | $0.15 | $0.60 |
Source: Google AI pricing.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU | 78.9% | Google — Gemini 1.5 Flash | 2024-05 |
| MATH | 77.9% | Google — Gemini 1.5 Flash | 2024-05 |
| HumanEval | 74.3% | Google — Gemini 1.5 Flash | 2024-05 |
Scores reported by Google. Not Benchgen measurements.
| Model | Context | Price (in/out per 1M) | Modality |
|---|---|---|---|
| Gemini 1.5 Flash | 1M | $0.075 / $0.30 | Multimodal |
| Gemini 1.5 Flash 8B | 1M | $0.0375 / $0.15 | Multimodal |
| Gemini 2.0 Flash | 1M | $0.10 / $0.40 | Multimodal |
| GPT-4o (May 2024) | 128K | $2.50 / $10 | Multimodal |
Gemini 1.5 Flash significantly undercuts GPT-4o on price while offering 8× the context. Gemini 2.0 Flash supersedes it with better quality at a slightly higher price. (Scores from respective announcements; not Benchgen measurements.)
Gemini 1.5 Flash is a large language model developed by Google.
This model isn’t on any benchmark leaderboard yet.