Quick answer: Gemini 1.5 Flash 8B (released October 3, 2024) is Google DeepMind's smallest and most affordable model with a 1-million-token context window. At $0.0375/$0.15 per million tokens, it costs roughly 8× less than Gemini 1.5 Flash and is designed for high-throughput tasks where latency and cost matter more than maximum reasoning depth.
Where Gemini 1.5 Flash 8B leads
Where it lags
Best for: high-volume pipelines, real-time classification, document summarisation over long contexts, and cost-optimised sub-agents.
Gemini 1.5 Flash 8B was released to address the extreme cost sensitivity of production AI deployments. By October 2024, the 1M-token context window had become a Google differentiator, but larger Flash and Pro models still cost more than many tasks justified. The 8B model brought the same long-context architecture down to a price point competitive with the cheapest models on the market.
The model supports all Gemini modalities: text, images, video frames, audio, and documents — making it a capable multimodal router even if its reasoning ceiling is lower than the full Flash or Pro.
| Field | Value |
|---|---|
| Organization | Google DeepMind |
| API identifier | gemini-1.5-flash-8b |
| Context window | 1,000,000 tokens |
| Max output | 8,192 tokens |
| License | Proprietary |
| Release date | October 3, 2024 |
| Knowledge cutoff | July 2024 |
| Modality | Multimodal (text, image, audio, video, documents) |
| Tier | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Up to 128K tokens/request | $0.0375 | $0.15 |
| Over 128K tokens/request | $0.075 | $0.30 |
Source: Google AI pricing.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU | 70.0% | Google — Gemini 1.5 Flash 8B | 2024-10 |
| MATH | 73.3% | Google — Gemini 1.5 Flash 8B | 2024-10 |
Scores reported by Google. Not Benchgen measurements.
| Model | Context | Price (in/out per 1M) | Modality |
|---|---|---|---|
| Gemini 1.5 Flash 8B | 1M | $0.0375 / $0.15 | Multimodal |
| Gemini 1.5 Flash | 1M | $0.075 / $0.30 | Multimodal |
| Gemini 2.0 Flash Lite | 1M | $0.075 / $0.30 | Multimodal |
| Claude Haiku 4.5 | 200K | $1.00 / $5.00 | Multimodal |
Gemini 1.5 Flash 8B is the cheapest multimodal long-context option, though Claude Haiku 4.5 significantly leads on coding benchmarks at 13× higher cost. (Scores from respective announcements; not Benchgen measurements.)
Gemini 1.5 Flash 8B is a large language model developed by Google.
This model isn’t on any benchmark leaderboard yet.