Benchgen
Models/google/

Gemini 1.5 Flash 8B

DraftPublic

Model Details

Gemini 1.5 Flash 8B

Organization Context Pricing License Modality Released

Quick answer: Gemini 1.5 Flash 8B (released October 3, 2024) is Google DeepMind's smallest and most affordable model with a 1-million-token context window. At $0.0375/$0.15 per million tokens, it costs roughly 8× less than Gemini 1.5 Flash and is designed for high-throughput tasks where latency and cost matter more than maximum reasoning depth.

At a Glance

Where Gemini 1.5 Flash 8B leads

  • Cheapest 1M-context model at launch: $0.0375 input / $0.15 output per 1M tokens
  • Native multimodality across text, vision, audio, and video in one model
  • Optimised for real-time chat, classification, summarisation, and high-volume agentic sub-tasks
  • Available in Google AI Studio and the Gemini API with the same long-context capabilities as larger Gemini models

Where it lags

  • 8B parameter scale limits reasoning depth for hard coding, math, and multi-step problems
  • SWE-bench Verified score substantially below frontier models
  • Superseded by Gemini 2.0 Flash Lite and later Gemini 3 Flash for most new builds

Best for: high-volume pipelines, real-time classification, document summarisation over long contexts, and cost-optimised sub-agents.

What Gemini 1.5 Flash 8B Is

Gemini 1.5 Flash 8B was released to address the extreme cost sensitivity of production AI deployments. By October 2024, the 1M-token context window had become a Google differentiator, but larger Flash and Pro models still cost more than many tasks justified. The 8B model brought the same long-context architecture down to a price point competitive with the cheapest models on the market.

The model supports all Gemini modalities: text, images, video frames, audio, and documents — making it a capable multimodal router even if its reasoning ceiling is lower than the full Flash or Pro.

Specifications

FieldValue
OrganizationGoogle DeepMind
API identifiergemini-1.5-flash-8b
Context window1,000,000 tokens
Max output8,192 tokens
LicenseProprietary
Release dateOctober 3, 2024
Knowledge cutoffJuly 2024
ModalityMultimodal (text, image, audio, video, documents)

Pricing

TierInput (per 1M tokens)Output (per 1M tokens)
Up to 128K tokens/request$0.0375$0.15
Over 128K tokens/request$0.075$0.30

Source: Google AI pricing.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU70.0%Google — Gemini 1.5 Flash 8B2024-10
MATH73.3%Google — Gemini 1.5 Flash 8B2024-10

Scores reported by Google. Not Benchgen measurements.

Gemini 1.5 Flash 8B vs Alternatives

ModelContextPrice (in/out per 1M)Modality
Gemini 1.5 Flash 8B1M$0.0375 / $0.15Multimodal
Gemini 1.5 Flash1M$0.075 / $0.30Multimodal
Gemini 2.0 Flash Lite1M$0.075 / $0.30Multimodal
Claude Haiku 4.5200K$1.00 / $5.00Multimodal

Gemini 1.5 Flash 8B is the cheapest multimodal long-context option, though Claude Haiku 4.5 significantly leads on coding benchmarks at 13× higher cost. (Scores from respective announcements; not Benchgen measurements.)

Frequently Asked Questions

What is Gemini 1.5 Flash 8B? Gemini 1.5 Flash 8B is Google DeepMind's smallest Gemini model, offering a 1M-token context window at $0.0375/$0.15 per million tokens. Released October 2024, it is designed for cost-sensitive, high-volume workloads.
How much does Gemini 1.5 Flash 8B cost? $0.0375 per million input tokens and $0.15 per million output tokens (for requests up to 128K tokens). Doubles for longer requests.

Specs sourced from Google AI — Gemini 1.5 Flash 8B. Last updated 2026-06-19.

Gemini 1.5 Flash 8B

Gemini 1.5 Flash 8B is a large language model developed by Google.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.