Benchgen
Models/google/

Gemini 1.5 Flash

DraftPublic

Model Details

Gemini 1.5 Flash

Organization Context Pricing License Modality Released

Quick answer: Gemini 1.5 Flash (released May 14, 2024) is Google DeepMind's fast production model with a 1-million-token context window. At $0.075/$0.30 per million tokens, it is designed for high-throughput tasks requiring broad multimodal capabilities — text, images, audio, and video — at a fraction of Pro pricing.

At a Glance

Where Gemini 1.5 Flash leads

  • 1M-token context window — the first model to offer this at scale when launched
  • True native multimodality: processes text, images, audio, video, and PDFs in a single model call
  • Fast response times optimised for production-scale batch and real-time workloads
  • Used widely in Google's own products and Workspace integrations

Where it lags

  • Reasoning and coding benchmarks well below Gemini 1.5 Pro and the Gemini 3 family
  • Superseded by Gemini 2.0 Flash for most new builds requiring better quality

Best for: large-document analysis, audio/video understanding, high-volume classification, and cost-sensitive multimodal pipelines.

What Gemini 1.5 Flash Is

Gemini 1.5 Flash was released at Google I/O 2024 alongside Gemini 1.5 Pro. Where Pro targeted maximum quality, Flash targeted speed and scalability — built on the same architecture (Mixture of Experts) but with a narrower, faster-activating parameter budget. The result was a model that could process hour-long videos, thousands of pages of text, or mixed audio/text inputs at a price accessible for production deployments.

Flash's defining feature at launch was the full 1M-token context window (expandable to 2M in limited preview) — a 10× increase over GPT-4's 128K that opened entirely new use cases in document understanding, codebase analysis, and long-session agents.

Specifications

FieldValue
OrganizationGoogle DeepMind
API identifiergemini-1.5-flash
Context window1,000,000 tokens
Max output8,192 tokens
LicenseProprietary
Release dateMay 14, 2024
Knowledge cutoffMay 2024
ModalityMultimodal (text, image, audio, video, documents)
ArchitectureMixture of Experts

Pricing

TierInput (per 1M tokens)Output (per 1M tokens)
Up to 128K tokens/request$0.075$0.30
Over 128K tokens/request$0.15$0.60

Source: Google AI pricing.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU78.9%Google — Gemini 1.5 Flash2024-05
MATH77.9%Google — Gemini 1.5 Flash2024-05
HumanEval74.3%Google — Gemini 1.5 Flash2024-05

Scores reported by Google. Not Benchgen measurements.

Gemini 1.5 Flash vs Alternatives

ModelContextPrice (in/out per 1M)Modality
Gemini 1.5 Flash1M$0.075 / $0.30Multimodal
Gemini 1.5 Flash 8B1M$0.0375 / $0.15Multimodal
Gemini 2.0 Flash1M$0.10 / $0.40Multimodal
GPT-4o (May 2024)128K$2.50 / $10Multimodal

Gemini 1.5 Flash significantly undercuts GPT-4o on price while offering 8× the context. Gemini 2.0 Flash supersedes it with better quality at a slightly higher price. (Scores from respective announcements; not Benchgen measurements.)

Frequently Asked Questions

What is Gemini 1.5 Flash? Gemini 1.5 Flash is Google DeepMind's fast, production-scale model with a 1M-token context window. Released May 2024, it processes text, images, audio, and video at $0.075/$0.30 per million tokens.
How much does Gemini 1.5 Flash cost? $0.075 per million input tokens and $0.30 per million output tokens (for requests up to 128K tokens). Doubles for longer requests.

Specs sourced from Google — Gemini 1.5 Flash. Last updated 2026-06-19.

Gemini 1.5 Flash

Gemini 1.5 Flash is a large language model developed by Google.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.