Benchgen
Models/google/

Gemma 3 12B

DraftPublic

Model Details

Gemma 3 12B

Organization Parameters Context License Weights Released

Quick answer: Gemma 3 12B is Google's March 2025 mid-tier open-weight multimodal model, scoring 85.4% on HumanEval and 60.6% on MMLU-Pro. With a 128K context window under the Gemma Terms of Use, it provides strong coding and multimodal capability in a 12B parameter footprint — deployable on a single mid-range GPU.

At a Glance

Where Gemma 3 12B leads

  • 85.4% HumanEval — excellent function-level code completion for a 12B model
  • 60.6% MMLU-Pro — competitive academic knowledge at 12B scale
  • 128K context window
  • Multimodal: text and image inputs at 12B scale
  • Single-GPU deployable (in quantised form)
  • Open weights on Hugging Face

Where it lags

  • 60.6% MMLU-Pro — below Phi-4 (70.4% at 14B) despite similar size
  • Below Gemma 3 27B on all benchmarks (expected)
  • Gemma Terms of Use — not fully permissive (not Apache 2.0 / MIT)
  • Text+vision only

Best for: Resource-constrained open-weight multimodal deployments; edge and on-device inference; teams that need vision at 12B scale.

What Gemma 3 12B Is

Gemma 3 12B is the mid-tier model in Google's Gemma 3 family, released March 12, 2025 alongside the 1B, 4B, 12B, and 27B variants. At 12B parameters with multimodal capability and a 128K context window, it is designed for deployments where the 27B model's infrastructure requirements are too high.

Its 85.4% HumanEval score is strong for a 12B model — approaching the 87.8% of the 27B version despite having less than half the parameters. This reflects Google's training efficiency improvements in the Gemma 3 generation.

For pure text tasks requiring strong knowledge, Phi-4 (14B, 70.4% MMLU-Pro) outperforms Gemma 3 12B (60.6%) at a similar size and with a fully permissive Apache 2.0 license. Gemma 3 12B is preferred when multimodal vision capability is a requirement.

Specifications

FieldValue
OrganizationGoogle
Parameters12B (dense)
Context window128,000 tokens
LicenseGemma Terms of Use
HuggingFacegoogle/gemma-3-12b-it
Release dateMarch 12, 2025
Knowledge cutoffSeptember 2024
ModalityText + Vision (multimodal)

Pricing

Gemma 3 12B is available as open weights under the Gemma Terms of Use. Available via Google Cloud Vertex AI and third-party providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval85.4%Benchgen evaluation2025-07
MMLU-Pro60.6%Benchgen evaluation2025-07

Gemma 3 12B vs Alternatives

ModelHumanEvalMMLU-ProVisionLicense
Gemma 3 12B85.4%60.6%YesGemma ToU
Gemma 3 27B87.8%67.5%YesGemma ToU
Phi-482.6%70.4%NoApache 2.0
Qwen3 32BNoApache 2.0

For vision tasks at small scale: Gemma 3 12B is the recommended open-weight choice. For text-only tasks: Phi-4 achieves higher MMLU-Pro (70.4%) at similar size with Apache 2.0 license.

Frequently Asked Questions

What is Gemma 3 12B? Gemma 3 12B is Google's March 2025 open-weight multimodal 12B model, scoring 85.4% HumanEval and 60.6% MMLU-Pro with a 128K context window.
Is Gemma 3 12B open source? The weights are open under the Gemma Terms of Use on Hugging Face. Commercial use is permitted with conditions — review the license terms.
Does Gemma 3 12B support images? Yes. Gemma 3 12B supports text and image inputs (multimodal).
Should I use Gemma 3 12B or Gemma 3 27B? Gemma 3 27B achieves higher scores on all benchmarks (87.8% vs 85.4% HumanEval, 67.5% vs 60.6% MMLU-Pro). Choose 12B when the 27B infrastructure requirement (memory/compute) is the constraint.

Specs from Google's official Gemma 3 release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.