Quick answer: Gemma 3 12B is Google's March 2025 mid-tier open-weight multimodal model, scoring 85.4% on HumanEval and 60.6% on MMLU-Pro. With a 128K context window under the Gemma Terms of Use, it provides strong coding and multimodal capability in a 12B parameter footprint — deployable on a single mid-range GPU.
Where Gemma 3 12B leads
Where it lags
Best for: Resource-constrained open-weight multimodal deployments; edge and on-device inference; teams that need vision at 12B scale.
Gemma 3 12B is the mid-tier model in Google's Gemma 3 family, released March 12, 2025 alongside the 1B, 4B, 12B, and 27B variants. At 12B parameters with multimodal capability and a 128K context window, it is designed for deployments where the 27B model's infrastructure requirements are too high.
Its 85.4% HumanEval score is strong for a 12B model — approaching the 87.8% of the 27B version despite having less than half the parameters. This reflects Google's training efficiency improvements in the Gemma 3 generation.
For pure text tasks requiring strong knowledge, Phi-4 (14B, 70.4% MMLU-Pro) outperforms Gemma 3 12B (60.6%) at a similar size and with a fully permissive Apache 2.0 license. Gemma 3 12B is preferred when multimodal vision capability is a requirement.
| Field | Value |
|---|---|
| Organization | |
| Parameters | 12B (dense) |
| Context window | 128,000 tokens |
| License | Gemma Terms of Use |
| HuggingFace | google/gemma-3-12b-it |
| Release date | March 12, 2025 |
| Knowledge cutoff | September 2024 |
| Modality | Text + Vision (multimodal) |
Gemma 3 12B is available as open weights under the Gemma Terms of Use. Available via Google Cloud Vertex AI and third-party providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 85.4% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 60.6% | Benchgen evaluation | 2025-07 |
| Model | HumanEval | MMLU-Pro | Vision | License |
|---|---|---|---|---|
| Gemma 3 12B | 85.4% | 60.6% | Yes | Gemma ToU |
| Gemma 3 27B | 87.8% | 67.5% | Yes | Gemma ToU |
| Phi-4 | 82.6% | 70.4% | No | Apache 2.0 |
| Qwen3 32B | — | — | No | Apache 2.0 |
For vision tasks at small scale: Gemma 3 12B is the recommended open-weight choice. For text-only tasks: Phi-4 achieves higher MMLU-Pro (70.4%) at similar size with Apache 2.0 license.
Specs from Google's official Gemma 3 release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.