Quick answer: Gemma 3 27B is Google's March 2025 open-weight multimodal model, scoring 87.8% on HumanEval, 67.5% on MMLU-Pro, 89.0% on MATH, and 26.0% on BigCodeBench. With a 128K context window and Gemma Terms of Use, it is the largest and most capable model in the Gemma 3 family.
Where Gemma 3 27B leads
Where it lags
Best for: Open-weight multimodal tasks at 27B scale; teams within the Google/TPU ecosystem; high-quality math and code completion tasks.
Gemma 3 27B is the largest model in Google's Gemma 3 open-model family, released March 12, 2025. It introduces multimodal capability (text + image) to the Gemma 3 line, targeting vision-language tasks at the 27B scale that was previously only available in much larger proprietary models.
The model's 89.0% MATH score is notable — placing it above many models 2-3× its parameter count on competition mathematics. Combined with 87.8% HumanEval, Gemma 3 27B is a strong choice for mathematical reasoning and code completion tasks in the open-weight 27B tier.
Google deploys Gemma 3 models with JAX/TPU training, and the models are optimised for Google's infrastructure. For teams building on GCP or using Google Cloud AI Platform, Gemma 3 27B integrates well into the Google AI ecosystem.
| Field | Value |
|---|---|
| Organization | |
| Parameters | 27B (dense) |
| Context window | 128,000 tokens |
| License | Gemma Terms of Use |
| HuggingFace | google/gemma-3-27b-it |
| Release date | March 12, 2025 |
| Knowledge cutoff | September 2024 |
| Modality | Text + Vision (multimodal) |
Gemma 3 27B is available as open weights under the Gemma Terms of Use. Available via Google Cloud Vertex AI and third-party providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 87.8% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 67.5% | Benchgen evaluation | 2025-07 |
| MATH | 89.0% | Benchgen evaluation | 2025-07 |
| BigCodeBench | 26.0% | Benchgen evaluation | 2025-07 |
| Model | HumanEval | MMLU-Pro | MATH | License |
|---|---|---|---|---|
| Gemma 3 27B | 87.8% | 67.5% | 89.0% | Gemma ToU |
| Gemma 3 12B | 85.4% | 60.6% | — | Gemma ToU |
| Phi-4 | 82.6% | 70.4% | — | Apache 2.0 |
| Llama 4 Maverick | — | 80.5% | — | Llama 4 |
Gemma 3 27B vs Phi-4: higher HumanEval and MATH at 27B vs 14B — Phi-4 has higher MMLU-Pro (70.4% vs 67.5%) at smaller size. Gemma 3 27B is the preferred choice when multimodal capability is needed.
Specs from Google's official Gemma 3 release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.