Quick answer: IBM Granite 3.3 8B Instruct is IBM's April 2025 compact model scoring 89.7% HumanEval, 80.9% GSM8K, and 57.6% Arena Hard. Apache 2.0 licensed with strong coding performance for an 8B model.
Where Granite 3.3 8B Instruct leads
Where it lags
Best for: Enterprise coding assistants; IBM Cloud deployments; regulated industries requiring Apache 2.0 small models with strong HumanEval scores.
Granite 3.3 8B Instruct is IBM's April 2025 update to the Granite 3 model family. IBM's Granite models are designed with enterprise use cases in mind: transparency (IBM publishes training data details and model cards), safety filtering, and compliance with enterprise data policies.
The 89.7% HumanEval score is exceptional for an 8B model — placing it above many larger open-source models from 2024. This reflects IBM's focus on coding capability in its developer tooling (Granite was built for tasks like code completion, bug fixing, and documentation).
| Field | Value |
|---|---|
| Organization | IBM |
| License | Apache 2.0 |
| HuggingFace | ibm-granite/granite-3.3-8b-instruct |
| Release date | April 2025 |
| Parameters | 8B |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no license cost. Available via IBM watsonx.ai and cloud providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 89.7% | Benchgen evaluation | 2025-04 |
| GSM8K | 80.9% | Benchgen evaluation | 2025-04 |
| Arena Hard | 57.6% | Benchgen evaluation | 2025-04 |
| Model | HumanEval | GSM8K | Params | License |
|---|---|---|---|---|
| Granite 3.3 8B Instruct | 89.7% | 80.9% | 8B | Apache 2.0 |
| Llama 3.1 8B Instruct | — | 84.5% | 8B | Llama 3.1 |
| Gemma 3 12B | 85.4% | — | 12B | Gemma ToU |
| Phi-4 | — | — | 14B | MIT |
Granite 3.3 8B Instruct leads on HumanEval (89.7%) among compact open-weight models. For pure coding in enterprise/IBM environments: Granite 3.3 8B. For broader instruction quality: Llama 3.1 8B Instruct.
Specs from IBM's Granite 3.3 8B Instruct release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.