Benchgen
Models/ibm/

Granite 3.3 8B Instruct

DraftPublic

Model Details

Granite 3.3 8B Instruct

Organization License Released Params

Quick answer: IBM Granite 3.3 8B Instruct is IBM's April 2025 compact model scoring 89.7% HumanEval, 80.9% GSM8K, and 57.6% Arena Hard. Apache 2.0 licensed with strong coding performance for an 8B model.

At a Glance

Where Granite 3.3 8B Instruct leads

  • 89.7% HumanEval — exceptional coding for an 8B model (matches many larger models)
  • Apache 2.0 — fully open including commercial use
  • IBM enterprise trust: built for regulated enterprise environments
  • Competitive at its parameter count

Where it lags

  • 57.6% Arena Hard — moderate on open-ended instruction quality
  • 80.9% GSM8K — good but below leading 8B models
  • 8B parameters: limited capacity vs larger models
  • Smaller context window than MoE competitors

Best for: Enterprise coding assistants; IBM Cloud deployments; regulated industries requiring Apache 2.0 small models with strong HumanEval scores.

What Granite 3.3 8B Instruct Is

Granite 3.3 8B Instruct is IBM's April 2025 update to the Granite 3 model family. IBM's Granite models are designed with enterprise use cases in mind: transparency (IBM publishes training data details and model cards), safety filtering, and compliance with enterprise data policies.

The 89.7% HumanEval score is exceptional for an 8B model — placing it above many larger open-source models from 2024. This reflects IBM's focus on coding capability in its developer tooling (Granite was built for tasks like code completion, bug fixing, and documentation).

Specifications

FieldValue
OrganizationIBM
LicenseApache 2.0
HuggingFaceibm-granite/granite-3.3-8b-instruct
Release dateApril 2025
Parameters8B
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no license cost. Available via IBM watsonx.ai and cloud providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval89.7%Benchgen evaluation2025-04
GSM8K80.9%Benchgen evaluation2025-04
Arena Hard57.6%Benchgen evaluation2025-04

Granite 3.3 8B Instruct vs Alternatives

ModelHumanEvalGSM8KParamsLicense
Granite 3.3 8B Instruct89.7%80.9%8BApache 2.0
Llama 3.1 8B Instruct84.5%8BLlama 3.1
Gemma 3 12B85.4%12BGemma ToU
Phi-414BMIT

Granite 3.3 8B Instruct leads on HumanEval (89.7%) among compact open-weight models. For pure coding in enterprise/IBM environments: Granite 3.3 8B. For broader instruction quality: Llama 3.1 8B Instruct.

Frequently Asked Questions

What is Granite 3.3 8B Instruct? IBM's April 2025 compact 8B model scoring 89.7% HumanEval, 80.9% GSM8K, and 57.6% Arena Hard. Apache 2.0 with enterprise-focused design.
Why is Granite 3.3's HumanEval so high for an 8B model? IBM specifically optimised Granite for code tasks — it is a primary use case for IBM's developer tooling (code completion, documentation, bug fixing). The 89.7% HumanEval reflects this coding-first approach.

Specs from IBM's Granite 3.3 8B Instruct release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.