Benchgen
Models/ibm/

IBM Granite 4.2 30B

DraftPublic

Model Details

IBM Granite 4.2 30B

Organization Params License Released

Quick answer: IBM Granite 4.2 30B is the flagship model in IBM's reasoning-focused Granite 4.2 family — a dense, Apache 2.0 model with the deepest agentic-RL training of the three sizes (an extra second-phase SWE-focused SFT pass on top of the shared SWE/Terminal/Search RL block). It leads the family on every reported benchmark: 57.0% SWE-bench Verified, 33.3% SWE-bench Pro, 89.2% AIME25, and 61.4% BFCL-v4, with a native 128K-token context.

At a Glance

Where Granite 4.2 30B leads

  • Best-in-family agentic coding: 41.9% SWE-bench Multilingual, 33.3% SWE-bench Pro, 57.0% SWE-bench Verified, 29.2% Terminal-Bench 2.1 — the largest agentic-RL investment of the three sizes, including an extra 30B-only second-phase SFT pass focused on agentic coding
  • 89.2% AIME25 and 89.2% HMMT Feb25 — ties as the strongest competition-math reasoner in the family
  • 61.4% BFCL-v4 — most reliable native tool-calling of the three sizes

Where it lags

  • Still trails frontier closed models and larger MoE competitors on raw agentic-coding ceiling (e.g. SWE-bench Pro at 33.3% vs. 50%+ scores from much larger models)
  • Dense 30B architecture — no MoE efficiency gains; higher inference cost per token than similarly-capable MoE models
  • Text-only — no native vision/multimodal support

Best for: Enterprise flagship agentic coding, terminal-operation, and deep-research workflows — teams wanting IBM's strongest agentic-RL-trained model with full Apache 2.0 licensing.

What Granite 4.2 30B Is

Granite 4.2 30B is the largest and most capable model in IBM's three-model Granite 4.2 release. It shares the same dense decoder-only architecture and five-phase pre-training (~15T tokens) as its 3B and 8B siblings, but goes furthest up IBM's post-training ladder: three rounds of foundational RLVR (vs. two for the smaller sizes), the full agentic-RL block (SWE-agent → Terminal-agent → Search-agent, GRPO-trained on real sandboxed environments), and a 30B-exclusive second SFT phase that upsamples agentic/SWE/coding trajectories for an additional epoch before the final RLHF alignment pass.

Training data includes 1 trillion tokens of synthetic code from IBM's CodeAlchemy pipeline, plus a speculative-decoding layer for faster serving. Like all Granite 4.2 models, it supports a thinking/non-thinking/low-effort reasoning switch and native OpenAI-format tool calling, and integrates directly with OpenHands, OpenCode, and Pi agentic harnesses.

Specifications

FieldValue
OrganizationIBM
Parameters30B (dense)
ArchitectureDecoder-only dense transformer, GQA (32 attention heads / 8 KV heads, head size 128), RoPE (θ=10,000,000), SwiGLU MLP (hidden size 32,768), RMSNorm, 64 layers
Context window131,072 tokens (128K)
LicenseApache 2.0
Release dateAugust 2026
ModalityText only

Pricing

Open weights available on Hugging Face (ibm-granite/granite-4.2-30b) — free to self-host under Apache 2.0. Also available via CoreWeave Inference, DeepInfra, OpenRouter, Replicate, and watsonx. FP8, NVFP4, MXFP4, and GGUF quantized variants are published.

Context Window

Granite 4.2 30B natively supports 131,072 tokens (128K).

Public Benchmark Scores

BenchmarkScoreSourceDate
SWE-bench Multilingual41.89%Granite 4.2 launch blog (Hugging Face)2026-08
SWE-bench Pro33.29%Granite 4.2 launch blog (Hugging Face)2026-08
SWE-bench Verified57.00%Granite 4.2 launch blog (Hugging Face)2026-08
TerminalBench 2.129.24%Granite 4.2 launch blog (Hugging Face)2026-08
BFCL-v461.39%Granite 4.2 launch blog (Hugging Face)2026-08
Bird-SQL (dev)0.4185Granite 4.2 launch blog (Hugging Face)2026-08
GDPVal-AA v21225 (Elo)Granite 4.2 launch blog (Hugging Face)2026-08
AIME 202589.17%Granite 4.2 launch blog (Hugging Face)2026-08
HMMT 202589.17%Granite 4.2 launch blog (Hugging Face)2026-08
LiveCodeBench v60.7577Granite 4.2 launch blog (Hugging Face)2026-08
SciCode38.76%Granite 4.2 launch blog (Hugging Face)2026-08
MMLU-Pro77.60%Granite 4.2 launch blog (Hugging Face)2026-08
Arena-Hard-V20.6793Granite 4.2 launch blog (Hugging Face)2026-08
IFBench77.17%Granite 4.2 launch blog (Hugging Face)2026-08

Scores are self-reported by IBM Research on the official Hugging Face launch blog (huggingface.co/blog/ibm-granite/granite-4-2). Bird-SQL and GDPVal-AA v2 are reported by IBM under the shorthand names "BirdBench" (41.85%, converted to Benchgen's 0-1 fractional scale) and "GDPval" (1225, an Elo rating matching Benchgen's GDPVal-AA v2 anchored-Elo scale and score range). IFBench is reported as "IFBench (prompt)" — the prompt-level-loose submetric, matching Benchgen's ifbench metric exactly. TerminalBench maps to IBM's "Terminal-Bench 2.1" — Benchgen's terminalbench page tracks version 2.1 specifically, an exact match.

Additional benchmarks reported by IBM, not added to Benchgen:

  • τ³-bench (62.00) — no domain specified in IBM's table; Benchgen's only τ³ page (tau3-banking) is a much lower-ceiling banking-specific split (max observed 33.4), indicating a scale/scope mismatch — skipped.
  • ProfBench (42.90) — no Benchgen page exists.
  • GPQA (66.41, plain — not "GPQA Diamond") — Benchgen only tracks the Diamond subset; skipped to avoid conflating distinct question sets.
  • MMLU-ProX lite (IBM) (66.64) — explicitly labeled an IBM in-house variant.
  • RULER 64K (89.96) and RULER 128K (81.38) — no Benchgen page exists for RULER yet.

Frequently Asked Questions

What is IBM Granite 4.2 30B? IBM's flagship Granite 4.2 model — a dense, Apache 2.0 reasoning model with the deepest agentic-RL training in the family, built for enterprise coding, terminal-operation, and research-agent workflows.
How is Granite 4.2 30B different from the 8B model? Both complete the same agentic-RL block (SWE-agent, Terminal-agent, Search-agent), but the 30B model additionally runs a 30B-exclusive second SFT phase upsampling agentic/SWE data, plus an extra round of foundational RLVR — leading it on every reported benchmark.
Is IBM Granite 4.2 30B open source? Yes — released under Apache 2.0, with no restrictions on fine-tuning or commercial use.

Specs from IBM's official Hugging Face launch blog (huggingface.co/blog/ibm-granite/granite-4-2) and IBM Research blog (research.ibm.com/blog/introducing-granite-4-2). Last updated 2026-08-31.