Benchgen
Models/aleph-alpha/

Kolibri-1

DraftPublic

Model Details

Kolibri-1

Organization License Params Context Released

Quick answer: Kolibri-1 is Aleph Alpha's open-weight mixture-of-experts reasoning model with 78B total and 3.46B active parameters, released October 3, 2026 under Apache 2.0. It is trained from scratch on 20T tokens with a German/English focus, supports reasoning effort levels and tool calling, and a 262K-token native context validated up to 1M tokens. Aleph Alpha reports 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6 and 66.4 on SWE-bench Verified.

At a Glance

Where it leads: Strong math and coding for its active size (AIME 2026 96.0, LiveCodeBench v6 85.9) and German-language reasoning, with FP8 weights that fit on a single H200 or B200.

Where it lags: Behind Qwen3.8-27B on most of Aleph Alpha's own comparison table (e.g. Terminal-Bench 2.1 27.7 vs 76.8, Humanity's Last Exam 21.5 vs 35.6), and weak on AA-Omniscience factual recall (accuracy 14.8).

Best for: Sovereign German/English assistants, retrieval-augmented generation and agentic tool calling where inference cost per token matters.

What Kolibri-1 Is

Kolibri-1 is a 50-layer MoE transformer with 384 experts per layer (top-6 routed plus one shared) and a 4:1 sliding-window-to-full-attention layout, trained with Muon on 768 NVIDIA B200 GPUs (about 21 days of pre-training, 6.4e23 FLOPs). Post-training combined supervised fine-tuning with asynchronous RL across reasoning, tool-calling, software-engineering, terminal and retrieval environments, including German-language environments with a language-consistency reward.

It exposes explicit reasoning effort (none, low, medium, high) through the chat template, ships with a custom German-aware tokenizer (UniBPE, 128K vocabulary), and is released as FP8 weights with a BF16 variant. Aleph Alpha is a signatory of the EU GPAI Code of Practice.

Specifications

FieldValue
OrganizationAleph Alpha
Hugging FaceAleph-Alpha/Kolibri-1
ArchitectureMixture-of-experts, 78B total / 3.46B active
LanguagesGerman, English
Context length262,144 native; validated to 1,048,576
LicenseApache 2.0
Release date2026-10-03
Knowledge cutoff2026-06-18
Reasoning modeYes (none / low / medium / high)

Public Benchmark Scores

BenchmarkScoreSourceDate
GPQA Diamond84.3Model card2026-10
Humanity's Last Exam21.5Model card2026-10
MMLU-Pro (CoT)80.0Model card2026-10
AIME 202696.0Model card2026-10
AIME 202596.9Model card2026-10
LiveCodeBench v685.9Model card2026-10
HumanEval+92.7Model card2026-10
SWE-bench Verified66.4Model card2026-10
Terminal-Bench 2.127.7Model card2026-10
IFBench (loose-prompt)78.1Model card2026-10
BFCL v4 (overall)61.4Model card2026-10
AA-LCR68.3Model card2026-10
Tau2-Bench Telecom94.7Model card2026-10
Tau2-Bench Retail69.9Model card2026-10
Tau2-Bench Airline76.7Model card2026-10

Self-reported by Aleph Alpha; not independent Benchgen measurements. The card also reports BrowseComp 29.4, Tau3-Bench Banking 38.1, FRAMES, SealQA, MuSiQue, LongBench Pro, RULER, HELMET and German-language variants of several benchmarks; these are not added because the setup differs from the Benchgen pages (e.g. Qwen3.5-35B-A3B's BrowseComp is 36.5 on this card vs 61 on Benchgen) or no matching page exists.

Frequently Asked Questions

What is Kolibri-1? Kolibri-1 is Aleph Alpha's open-weight 78B mixture-of-experts reasoning model (3.46B active parameters) for German and English, released October 3, 2026 under Apache 2.0.
What is Kolibri-1's context length? 262,144 tokens natively, with quality and serving efficiency validated up to 1,048,576 tokens.
Can I run Kolibri-1 locally? Yes. FP8 weights need about 78 GB of memory, for example 2x H100, one H200, or one B200, using the aleph-alpha-inference vLLM plugin.