Quick answer: Kolibri-1 is Aleph Alpha's open-weight mixture-of-experts reasoning model with 78B total and 3.46B active parameters, released October 3, 2026 under Apache 2.0. It is trained from scratch on 20T tokens with a German/English focus, supports reasoning effort levels and tool calling, and a 262K-token native context validated up to 1M tokens. Aleph Alpha reports 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6 and 66.4 on SWE-bench Verified.
Where it leads: Strong math and coding for its active size (AIME 2026 96.0, LiveCodeBench v6 85.9) and German-language reasoning, with FP8 weights that fit on a single H200 or B200.
Where it lags: Behind Qwen3.8-27B on most of Aleph Alpha's own comparison table (e.g. Terminal-Bench 2.1 27.7 vs 76.8, Humanity's Last Exam 21.5 vs 35.6), and weak on AA-Omniscience factual recall (accuracy 14.8).
Best for: Sovereign German/English assistants, retrieval-augmented generation and agentic tool calling where inference cost per token matters.
Kolibri-1 is a 50-layer MoE transformer with 384 experts per layer (top-6 routed plus one shared) and a 4:1 sliding-window-to-full-attention layout, trained with Muon on 768 NVIDIA B200 GPUs (about 21 days of pre-training, 6.4e23 FLOPs). Post-training combined supervised fine-tuning with asynchronous RL across reasoning, tool-calling, software-engineering, terminal and retrieval environments, including German-language environments with a language-consistency reward.
It exposes explicit reasoning effort (none, low, medium, high) through the chat template, ships with a custom German-aware tokenizer (UniBPE, 128K vocabulary), and is released as FP8 weights with a BF16 variant. Aleph Alpha is a signatory of the EU GPAI Code of Practice.
| Field | Value |
|---|---|
| Organization | Aleph Alpha |
| Hugging Face | Aleph-Alpha/Kolibri-1 |
| Architecture | Mixture-of-experts, 78B total / 3.46B active |
| Languages | German, English |
| Context length | 262,144 native; validated to 1,048,576 |
| License | Apache 2.0 |
| Release date | 2026-10-03 |
| Knowledge cutoff | 2026-06-18 |
| Reasoning mode | Yes (none / low / medium / high) |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GPQA Diamond | 84.3 | Model card | 2026-10 |
| Humanity's Last Exam | 21.5 | Model card | 2026-10 |
| MMLU-Pro (CoT) | 80.0 | Model card | 2026-10 |
| AIME 2026 | 96.0 | Model card | 2026-10 |
| AIME 2025 | 96.9 | Model card | 2026-10 |
| LiveCodeBench v6 | 85.9 | Model card | 2026-10 |
| HumanEval+ | 92.7 | Model card | 2026-10 |
| SWE-bench Verified | 66.4 | Model card | 2026-10 |
| Terminal-Bench 2.1 | 27.7 | Model card | 2026-10 |
| IFBench (loose-prompt) | 78.1 | Model card | 2026-10 |
| BFCL v4 (overall) | 61.4 | Model card | 2026-10 |
| AA-LCR | 68.3 | Model card | 2026-10 |
| Tau2-Bench Telecom | 94.7 | Model card | 2026-10 |
| Tau2-Bench Retail | 69.9 | Model card | 2026-10 |
| Tau2-Bench Airline | 76.7 | Model card | 2026-10 |
Self-reported by Aleph Alpha; not independent Benchgen measurements. The card also reports BrowseComp 29.4, Tau3-Bench Banking 38.1, FRAMES, SealQA, MuSiQue, LongBench Pro, RULER, HELMET and German-language variants of several benchmarks; these are not added because the setup differs from the Benchgen pages (e.g. Qwen3.5-35B-A3B's BrowseComp is 36.5 on this card vs 61 on Benchgen) or no matching page exists.