Quick answer: Kimi K2 Instruct is Moonshot AI's July 2025 open-weight model scoring 93.3% on HumanEval, 81.1% on MMLU-Pro, 59.1% on BFCL-v3, and 9.8% on AetherCode. Under Apache 2.0 with open weights, it delivers top-tier function-calling and academic knowledge benchmarks for an open model.
Where Kimi K2 Instruct leads
Where it lags
Best for: Open-weight deployments requiring strong function calling (59.1% BFCL-v3), high coding completion (93.3% HumanEval), and broad academic knowledge (81.1% MMLU-Pro).
Kimi K2 Instruct is Moonshot AI's July 2025 instruction-tuned open-weight model. Moonshot AI, the Chinese AI lab behind the Kimi consumer product, released K2 with Apache 2.0 licensing — a notable choice that enables unrestricted commercial deployment.
The model's 93.3% HumanEval score is exceptional for an open-weight model, and its 59.1% BFCL-v3 result demonstrates strong function-calling capability — important for tool-using agent deployments. At 81.1% MMLU-Pro, it matches Llama 4 Maverick on academic knowledge benchmarks.
The 9.8% AetherCode score is lower than expected given its HumanEval performance, suggesting that Kimi K2 Instruct is strong at code completion and synthesis tasks but less effective at the end-to-end agentic software engineering measured by AetherCode.
| Field | Value |
|---|---|
| Organization | Moonshot AI |
| Parameters | Undisclosed |
| License | Apache 2.0 |
| HuggingFace | Kimi/K2-Instruct |
| Release date | July 1, 2025 |
| Knowledge cutoff | March 2025 |
| Modality | Text only |
Kimi K2 Instruct is available as open weights under Apache 2.0 — free to self-host. Available via Moonshot AI API and third-party providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 93.3% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 81.1% | Benchgen evaluation | 2025-07 |
| BFCL v3 | 59.1% | Benchgen evaluation | 2025-07 |
| AetherCode | 9.8% | Benchgen evaluation | 2025-07 |
| Model | HumanEval | MMLU-Pro | BFCL-v3 | License |
|---|---|---|---|---|
| Kimi K2 Instruct | 93.3% | 81.1% | 59.1% | Apache 2.0 |
| Llama 4 Maverick | — | 80.5% | — | Llama 4 |
| Phi-4 | 82.6% | 70.4% | — | Apache 2.0 |
| Qwen3 32B | — | — | — | Apache 2.0 |
Kimi K2's 93.3% HumanEval is the highest we track for any open-weight model, and its 59.1% BFCL-v3 indicates strong function-calling performance. For applications requiring agentic end-to-end coding, its 9.8% AetherCode is a limitation; Qwen3 32B (16.3%) performs better on that benchmark.
Specs from Moonshot AI's official Kimi K2 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.