Benchgen
Models/moonshot-ai/

Kimi K2 Instruct

DraftPublic

Model Details

Kimi K2 Instruct

Organization License Weights Released

Quick answer: Kimi K2 Instruct is Moonshot AI's July 2025 open-weight model scoring 93.3% on HumanEval, 81.1% on MMLU-Pro, 59.1% on BFCL-v3, and 9.8% on AetherCode. Under Apache 2.0 with open weights, it delivers top-tier function-calling and academic knowledge benchmarks for an open model.

At a Glance

Where Kimi K2 Instruct leads

  • 93.3% HumanEval — highest among open-weight models on this benchmark
  • 81.1% MMLU-Pro — matches Llama 4 Maverick (80.5%) on academic knowledge
  • 59.1% BFCL-v3 — strong tool use and function calling capability
  • Apache 2.0 — unrestricted commercial use
  • March 2025 knowledge cutoff — more recent than most 2024-era open models

Where it lags

  • 9.8% AetherCode — below Qwen3 32B (16.3%) on agentic coding tasks
  • Architecture/parameters undisclosed — harder to predict inference requirements
  • Moonshot AI is less established than Meta/Alibaba in the open-weight ecosystem

Best for: Open-weight deployments requiring strong function calling (59.1% BFCL-v3), high coding completion (93.3% HumanEval), and broad academic knowledge (81.1% MMLU-Pro).

What Kimi K2 Instruct Is

Kimi K2 Instruct is Moonshot AI's July 2025 instruction-tuned open-weight model. Moonshot AI, the Chinese AI lab behind the Kimi consumer product, released K2 with Apache 2.0 licensing — a notable choice that enables unrestricted commercial deployment.

The model's 93.3% HumanEval score is exceptional for an open-weight model, and its 59.1% BFCL-v3 result demonstrates strong function-calling capability — important for tool-using agent deployments. At 81.1% MMLU-Pro, it matches Llama 4 Maverick on academic knowledge benchmarks.

The 9.8% AetherCode score is lower than expected given its HumanEval performance, suggesting that Kimi K2 Instruct is strong at code completion and synthesis tasks but less effective at the end-to-end agentic software engineering measured by AetherCode.

Specifications

FieldValue
OrganizationMoonshot AI
ParametersUndisclosed
LicenseApache 2.0
HuggingFaceKimi/K2-Instruct
Release dateJuly 1, 2025
Knowledge cutoffMarch 2025
ModalityText only

Pricing

Kimi K2 Instruct is available as open weights under Apache 2.0 — free to self-host. Available via Moonshot AI API and third-party providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval93.3%Benchgen evaluation2025-07
MMLU-Pro81.1%Benchgen evaluation2025-07
BFCL v359.1%Benchgen evaluation2025-07
AetherCode9.8%Benchgen evaluation2025-07

Kimi K2 Instruct vs Alternatives

ModelHumanEvalMMLU-ProBFCL-v3License
Kimi K2 Instruct93.3%81.1%59.1%Apache 2.0
Llama 4 Maverick80.5%Llama 4
Phi-482.6%70.4%Apache 2.0
Qwen3 32BApache 2.0

Kimi K2's 93.3% HumanEval is the highest we track for any open-weight model, and its 59.1% BFCL-v3 indicates strong function-calling performance. For applications requiring agentic end-to-end coding, its 9.8% AetherCode is a limitation; Qwen3 32B (16.3%) performs better on that benchmark.

Frequently Asked Questions

What is Kimi K2 Instruct? Kimi K2 Instruct is Moonshot AI's July 2025 Apache 2.0 open-weight model, scoring 93.3% HumanEval, 81.1% MMLU-Pro, and 59.1% BFCL-v3. Designed for strong code completion and function calling.
Is Kimi K2 Instruct open source? Yes. Kimi K2 Instruct is released under Apache 2.0 — free for commercial use without restrictions.
What is Kimi K2 Instruct strong at? HumanEval (93.3%), MMLU-Pro (81.1%), and BFCL-v3 function calling (59.1%). It is weaker at AetherCode (9.8%) — end-to-end agentic coding tasks.

Specs from Moonshot AI's official Kimi K2 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.