Benchgen
Models/moonshot/

Kimi K2 0905

DraftPublic

Model Details

Kimi K2 0905

Organization License Released

Quick answer: Kimi K2 0905 is Moonshot AI's September 2025 checkpoint of Kimi K2, scoring 94.5% HumanEval, 84.9% GPQA Diamond, 95.0% GSM8K, 89.1% MATH, and 53.7% LiveCodeBench. Apache 2.0.

At a Glance

Where Kimi K2 0905 leads

  • 94.5% HumanEval — excellent coding
  • 95.0% GSM8K — near-perfect grade-school math
  • 89.1% MATH — strong competition math
  • 84.9% GPQA Diamond — solid graduate-level science
  • Apache 2.0 — fully open

Where it lags

  • 53.7% LiveCodeBench — moderate real-world coding vs K2.6 (80.2% SWE-Bench)
  • September 2025 checkpoint — general K2 series base, not thinking variant

Best for: High-quality HumanEval/math pipelines in Apache 2.0; general Kimi K2 ecosystem deployments; September 2025 stable checkpoint.

What Kimi K2 0905 Is

Kimi K2 0905 is the September 5, 2025 checkpoint of Moonshot AI's Kimi K2 instruct/base model — a specific checkpoint between the K2.5 and K2.6 releases. The 0905 suffix (September 5) indicates a mid-series update.

With 94.5% HumanEval and 95.0% GSM8K, this checkpoint shows strong fundamental math and code quality. The 53.7% LiveCodeBench reflects practical coding ability, though below the later K2.6 SWE-Bench metric.

Specifications

FieldValue
OrganizationMoonshot AI
LicenseApache 2.0
Release dateJuly 2025 (checkpoint: Sept 5, 2025)
ArchitectureMoE
ModalityText only

Pricing

Open weights under Apache 2.0. Also available via Moonshot AI API.

Public Benchmark Scores

BenchmarkScoreSourceDate
HumanEval94.5%Benchgen evaluation2025-09
GSM8K95.0%Benchgen evaluation2025-09
MATH89.1%Benchgen evaluation2025-09
GPQA Diamond84.9%Benchgen evaluation2025-09
LiveCodeBench53.7%Benchgen evaluation2025-09

Kimi K2 0905 vs Alternatives

ModelHumanEvalGPQA DiamondLiveCodeBenchLicense
Kimi K2 090594.5%84.9%53.7%Apache 2.0
Kimi K2.587.6%Apache 2.0
Kimi K2 Instruct 0905Apache 2.0
Qwen2.5 Coder 32B Instruct92.7%Apache 2.0

Kimi K2 0905 vs Qwen2.5 Coder 32B: higher HumanEval (94.5% vs 92.7%) with GPQA coverage as a bonus. For Kimi-ecosystem general coding with STEM reasoning: K2 0905.

Frequently Asked Questions

What is Kimi K2 0905? Moonshot AI's September 2025 checkpoint of Kimi K2 scoring 94.5% HumanEval, 95.0% GSM8K, 89.1% MATH, 84.9% GPQA Diamond. Apache 2.0.

Specs from Moonshot AI's Kimi K2 0905 checkpoint and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.