Benchgen
Models/moonshot/

Kimi K2-Thinking-0905

DraftPublic

Model Details

Kimi K2 Thinking 0905

Organization License Released Reasoning

Quick answer: Kimi K2 Thinking 0905 is Moonshot AI's September 2025 reasoning checkpoint of Kimi K2, scoring 60.2% BrowseComp, 51.0% HLE, 84.6% MMLU-Pro, 47.1% TerminalBench, and 44.8% SciCode. Apache 2.0 — notable as an open-weight model with these frontier-class reasoning scores.

At a Glance

Where Kimi K2 Thinking 0905 leads

  • 60.2% BrowseComp — strong web navigation/research
  • 51.0% HLE — among top open-weight scores on Humanity's Last Exam
  • 84.6% MMLU-Pro — strong academic knowledge with reasoning
  • 47.1% TerminalBench — solid terminal task performance
  • Apache 2.0 — fully open for commercial use

Where it lags

  • 44.8% SciCode — moderate scientific coding
  • MoE with thinking overhead — higher latency than instruct variants
  • Text-only
  • July 2025 release — pre-dates some newer frontier models

Best for: Open-source frontier-class reasoning; web research pipelines; academic knowledge tasks; teams needing Apache 2.0 models with o3-tier reasoning.

What Kimi K2 Thinking 0905 Is

Kimi K2 Thinking 0905 is the thinking (chain-of-thought reasoning) variant of Moonshot AI's Kimi K2 model family, at the September 5, 2025 checkpoint. Kimi K2 is a large MoE model; the "Thinking" variant extends it with a reasoning mode similar to DeepSeek-R1 and QwQ-32B.

The BrowseComp score (60.2%) is notable — this benchmark tests multi-step web research ability, and 60% is in the frontier tier. Combined with 51.0% HLE, this is one of the strongest open-weight reasoning models from mid-2025.

Specifications

FieldValue
OrganizationMoonshot AI
LicenseApache 2.0
Release dateJuly 2025 (checkpoint: Sept 2025)
ArchitectureMoE with chain-of-thought reasoning
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no license cost. Also available via Moonshot AI API.

Public Benchmark Scores

BenchmarkScoreSourceDate
BrowseComp60.2%Benchgen evaluation2025-09
Humanity's Last Exam51.0%Benchgen evaluation2025-09
MMLU-Pro84.6%Benchgen evaluation2025-09
TerminalBench47.1%Benchgen evaluation2025-09
SciCode44.8%Benchgen evaluation2025-09

Kimi K2 Thinking 0905 vs Alternatives

ModelHLEBrowseCompMMLU-ProLicense
Kimi K2 Thinking 090551.0%60.2%84.6%Apache 2.0
Kimi K2 Instruct 090581.1%Apache 2.0
GLM-5.254.7%Proprietary
GPT-5.541.4%84.4%Proprietary

Kimi K2 Thinking 0905 leads open-weight models on HLE (51.0%) and BrowseComp (60.2%). It beats GPT-5.5 on HLE (51.0% vs 41.4%) while remaining Apache 2.0. GLM-5.2 leads on HLE (54.7%) but is proprietary.

Frequently Asked Questions

What is Kimi K2 Thinking 0905? Moonshot AI's September 2025 reasoning checkpoint of Kimi K2, scoring 60.2% BrowseComp, 51.0% HLE, and 84.6% MMLU-Pro. Apache 2.0 MoE with chain-of-thought reasoning.
How does Kimi K2 Thinking 0905 compare to GLM-5.2 on HLE? GLM-5.2 leads (54.7% vs 51.0%) but is proprietary. Kimi K2 Thinking 0905 is the top open-weight (Apache 2.0) model for HLE-class reasoning tasks.

Specs from Moonshot AI's Kimi K2 Thinking 0905 release and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.