Benchgen
Models/moonshot/

Kimi K2.5

DraftPublic

Model Details

Kimi K2.5

Organization License Released

Quick answer: Kimi K2.5 is Moonshot AI's July 2025 model scoring 74.9% BrowseComp, 76.8% SWE-Bench Verified, 95.8% AIME 2026, 87.6% GPQA Diamond, and 87.1% MMLU-Pro. Apache 2.0 — strong baseline before K2.6 release.

At a Glance

Where Kimi K2.5 leads

  • 74.9% BrowseComp — strong web research
  • 76.8% SWE-Bench Verified — competitive software engineering
  • 95.8% AIME 2026 — excellent competition mathematics
  • 87.1% MMLU-Pro — strong academic knowledge
  • Apache 2.0 — fully open

Where it lags

  • Superseded by Kimi K2.6 on all key metrics
  • 24.4% HLE — limited frontier reasoning
  • 36.9% SimpleQA — limited factual accuracy

Best for: Teams locked to a K2.5 checkpoint; open-source deployments requiring strong BrowseComp/SWE balance before K2.6.

What Kimi K2.5 Is

Kimi K2.5 is the July 2025 checkpoint of Moonshot AI's Kimi K2 MoE model, predating the K2.6 release. It shows the strong capabilities of the K2 family before the K2.6 improvements: 74.9% BrowseComp, 76.8% SWE-Bench, and 95.8% AIME represent a highly competitive set of open-weight benchmarks.

The model is superseded by K2.6 for most use cases. K2.5 may be preferred where a stable, vetted checkpoint is required or where K2.6's additional LHTB/safety scores have not yet been validated.

Specifications

FieldValue
OrganizationMoonshot AI
LicenseApache 2.0
Release dateJuly 2025
ArchitectureMoE
ModalityText only

Pricing

Open weights under Apache 2.0. Also available via Moonshot AI API.

Public Benchmark Scores

BenchmarkScoreSourceDate
BrowseComp74.9%Benchgen evaluation2025-07
SWE-Bench Verified76.8%Benchgen evaluation2025-07
AIME 202695.8%Benchgen evaluation2025-07
GPQA Diamond87.6%Benchgen evaluation2025-07
Humanity's Last Exam24.4%Benchgen evaluation2025-07
MMLU-Pro87.1%Benchgen evaluation2025-07
SimpleQA36.9%Benchgen evaluation2025-07

Kimi K2.5 vs Alternatives

ModelBrowseCompSWE-BenchMMLU-ProLicense
Kimi K2.574.9%76.8%87.1%Apache 2.0
Kimi K2.683.2%80.2%Apache 2.0
Kimi K2 Thinking 090560.2%84.6%Apache 2.0
DeepSeek-R1 052885%MIT

Kimi K2.5 vs K2.6: K2.6 is better on all metrics. Use K2.5 only when K2.6 is not available or when a specific checkpoint is required.

Frequently Asked Questions

What is Kimi K2.5? Moonshot AI's July 2025 MoE model scoring 74.9% BrowseComp, 76.8% SWE-Bench, 95.8% AIME, and 87.1% MMLU-Pro. Apache 2.0.
Should I use Kimi K2.5 or Kimi K2.6? Kimi K2.6 for all new deployments — it improves on K2.5 across all benchmarks. K2.5 only if K2.6 is not yet available in your deployment environment.

Specs from Moonshot AI's Kimi K2.5 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.