Benchgen

CritPt — Results

RankModelScore
1kimi-k323.4
C

CritPt

1 phaseActive

Reasoning benchmark targeting critical decision points where careless analysis leads to error. Metric: % accuracy.

Overview

CritPt

Category Metric Saturation Created

Quick answer: CritPt is a reasoning benchmark that targets "critical points" in multi-step problems — decision junctures where careless or shallow analysis leads models astray, even if they can handle the surrounding steps correctly. Kimi K3 scores 23.4% as of July 2026, reflecting the benchmark's deliberately high difficulty.

At a Glance

What it tests: A model's ability to correctly navigate critical decision points within complex reasoning chains, where small analytical errors compound into wrong final answers.

Why it matters: Aggregate accuracy on general reasoning benchmarks can mask specific failure points. CritPt isolates the moments where reasoning most commonly breaks down, providing a more diagnostic signal.

Known limitations: As an emerging benchmark cited in Kimi K3's own technical report, independent documentation of its exact task composition and methodology is not yet widely published.

What CritPt Measures

CritPt evaluates whether a model can identify and correctly resolve critical points within multi-step reasoning tasks — junctures where a plausible but incorrect line of reasoning is especially tempting. Rather than scoring only the final answer, the design of critical-point-style benchmarks emphasizes correctness at the specific step where most reasoning failures actually occur, making low scores (like Kimi K3's 23.4%) expected even for frontier models.

Benchmark Specifications

FieldValue
Task categoryReasoning
Metric% accuracy
SaturationLow
Created byNot yet independently documented

How CritPt Is Scored

Models are scored on % accuracy for correctly resolving the critical reasoning point within each task, rather than partial credit for surrounding steps.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K323.4%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

CritPt on Benchgen

No Benchgen results yet — be the first to run CritPt.

CritPt vs Other Benchmarks

BenchmarkWhat it testsSaturation
CritPtCritical-point reasoning accuracyLow
Humanity's Last ExamFrontier-hard knowledge & reasoningLow
GPQA DiamondGraduate-level science reasoningLow
ARC-AGI v2Abstract reasoningLow

Run CritPt on Your Model

Benchgen lets you run CritPt against your own model, tracking accuracy at critical reasoning junctures to diagnose where your model's reasoning chains are most likely to break down.

Frequently Asked Questions

What is CritPt? CritPt is a reasoning benchmark that isolates "critical points" in multi-step problems — the specific junctures where careless analysis most commonly causes models to fail.
What does a good score look like on CritPt? Kimi K3 reports 23.4% as of July 2026. Low absolute scores are expected given the benchmark's focus on the hardest failure points in reasoning chains.
Who created CritPt? CritPt's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.