| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 23.4 |
1 phaseActive
Reasoning benchmark targeting critical decision points where careless analysis leads to error. Metric: % accuracy.
Quick answer: CritPt is a reasoning benchmark that targets "critical points" in multi-step problems — decision junctures where careless or shallow analysis leads models astray, even if they can handle the surrounding steps correctly. Kimi K3 scores 23.4% as of July 2026, reflecting the benchmark's deliberately high difficulty.
What it tests: A model's ability to correctly navigate critical decision points within complex reasoning chains, where small analytical errors compound into wrong final answers.
Why it matters: Aggregate accuracy on general reasoning benchmarks can mask specific failure points. CritPt isolates the moments where reasoning most commonly breaks down, providing a more diagnostic signal.
Known limitations: As an emerging benchmark cited in Kimi K3's own technical report, independent documentation of its exact task composition and methodology is not yet widely published.
CritPt evaluates whether a model can identify and correctly resolve critical points within multi-step reasoning tasks — junctures where a plausible but incorrect line of reasoning is especially tempting. Rather than scoring only the final answer, the design of critical-point-style benchmarks emphasizes correctness at the specific step where most reasoning failures actually occur, making low scores (like Kimi K3's 23.4%) expected even for frontier models.
| Field | Value |
|---|---|
| Task category | Reasoning |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models are scored on % accuracy for correctly resolving the critical reasoning point within each task, rather than partial credit for surrounding steps.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 23.4% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run CritPt.
| Benchmark | What it tests | Saturation |
|---|---|---|
| CritPt | Critical-point reasoning accuracy | Low |
| Humanity's Last Exam | Frontier-hard knowledge & reasoning | Low |
| GPQA Diamond | Graduate-level science reasoning | Low |
| ARC-AGI v2 | Abstract reasoning | Low |
Benchgen lets you run CritPt against your own model, tracking accuracy at critical reasoning junctures to diagnose where your model's reasoning chains are most likely to break down.