| Rank | Model | Score |
|---|---|---|
| 1 | clef | 61.21 |
| 2 | clef-flash | 57.07 |
1 phaseActive
Community benchmark suite for decision models — typed questions in, calibrated probabilities out — reported as a chance-corrected aggregate (0 = random, 100 = perfect).
Quick answer: Decision Index (version 0.2.1) is a community-maintained suite that scores "decision models" — models that read a state plus a schema of typed questions and return a calibrated probability for every option in a single forward pass. Results are a chance-corrected aggregate over 38 sub-benchmarks (0 = random guessing, 100 = perfect). Cloudflare's Clef leads at 61.21, self-reported.
What it tests: How accurately and how well-calibrated a model is when it must make bounded, structured decisions (choice, score, true/false questions) rather than write free-form text — across tool selection, intent classification, NLI, reasoning, retrieval and safety sub-benchmarks.
Why it matters: Decision models are a new class of small, fast classifiers designed to sit inside agent workflows. Decision Index gives them a shared yardstick that also reports latency, so accuracy can be weighed against speed.
Known limitations: The suite is community-run and unofficial. Clef and Clef-Flash were run by their authors (Cloudflare) on 36 of 38 benchmarks, with the missing ones scored as 0, and have not been reproduced by the upstream board.
Decision models, popularised by TypeSafe AI's Jev, return typed answers — a named choice, an ordinal score, or a true/false probability — for every question asked about a given state. Decision Index aggregates dozens of existing public benchmarks (BFCL, ToolRet, API-Bank, BANKING77, CLINC150, MMLU, GPQA Diamond, GSM8K, MuSR, ForecastBench and others) adapted to this answer format.
Scores are chance-corrected so that a model guessing at random scores 0 and a perfect model scores 100, which keeps easy and hard sub-benchmarks comparable. Per-benchmark breakdowns and median/p95 request latency are reported alongside the aggregate.
| Field | Value |
|---|---|
| Task category | Decision / classification |
| Metric | Chance-corrected aggregate score (0-100) |
| Version | 0.2.1 |
| Sub-benchmarks | 38 |
| Created by | multimodalart and community contributors |
| Project | Hugging Face Space |
Each sub-benchmark's metric (accuracy, macro-F1, nDCG@10, Brier score) is converted to a chance-corrected score, and the aggregate averages across categories. Latency (median and p95 ms per request) is reported separately and does not enter the score.
| Rank | Model | Decision Index | Source | Date |
|---|---|---|---|---|
| 1 | Clef | 61.21 (self-reported) | Cloudflare decision leaderboard | 2026-10 |
| 2 | Jev (TypeSafe AI, closed API) | 57.91 | Cloudflare decision leaderboard | 2026-10 |
| 3 | Surogate Rune 26B-A4B v3 | 57.44 | Cloudflare decision leaderboard | 2026-10 |
| 4 | Decider chat (Gemma-4-31B) | 57.33 | Cloudflare decision leaderboard | 2026-10 |
| 5 | Clef-Flash | 57.07 (self-reported) | Cloudflare decision leaderboard | 2026-10 |
Community results snapshot 28 September 2026, mirrored on Cloudflare's leaderboard (built 1 October 2026). Only models with a Benchgen page are linked; the rest are shown for context.
No Benchgen results yet — be the first to run Decision Index.
| Benchmark | What it tests | Output format |
|---|---|---|
| Decision Index | Calibrated structured decisions across 38 sub-benchmarks | Typed probabilities |
| JevBench | Decision accuracy and calibration on a fixed item set | Typed probabilities |
| BFCL v4 | Function-calling accuracy | Generated tool calls |
Benchgen can track decision-model performance over time alongside the community-reported scores above.