Benchgen

Decision Index — Results

RankModelScore
1clef61.21
2clef-flash57.07
D

Decision Index

1 phaseActive

Community benchmark suite for decision models — typed questions in, calibrated probabilities out — reported as a chance-corrected aggregate (0 = random, 100 = perfect).

Overview

Decision Index

Category Metric Tasks Saturation

Quick answer: Decision Index (version 0.2.1) is a community-maintained suite that scores "decision models" — models that read a state plus a schema of typed questions and return a calibrated probability for every option in a single forward pass. Results are a chance-corrected aggregate over 38 sub-benchmarks (0 = random guessing, 100 = perfect). Cloudflare's Clef leads at 61.21, self-reported.

At a Glance

What it tests: How accurately and how well-calibrated a model is when it must make bounded, structured decisions (choice, score, true/false questions) rather than write free-form text — across tool selection, intent classification, NLI, reasoning, retrieval and safety sub-benchmarks.

Why it matters: Decision models are a new class of small, fast classifiers designed to sit inside agent workflows. Decision Index gives them a shared yardstick that also reports latency, so accuracy can be weighed against speed.

Known limitations: The suite is community-run and unofficial. Clef and Clef-Flash were run by their authors (Cloudflare) on 36 of 38 benchmarks, with the missing ones scored as 0, and have not been reproduced by the upstream board.

What Decision Index Measures

Decision models, popularised by TypeSafe AI's Jev, return typed answers — a named choice, an ordinal score, or a true/false probability — for every question asked about a given state. Decision Index aggregates dozens of existing public benchmarks (BFCL, ToolRet, API-Bank, BANKING77, CLINC150, MMLU, GPQA Diamond, GSM8K, MuSR, ForecastBench and others) adapted to this answer format.

Scores are chance-corrected so that a model guessing at random scores 0 and a perfect model scores 100, which keeps easy and hard sub-benchmarks comparable. Per-benchmark breakdowns and median/p95 request latency are reported alongside the aggregate.

Benchmark Specifications

FieldValue
Task categoryDecision / classification
MetricChance-corrected aggregate score (0-100)
Version0.2.1
Sub-benchmarks38
Created bymultimodalart and community contributors
ProjectHugging Face Space

How Decision Index Is Scored

Each sub-benchmark's metric (accuracy, macro-F1, nDCG@10, Brier score) is converted to a chance-corrected score, and the aggregate averages across categories. Latency (median and p95 ms per request) is reported separately and does not enter the score.

State-of-the-Art Results

RankModelDecision IndexSourceDate
1Clef61.21 (self-reported)Cloudflare decision leaderboard2026-10
2Jev (TypeSafe AI, closed API)57.91Cloudflare decision leaderboard2026-10
3Surogate Rune 26B-A4B v357.44Cloudflare decision leaderboard2026-10
4Decider chat (Gemma-4-31B)57.33Cloudflare decision leaderboard2026-10
5Clef-Flash57.07 (self-reported)Cloudflare decision leaderboard2026-10

Community results snapshot 28 September 2026, mirrored on Cloudflare's leaderboard (built 1 October 2026). Only models with a Benchgen page are linked; the rest are shown for context.

Decision Index on Benchgen

No Benchgen results yet — be the first to run Decision Index.

Decision Index vs Other Benchmarks

BenchmarkWhat it testsOutput format
Decision IndexCalibrated structured decisions across 38 sub-benchmarksTyped probabilities
JevBenchDecision accuracy and calibration on a fixed item setTyped probabilities
BFCL v4Function-calling accuracyGenerated tool calls

Run Decision Index on Your Model

Benchgen can track decision-model performance over time alongside the community-reported scores above.

Frequently Asked Questions

What is the Decision Index? A community suite that scores decision models — models that turn a state and typed questions into calibrated probabilities — as a chance-corrected aggregate over 38 sub-benchmarks.
What does a good Decision Index score look like? Clef leads at 61.21 (self-reported) and TypeSafe AI's closed Jev scores 57.91; most small open decision models score between roughly 25 and 55.
Who created the Decision Index? It is maintained by multimodalart and community contributors on Hugging Face, and is unofficial and not affiliated with TypeSafe AI.