| Rank | Model | Score |
|---|---|---|
| 1 | k2-type-0-9b | 76.2 |
1 phaseActive
Benchmark for Jev-style decision models: accuracy and calibration (Brier, ECE) on a public set of 231 decision items plus a sealed hold-out.
Quick answer: JevBench measures how accurately and how well-calibrated a decision model is — a model that returns typed probabilities (choice, score, true/false) instead of free text. The public set has 231 items split into standard (72), easy (48) and hard (111); the official score additionally uses sealed items. K2-Type-0.9B reaches 76.2% on the public set.
What it tests: Decision accuracy, Brier score and expected calibration error (ECE) for models speaking the Jev/SystemOne /v1/systemone API.
Why it matters: It was the first shared yardstick for the wave of small decision models released after TypeSafe AI's Jev, and reports calibration alongside accuracy.
Known limitations: Public-set accuracy is higher than the official score because the official score also uses sealed items, on which every listed system scores well below its public accuracy. Scores here are self-reported by model authors.
Each item presents a state and one or more typed questions; the model returns a probability for every option. Accuracy is the share of items where the highest-probability option matches the label. Calibration is reported as Brier score and ECE, and latency per decision is tracked in some reports.
Reported public-set accuracy for other systems on the JevBench v1.4 board includes Jev 1.13.0 (0.866), Qwen3.5-4B-based entries (0.74-0.82), Gemma 4 E2B + LoRA (0.732), decider-2b (0.710), Kev 0.6B (0.667) and Kev 4B (0.662). Only models with a Benchgen page are listed in the table below.
| Field | Value |
|---|---|
| Task category | Decision / classification |
| Metric | % accuracy (public set) |
| Public items | 231 (standard 72, easy 48, hard 111) |
| Created by | fstandhartinger |
| Project | github.com/fstandhartinger/jevbench |
Models are run through the typesafe adapter against their own server on a single GPU, and accuracy over the 231 public items is reported with Brier score and ECE. The official leaderboard score also includes sealed items.
| Rank | Model | Public-set accuracy | Source | Date |
|---|---|---|---|---|
| 1 | K2-Type-0.9B | 76.2% (176/231; Brier 0.328, ECE 0.065) | IFM model card | 2026-09 |
Self-reported by the model's authors on the public set; not independently reproduced.
No Benchgen results yet — be the first to run JevBench.
| Benchmark | What it tests | Output format |
|---|---|---|
| JevBench | Decision accuracy + calibration, fixed item set | Typed probabilities |
| Decision Index | Aggregate over 38 sub-benchmarks | Typed probabilities |
Run the public set with the JevBench repository's typesafe adapter against your decision model's server.