Benchgen

JevBench — Results

RankModelScore
1k2-type-0-9b76.2
J

JevBench

1 phaseActive

Benchmark for Jev-style decision models: accuracy and calibration (Brier, ECE) on a public set of 231 decision items plus a sealed hold-out.

Overview

JevBench

Category Metric Tasks Saturation

Quick answer: JevBench measures how accurately and how well-calibrated a decision model is — a model that returns typed probabilities (choice, score, true/false) instead of free text. The public set has 231 items split into standard (72), easy (48) and hard (111); the official score additionally uses sealed items. K2-Type-0.9B reaches 76.2% on the public set.

At a Glance

What it tests: Decision accuracy, Brier score and expected calibration error (ECE) for models speaking the Jev/SystemOne /v1/systemone API.

Why it matters: It was the first shared yardstick for the wave of small decision models released after TypeSafe AI's Jev, and reports calibration alongside accuracy.

Known limitations: Public-set accuracy is higher than the official score because the official score also uses sealed items, on which every listed system scores well below its public accuracy. Scores here are self-reported by model authors.

What JevBench Measures

Each item presents a state and one or more typed questions; the model returns a probability for every option. Accuracy is the share of items where the highest-probability option matches the label. Calibration is reported as Brier score and ECE, and latency per decision is tracked in some reports.

Reported public-set accuracy for other systems on the JevBench v1.4 board includes Jev 1.13.0 (0.866), Qwen3.5-4B-based entries (0.74-0.82), Gemma 4 E2B + LoRA (0.732), decider-2b (0.710), Kev 0.6B (0.667) and Kev 4B (0.662). Only models with a Benchgen page are listed in the table below.

Benchmark Specifications

FieldValue
Task categoryDecision / classification
Metric% accuracy (public set)
Public items231 (standard 72, easy 48, hard 111)
Created byfstandhartinger
Projectgithub.com/fstandhartinger/jevbench

How JevBench Is Scored

Models are run through the typesafe adapter against their own server on a single GPU, and accuracy over the 231 public items is reported with Brier score and ECE. The official leaderboard score also includes sealed items.

State-of-the-Art Results

RankModelPublic-set accuracySourceDate
1K2-Type-0.9B76.2% (176/231; Brier 0.328, ECE 0.065)IFM model card2026-09

Self-reported by the model's authors on the public set; not independently reproduced.

JevBench on Benchgen

No Benchgen results yet — be the first to run JevBench.

JevBench vs Other Benchmarks

BenchmarkWhat it testsOutput format
JevBenchDecision accuracy + calibration, fixed item setTyped probabilities
Decision IndexAggregate over 38 sub-benchmarksTyped probabilities

Run JevBench on Your Model

Run the public set with the JevBench repository's typesafe adapter against your decision model's server.

Frequently Asked Questions

What is JevBench? A benchmark of decision items for models that return typed probabilities, reporting accuracy plus calibration (Brier score and ECE).
What is a good JevBench score? On the public set, Jev 1.13.0 reports 86.6% and small open models typically fall between about 66% and 82%; K2-Type-0.9B scores 76.2%.
Why is the official score lower than public accuracy? The official score also uses sealed items, and every listed system scores well below its public-set accuracy on them.