Benchgen
Models/cloudflare/

Clef

DraftPublic

Model Details

Clef

Organization License Params Type Released

Quick answer: Clef is Cloudflare's open decision model, released October 1, 2026. It is a 27B multimodal model post-trained from Qwen3.8-27B that turns a state (text, JSON, images or video) and a schema of typed questions into a calibrated probability for every option in one forward pass, with no free-form text generation. It scores 61.21 on the Decision Index (self-reported), ahead of TypeSafe AI's closed Jev (57.91), with about 209 ms median latency. Apache 2.0, available on Hugging Face and Workers AI.

At a Glance

What it is: A decision model — a classifier-style model that returns typed answers (choice, score, noul true/false) with probabilities, designed to be called in the hot path of an agent to route, triage or escalate.

Where it leads: Top of the Decision Index at 61.21, with strong results on intent classification (BANKING77 94.2, CLINC150+OOS 97.4 macro-F1), tool retrieval (ToolRet 69.2 nDCG@10) and BFCL (98.5 case-exact accuracy).

Where it lags: Weaker than Jev on GPQA Diamond-style reasoning in decision format (48.0 vs 78.3), MMLU-Pro (65.9 vs 82.7) and BBH (73.7 vs 92.9); results are self-reported and the missing 2 of 38 benchmarks were scored as 0.

What Clef Is

Clef freezes the Qwen3.8-27B backbone and jointly trains a joint schema head — a small transformer that routes evidence from the state to each question and scores every option of every question together — alongside rank-256 low-rank adapters. Training used label-smoothed cross-entropy with a Brier loss for calibration on internal synthetic data, plus a reinforcement-learning stage called RLCD (Reinforcement Learning for Calibrated Decisions) that gives partial credit to adjacent ordinal choices.

It is fully API-compatible with TypeSafe AI's Jev/SystemOne /v1/systemone format, accepts images and video, and supports a 64K-token context. Cloudflare reports that on its Threat Intelligence workflow Clef classified a website in 2.2 s versus 4.7 s for gpt-oss-120b. Smaller sibling: Clef-Flash.

Specifications

FieldValue
OrganizationCloudflare
Hugging FaceCloudflare/clef
Base modelQwen/Qwen3.8-27B
Parameters27B
OutputTyped probabilities (no text generation)
Context window64K tokens
LicenseApache 2.0
Release date2026-10-01
AvailabilityHugging Face weights; Cloudflare Workers AI (@cf/cloudflare/clef)

Public Benchmark Scores

BenchmarkScoreSourceDate
Decision Index (v0.2.1, chance-corrected)61.21Cloudflare decision leaderboard2026-10

Self-reported by Cloudflare on 36 of 38 Decision Index benchmarks (missing benchmarks scored as 0) and not yet reproduced by the upstream board. Cloudflare's model card also lists per-benchmark results (for example BFCL 98.5, ToolRet 69.2, BANKING77 94.2, CLINC150+OOS 97.4, MMLU 90.3, GPQA Diamond 48.0, GSM8K 80.8, median latency 209.3 ms); these use a decision (option-probability) format and are not comparable with the generative scores on Benchgen's MMLU, GPQA or BFCL pages, so they are not added.

Frequently Asked Questions

What is Clef? Cloudflare's open 27B multimodal decision model that outputs calibrated probabilities for typed questions instead of generating text.
How is Clef different from an LLM? It scores every allowed option in a single forward pass, so there is no token-by-token generation and no output parsing, which makes it faster and deterministic for routing and classification.
Is Clef compatible with Jev? Yes, it follows the same Jev/SystemOne /v1/systemone request and response format.