Quick answer: Clef is Cloudflare's open decision model, released October 1, 2026. It is a 27B multimodal model post-trained from Qwen3.8-27B that turns a state (text, JSON, images or video) and a schema of typed questions into a calibrated probability for every option in one forward pass, with no free-form text generation. It scores 61.21 on the Decision Index (self-reported), ahead of TypeSafe AI's closed Jev (57.91), with about 209 ms median latency. Apache 2.0, available on Hugging Face and Workers AI.
What it is: A decision model — a classifier-style model that returns typed answers (choice, score, noul true/false) with probabilities, designed to be called in the hot path of an agent to route, triage or escalate.
Where it leads: Top of the Decision Index at 61.21, with strong results on intent classification (BANKING77 94.2, CLINC150+OOS 97.4 macro-F1), tool retrieval (ToolRet 69.2 nDCG@10) and BFCL (98.5 case-exact accuracy).
Where it lags: Weaker than Jev on GPQA Diamond-style reasoning in decision format (48.0 vs 78.3), MMLU-Pro (65.9 vs 82.7) and BBH (73.7 vs 92.9); results are self-reported and the missing 2 of 38 benchmarks were scored as 0.
Clef freezes the Qwen3.8-27B backbone and jointly trains a joint schema head — a small transformer that routes evidence from the state to each question and scores every option of every question together — alongside rank-256 low-rank adapters. Training used label-smoothed cross-entropy with a Brier loss for calibration on internal synthetic data, plus a reinforcement-learning stage called RLCD (Reinforcement Learning for Calibrated Decisions) that gives partial credit to adjacent ordinal choices.
It is fully API-compatible with TypeSafe AI's Jev/SystemOne /v1/systemone format, accepts images and video, and supports a 64K-token context. Cloudflare reports that on its Threat Intelligence workflow Clef classified a website in 2.2 s versus 4.7 s for gpt-oss-120b. Smaller sibling: Clef-Flash.
| Field | Value |
|---|---|
| Organization | Cloudflare |
| Hugging Face | Cloudflare/clef |
| Base model | Qwen/Qwen3.8-27B |
| Parameters | 27B |
| Output | Typed probabilities (no text generation) |
| Context window | 64K tokens |
| License | Apache 2.0 |
| Release date | 2026-10-01 |
| Availability | Hugging Face weights; Cloudflare Workers AI (@cf/cloudflare/clef) |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Decision Index (v0.2.1, chance-corrected) | 61.21 | Cloudflare decision leaderboard | 2026-10 |
Self-reported by Cloudflare on 36 of 38 Decision Index benchmarks (missing benchmarks scored as 0) and not yet reproduced by the upstream board. Cloudflare's model card also lists per-benchmark results (for example BFCL 98.5, ToolRet 69.2, BANKING77 94.2, CLINC150+OOS 97.4, MMLU 90.3, GPQA Diamond 48.0, GSM8K 80.8, median latency 209.3 ms); these use a decision (option-probability) format and are not comparable with the generative scores on Benchgen's MMLU, GPQA or BFCL pages, so they are not added.
/v1/systemone request and response format.