Benchgen

Claw-Eval — Results

RankModelScore
1lfm2-5-2-6b62.85
C

Claw-Eval

1 phaseActive

300-task cross-lab reference benchmark using Pass^3 methodology to reduce single-run noise in agent evaluation.

Overview

Claw-Eval

Category Metric Tasks Saturation Created

Quick answer: Claw-Eval is a 300-task, 9-category agent benchmark that has become a cross-lab reference point — it's referenced or used internally by Meta, Kimi, Qwen, and Tencent as a "trustworthy" evaluation baseline. It uses Pass^3 scoring (each task is attempted 3 independent times and judged on consistency across trials) specifically to reduce the noise that single-run pass/fail evaluation introduces into agent benchmarking.

At a Glance

What it tests: General agent task-completion ability across 300 human-verified tasks spanning 9 categories, scored with a 3-trial consistency methodology rather than a single pass/fail attempt.

Why it matters: Single-run agent evaluation is notoriously noisy — the same agent can pass or fail an identical task on different attempts due to sampling variance. Claw-Eval's Pass^3 methodology directly addresses this, which is likely why it has been adopted as a reference/sanity-check benchmark across multiple frontier labs.

Known limitations: No single originating paper with full authorship has been consistently cited alongside Claw-Eval; it functions more as a shared cross-lab reference set than a benchmark with one canonical public leaderboard.

What Claw-Eval Measures

Claw-Eval evaluates general agent capability across 300 human-verified tasks organized into 9 categories, designed to be broad enough to serve as a cross-lab sanity check rather than a narrow, single-domain test. Its defining methodological choice is Pass^3 scoring: rather than running each task once and recording pass/fail, Claw-Eval runs each task 3 independent times and incorporates trial-to-trial consistency into the score. This directly targets a known weakness of single-run agent benchmarks — an agent that passes a task once by chance is scored differently from one that reliably passes it every time.

Because of this trust-oriented design, Claw-Eval is cited as a reference benchmark by multiple frontier labs (Meta, Kimi, Qwen, and Tencent), typically as one data point among several rather than the sole basis for capability claims — a role similar to how AlpacaEval or MT-Bench function as widely trusted secondary checks alongside a lab's primary benchmark suite.

Benchmark Specifications

FieldValue
Task categoryAgent (general capability, cross-lab reference)
MetricPass^3 (pass rate across 3 independent trials)
Number of tasks300, human-verified
Categories9
Adopted/cited byMeta, Kimi, Qwen, Tencent
SaturationLow
Created byCross-lab agent evaluation reference
PaperNot consistently attributed to a single public paper
DatasetNot public

How Claw-Eval Is Scored

Each of the 300 tasks is attempted 3 independent times per agent under test. The Pass^3 metric credits an agent based on how consistently it succeeds across those 3 trials, rather than treating a single successful attempt as full credit. This makes Claw-Eval more resistant to sampling-variance noise than one-shot pass/fail benchmarks, at the cost of requiring 3× the inference compute per task to produce a score.

Claw-Eval on Benchgen

No Benchgen results yet — be the first to run Claw-Eval.

Claw-Eval vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
Claw-EvalGeneral agent capability, 3-trial consistency scoring300Low
Harness-BenchHarness quality, holding model & task fixed106Low
MT-BenchMulti-turn conversational qualityHigh
Arena-Hard v2Preference-based model comparisonMedium

Claw-Eval plays a similar cross-lab "trust check" role to MT-Bench and Arena-Hard, but applied to agent task completion rather than conversational quality, and with a multi-trial (Pass^3) methodology specifically built to reduce single-run noise.

Run Claw-Eval on Your Model

Benchgen lets teams run multi-trial (Pass^k) agent evaluations against their own models and harnesses, applying the same consistency-scoring approach Claw-Eval uses to distinguish reliable task completion from lucky single-run passes.

Frequently Asked Questions

What is Claw-Eval? Claw-Eval is a 300-task, 9-category agent benchmark that uses Pass^3 (3-trial) scoring to reduce single-run noise, and is referenced as a cross-lab evaluation baseline by Meta, Kimi, Qwen, and Tencent.
What does Pass^3 scoring mean? Each task is attempted 3 independent times, and the agent's score reflects how consistently it succeeds across all 3 trials rather than crediting a single successful attempt — reducing the effect of sampling variance on the final score.
Who created Claw-Eval? Claw-Eval does not have one consistently cited originating paper; it functions as a shared cross-lab reference benchmark used internally by multiple frontier labs.
Is Claw-Eval saturated? Public score data is limited, but its 9-category, 300-task, 3-trial design is intended to remain a meaningful differentiator across agent capability levels rather than a fully solved benchmark.