Benchgen

JobBench — Results

RankModelScore
1kimi-k354.3
2qwen3-8-max53.4
J

JobBench

1 phaseActive

Workforce-oriented benchmark testing AI agents on tasks drawn from real job roles across occupations. Metric: % task success.

Overview

JobBench

Category Metric Saturation Created

Quick answer: JobBench evaluates AI agents on tasks drawn directly from real job roles across a range of occupations, testing whether a model can perform the actual day-to-day work associated with specific professions. Kimi K3 scores 54.3% as of July 2026.

At a Glance

What it tests: An agent's ability to complete authentic job-role tasks — the kind of work assigned to professionals in specific occupations — rather than abstract or synthetic test items.

Why it matters: As AI agents are increasingly evaluated for workforce automation potential, JobBench offers a task-grounded way to measure real occupational capability rather than proxy metrics.

Known limitations: As an emerging benchmark, the specific occupations and task sourcing methodology are not yet independently published outside its citation by Moonshot AI.

What JobBench Measures

JobBench evaluates an agent's ability to perform tasks sourced from real job roles across a range of occupations, testing occupational competence directly rather than through proxy academic tasks. This complements broader economically-valuable-task benchmarks (like Agents' Last Exam) by focusing specifically on job-role-grounded task completion.

Benchmark Specifications

FieldValue
Task categoryAgent / workforce automation
Metric% task success
SaturationLow
Created byNot yet independently documented

How JobBench Is Scored

Agents attempt tasks drawn from real job roles, with outcomes verified against expected deliverables, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K354.3%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

JobBench on Benchgen

No Benchgen results yet — be the first to run JobBench.

JobBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
JobBenchReal job-role task completionLow
Agents' Last ExamEconomically valuable long-horizon agent tasksLow
APEX-AgentsProfessional-grade knowledge-work agent tasksLow
AutomationBenchBusiness workflow automationLow

Run JobBench on Your Model

Benchgen lets you run JobBench against your own agent, tracking real job-role task success rates over time to assess workforce-automation readiness.

Frequently Asked Questions

What is JobBench? JobBench is a benchmark testing AI agents on tasks drawn from real job roles across a range of occupations.
What does a good score look like on JobBench? Kimi K3 reports 54.3% as of July 2026, a solid frontier-level result reflecting genuine occupational task competence.
Who created JobBench? JobBench's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.