| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 54.3 |
| 2 | qwen3-8-max | 53.4 |
1 phaseActive
Workforce-oriented benchmark testing AI agents on tasks drawn from real job roles across occupations. Metric: % task success.
Quick answer: JobBench evaluates AI agents on tasks drawn directly from real job roles across a range of occupations, testing whether a model can perform the actual day-to-day work associated with specific professions. Kimi K3 scores 54.3% as of July 2026.
What it tests: An agent's ability to complete authentic job-role tasks — the kind of work assigned to professionals in specific occupations — rather than abstract or synthetic test items.
Why it matters: As AI agents are increasingly evaluated for workforce automation potential, JobBench offers a task-grounded way to measure real occupational capability rather than proxy metrics.
Known limitations: As an emerging benchmark, the specific occupations and task sourcing methodology are not yet independently published outside its citation by Moonshot AI.
JobBench evaluates an agent's ability to perform tasks sourced from real job roles across a range of occupations, testing occupational competence directly rather than through proxy academic tasks. This complements broader economically-valuable-task benchmarks (like Agents' Last Exam) by focusing specifically on job-role-grounded task completion.
| Field | Value |
|---|---|
| Task category | Agent / workforce automation |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents attempt tasks drawn from real job roles, with outcomes verified against expected deliverables, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 54.3% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run JobBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| JobBench | Real job-role task completion | Low |
| Agents' Last Exam | Economically valuable long-horizon agent tasks | Low |
| APEX-Agents | Professional-grade knowledge-work agent tasks | Low |
| AutomationBench | Business workflow automation | Low |
Benchgen lets you run JobBench against your own agent, tracking real job-role task success rates over time to assess workforce-automation readiness.