Benchgen

AutomationBench — Results

RankModelScore
1kimi-k330.8
2qwen3-8-max27.3
A

AutomationBench

1 phaseActive

Business workflow automation benchmark testing AI agents on multi-step, cross-system process automation. Metric: % task success.

Overview

AutomationBench

Category Metric Saturation Created

Quick answer: AutomationBench tests AI agents on real business workflow automation tasks — multi-step processes that span multiple applications or systems, similar to what a robotic process automation (RPA) tool or business ops team would handle. Kimi K3 scores 30.8% as of July 2026, reflecting the difficulty of cross-system business automation.

At a Glance

What it tests: An agent's ability to automate multi-step business processes across different applications and systems, completing them reliably end-to-end.

Why it matters: Business process automation is a major commercial use case for AI agents. AutomationBench's low current scores (30.8% for a frontier model like Kimi K3) highlight how much headroom remains for reliable enterprise automation.

Known limitations: As an emerging benchmark, exact task composition and business-process templates are not yet independently published outside its citation by Moonshot AI.

What AutomationBench Measures

AutomationBench evaluates an agent's ability to automate realistic business workflows that span multiple systems or applications — for example, extracting data from one tool, transforming it, and populating another. Tasks emphasize reliability and correctness across the full multi-step process rather than any single isolated step, making the benchmark a useful proxy for real enterprise-automation readiness.

Benchmark Specifications

FieldValue
Task categoryAgent / business process automation
Metric% task success
SaturationLow
Created byNot yet independently documented

How AutomationBench Is Scored

Agents attempt multi-step business automation workflows, with outcomes verified against expected end states across the involved systems, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K330.8%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

AutomationBench on Benchgen

No Benchgen results yet — be the first to run AutomationBench.

AutomationBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
AutomationBenchBusiness workflow automationLow
SaaS-BenchSaaS product usage workflowsLow
JobBenchJob/workforce-oriented agent tasksLow
Tau3 BankingCustomer-service agentic tasksLow

Run AutomationBench on Your Model

Benchgen lets you run AutomationBench against your own agent, tracking cross-system business automation task success rates before production rollout.

Frequently Asked Questions

What is AutomationBench? AutomationBench is a benchmark testing AI agents on real business workflow automation tasks spanning multiple applications and systems.
What does a good score look like on AutomationBench? Kimi K3 reports 30.8% as of July 2026. Given the cross-system reliability required, this reflects a frontier-level result on a still-difficult, unsaturated benchmark.
Who created AutomationBench? AutomationBench's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.