| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 30.8 |
| 2 | qwen3-8-max | 27.3 |
1 phaseActive
Business workflow automation benchmark testing AI agents on multi-step, cross-system process automation. Metric: % task success.
Quick answer: AutomationBench tests AI agents on real business workflow automation tasks — multi-step processes that span multiple applications or systems, similar to what a robotic process automation (RPA) tool or business ops team would handle. Kimi K3 scores 30.8% as of July 2026, reflecting the difficulty of cross-system business automation.
What it tests: An agent's ability to automate multi-step business processes across different applications and systems, completing them reliably end-to-end.
Why it matters: Business process automation is a major commercial use case for AI agents. AutomationBench's low current scores (30.8% for a frontier model like Kimi K3) highlight how much headroom remains for reliable enterprise automation.
Known limitations: As an emerging benchmark, exact task composition and business-process templates are not yet independently published outside its citation by Moonshot AI.
AutomationBench evaluates an agent's ability to automate realistic business workflows that span multiple systems or applications — for example, extracting data from one tool, transforming it, and populating another. Tasks emphasize reliability and correctness across the full multi-step process rather than any single isolated step, making the benchmark a useful proxy for real enterprise-automation readiness.
| Field | Value |
|---|---|
| Task category | Agent / business process automation |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents attempt multi-step business automation workflows, with outcomes verified against expected end states across the involved systems, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 30.8% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run AutomationBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| AutomationBench | Business workflow automation | Low |
| SaaS-Bench | SaaS product usage workflows | Low |
| JobBench | Job/workforce-oriented agent tasks | Low |
| Tau3 Banking | Customer-service agentic tasks | Low |
Benchgen lets you run AutomationBench against your own agent, tracking cross-system business automation task success rates before production rollout.