| Rank | Model | Score |
|---|---|---|
| 1 | deepseek-v4-1-flash | 54.8 |
| 2 | nex-n2-5-max | 50.2 |
| 3 | nex-n2-5-pro | 44.2 |
| 4 | gpt-6-astra | 41.4 |
| 5 | deepseek-v4-pro | 31.8 |
| 6 | claude-fable-5-1 | 31.4 |
| 7 | kimi-k3 | 30.8 |
| 8 | qwen3-8-max | 27.3 |
| 9 | deepseek-v4-flash-vision-exp | 25.7 |
| 10 | k2-horizon-375b-a23b | 25.3 |
1 phaseActive
Business workflow automation benchmark testing AI agents on multi-step, cross-system process automation. Metric: % task success.
Quick answer: AutomationBench tests AI agents on real business workflow automation tasks — multi-step processes that span multiple applications or systems, similar to what a robotic process automation (RPA) tool or business ops team would handle. Kimi K3 scores 30.8% as of July 2026, reflecting the difficulty of cross-system business automation.
What it tests: An agent's ability to automate multi-step business processes across different applications and systems, completing them reliably end-to-end.
Why it matters: Business process automation is a major commercial use case for AI agents. AutomationBench's low current scores (30.8% for a frontier model like Kimi K3) highlight how much headroom remains for reliable enterprise automation.
Known limitations: As an emerging benchmark, exact task composition and business-process templates are not yet independently published outside its citation by Moonshot AI.
AutomationBench evaluates an agent's ability to automate realistic business workflows that span multiple systems or applications — for example, extracting data from one tool, transforming it, and populating another. Tasks emphasize reliability and correctness across the full multi-step process rather than any single isolated step, making the benchmark a useful proxy for real enterprise-automation readiness.
| Field | Value |
|---|---|
| Task category | Agent / business process automation |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents attempt multi-step business automation workflows, with outcomes verified against expected end states across the involved systems, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 30.8% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run AutomationBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| AutomationBench | Business workflow automation | Low |
| SaaS-Bench | SaaS product usage workflows | Low |
| JobBench | Job/workforce-oriented agent tasks | Low |
| Tau3 Banking | Customer-service agentic tasks | Low |
Benchgen lets you run AutomationBench against your own agent, tracking cross-system business automation task success rates before production rollout.