| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 60.1 |
1 phaseActive
SaaS product usage benchmark testing AI agents on realistic workflows within common business web applications. Metric: % task success.
Quick answer: SaaS-Bench tests AI agents on realistic workflows within common business SaaS web applications — CRMs, project-management tools, and similar products — evaluating whether an agent can navigate and operate these tools as a real user would. Kimi K3 scores 60.1% as of July 2026.
What it tests: An agent's ability to navigate and complete tasks within real SaaS web applications, mirroring how a human employee would use business software.
Why it matters: SaaS automation is one of the largest near-term commercial opportunities for AI agents. SaaS-Bench measures readiness for this specific, high-value deployment category.
Known limitations: As an emerging benchmark, the specific set of SaaS applications tested and task-sourcing methodology are not yet independently published outside its citation by Moonshot AI.
SaaS-Bench evaluates an agent's ability to complete realistic tasks within common business SaaS applications — such as configuring settings, entering and retrieving data, and completing multi-step workflows across app UIs. This complements broader GUI-agent benchmarks (like OSWorld) by focusing specifically on business-software usage patterns.
| Field | Value |
|---|---|
| Task category | Agent / SaaS product usage |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents attempt tasks within SaaS applications, with outcomes verified against the expected application state, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 60.1% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run SaaS-Bench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| SaaS-Bench | SaaS product usage workflows | Low |
| OSWorld 2.0 | Expanded, harder GUI agent workflows | Low |
| AutomationBench | Business workflow automation | Low |
| MCP Atlas | MCP tool-use agentic workflows | Low |
Benchgen lets you run SaaS-Bench against your own agent, tracking SaaS product usage task success rates before rolling out agentic automation to business software.