Benchgen

SaaS-Bench — Results

RankModelScore
1kimi-k360.1
S

SaaS-Bench

1 phaseActive

SaaS product usage benchmark testing AI agents on realistic workflows within common business web applications. Metric: % task success.

Overview

SaaS-Bench

Category Metric Saturation Created

Quick answer: SaaS-Bench tests AI agents on realistic workflows within common business SaaS web applications — CRMs, project-management tools, and similar products — evaluating whether an agent can navigate and operate these tools as a real user would. Kimi K3 scores 60.1% as of July 2026.

At a Glance

What it tests: An agent's ability to navigate and complete tasks within real SaaS web applications, mirroring how a human employee would use business software.

Why it matters: SaaS automation is one of the largest near-term commercial opportunities for AI agents. SaaS-Bench measures readiness for this specific, high-value deployment category.

Known limitations: As an emerging benchmark, the specific set of SaaS applications tested and task-sourcing methodology are not yet independently published outside its citation by Moonshot AI.

What SaaS-Bench Measures

SaaS-Bench evaluates an agent's ability to complete realistic tasks within common business SaaS applications — such as configuring settings, entering and retrieving data, and completing multi-step workflows across app UIs. This complements broader GUI-agent benchmarks (like OSWorld) by focusing specifically on business-software usage patterns.

Benchmark Specifications

FieldValue
Task categoryAgent / SaaS product usage
Metric% task success
SaturationLow
Created byNot yet independently documented

How SaaS-Bench Is Scored

Agents attempt tasks within SaaS applications, with outcomes verified against the expected application state, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K360.1%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

SaaS-Bench on Benchgen

No Benchgen results yet — be the first to run SaaS-Bench.

SaaS-Bench vs Other Benchmarks

BenchmarkWhat it testsSaturation
SaaS-BenchSaaS product usage workflowsLow
OSWorld 2.0Expanded, harder GUI agent workflowsLow
AutomationBenchBusiness workflow automationLow
MCP AtlasMCP tool-use agentic workflowsLow

Run SaaS-Bench on Your Model

Benchgen lets you run SaaS-Bench against your own agent, tracking SaaS product usage task success rates before rolling out agentic automation to business software.

Frequently Asked Questions

What is SaaS-Bench? SaaS-Bench is a benchmark testing AI agents on realistic workflows within common business SaaS web applications, such as CRMs and project-management tools.
What does a good score look like on SaaS-Bench? Kimi K3 reports 60.1% as of July 2026, a solid frontier-level result reflecting reasonable reliability navigating real business software.
Who created SaaS-Bench? SaaS-Bench's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.