| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 76.5 |
| 2 | qwen3-8-max | 72.5 |
1 phaseActive
Verified multi-tool agentic benchmark testing sustained tool-use workflows across diverse tool sets. Metric: % task success.
Quick answer: Toolathlon-Verified is a verified benchmark testing AI agents on multi-tool, multi-step workflows that require chaining together diverse tools to complete a task — an "athlon" of sequential tool-use challenges. Kimi K3 scores 76.5% as of July 2026.
What it tests: An agent's ability to select, sequence, and correctly invoke multiple different tools across a multi-step workflow to reach a verified final outcome.
Why it matters: Real-world agent deployments often require juggling many tools (search, code execution, file systems, APIs). Toolathlon-Verified specifically stresses this multi-tool coordination capability.
Known limitations: As an emerging benchmark, exact tool inventory and task composition are not yet independently published outside its citation by Moonshot AI.
Toolathlon-Verified evaluates an agent's ability to coordinate multiple distinct tools across a sequential, multi-step workflow, with task outcomes verified programmatically. The "Verified" designation indicates that task setups and success criteria have been validated/corrected, similar to the "Verified" convention used in benchmarks like SWE-Bench Verified and OSWorld-Verified.
| Field | Value |
|---|---|
| Task category | Agent / multi-tool workflows |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents complete multi-tool workflows, with final outcomes verified against expected results, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 76.5% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run Toolathlon-Verified.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Toolathlon-Verified | Multi-tool, multi-step agentic workflows | Low |
| MCP Atlas | MCP tool-use agentic workflows | Low |
| MCPMark-Verified | MCP protocol tool-use benchmark | Low |
| BFCL v3 | Function-calling accuracy | Medium |
Benchgen lets you run Toolathlon-Verified against your own agent harness, tracking multi-tool task success rates as you expand your agent's tool inventory.