| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 94.5 |
1 phaseActive
Verified benchmark for Model Context Protocol (MCP) tool-use agentic tasks, testing MCP server/tool interaction correctness. Metric: % task success.
Quick answer: MCPMark-Verified is a benchmark testing AI agents on tasks that require correctly interacting with Model Context Protocol (MCP) servers and tools — connecting to external systems, invoking MCP tool calls, and completing verified end-to-end workflows. Kimi K3 scores 94.5% as of July 2026.
What it tests: An agent's ability to correctly discover, connect to, and invoke tools exposed via the Model Context Protocol (MCP) to complete real tasks.
Why it matters: MCP has become a standard way for AI agents to connect to external tools and data sources. MCPMark-Verified measures how reliably a model can operate within this now-widespread protocol.
Known limitations: As an emerging benchmark, the exact MCP server/tool inventory and task composition are not yet independently published outside its citation by Moonshot AI.
MCPMark-Verified evaluates an agent's ability to interact correctly with tools and data sources exposed through the Model Context Protocol (MCP) — an increasingly standard interface for connecting LLM agents to external systems. Tasks require correct tool discovery, parameter construction, and multi-step MCP tool invocation, with outcomes verified against expected results.
| Field | Value |
|---|---|
| Task category | Agent / MCP tool-use |
| Metric | % task success |
| Saturation | Low |
| Created by | Not yet independently documented |
Agents complete tasks requiring MCP tool interaction, with outcomes verified against expected results, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 94.5% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run MCPMark-Verified.
| Benchmark | What it tests | Saturation |
|---|---|---|
| MCPMark-Verified | MCP protocol tool-use correctness | Low |
| MCP Atlas | MCP tool-use agentic workflows | Low |
| Toolathlon-Verified | Multi-tool, multi-step agentic workflows | Low |
| BFCL v3 | Function-calling accuracy | Medium |
Benchgen lets you run MCPMark-Verified against your own agent's MCP integration, tracking tool-use task success rates over time.