Benchgen

MCPMark-Verified — Results

RankModelScore
1kimi-k394.5
M

MCPMark-Verified

1 phaseActive

Verified benchmark for Model Context Protocol (MCP) tool-use agentic tasks, testing MCP server/tool interaction correctness. Metric: % task success.

Overview

MCPMark-Verified

Category Metric Saturation Created

Quick answer: MCPMark-Verified is a benchmark testing AI agents on tasks that require correctly interacting with Model Context Protocol (MCP) servers and tools — connecting to external systems, invoking MCP tool calls, and completing verified end-to-end workflows. Kimi K3 scores 94.5% as of July 2026.

At a Glance

What it tests: An agent's ability to correctly discover, connect to, and invoke tools exposed via the Model Context Protocol (MCP) to complete real tasks.

Why it matters: MCP has become a standard way for AI agents to connect to external tools and data sources. MCPMark-Verified measures how reliably a model can operate within this now-widespread protocol.

Known limitations: As an emerging benchmark, the exact MCP server/tool inventory and task composition are not yet independently published outside its citation by Moonshot AI.

What MCPMark-Verified Measures

MCPMark-Verified evaluates an agent's ability to interact correctly with tools and data sources exposed through the Model Context Protocol (MCP) — an increasingly standard interface for connecting LLM agents to external systems. Tasks require correct tool discovery, parameter construction, and multi-step MCP tool invocation, with outcomes verified against expected results.

Benchmark Specifications

FieldValue
Task categoryAgent / MCP tool-use
Metric% task success
SaturationLow
Created byNot yet independently documented

How MCPMark-Verified Is Scored

Agents complete tasks requiring MCP tool interaction, with outcomes verified against expected results, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K394.5%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

MCPMark-Verified on Benchgen

No Benchgen results yet — be the first to run MCPMark-Verified.

MCPMark-Verified vs Other Benchmarks

BenchmarkWhat it testsSaturation
MCPMark-VerifiedMCP protocol tool-use correctnessLow
MCP AtlasMCP tool-use agentic workflowsLow
Toolathlon-VerifiedMulti-tool, multi-step agentic workflowsLow
BFCL v3Function-calling accuracyMedium

Run MCPMark-Verified on Your Model

Benchgen lets you run MCPMark-Verified against your own agent's MCP integration, tracking tool-use task success rates over time.

Frequently Asked Questions

What is MCPMark-Verified? MCPMark-Verified is a benchmark testing AI agents on Model Context Protocol (MCP) tool-use tasks, with verified task setups and success criteria.
What does a good score look like on MCPMark-Verified? Kimi K3 reports 94.5% as of July 2026, a very strong result reflecting reliable MCP tool-use capability.
Who created MCPMark-Verified? MCPMark-Verified's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.