Benchgen

APEX-Agents — Results

RankModelScore
1grok-4-657.5
2kimi-k341
3solar-pro-418.7
A

APEX-Agents

1 phaseActive

Mercor's professional-grade agentic task benchmark, testing AI agents on economically valuable knowledge-work workflows. Metric: % task success.

Overview

APEX-Agents

Category Metric Saturation Created

Dataset

Quick answer: APEX-Agents is a professional-grade agentic task benchmark from Mercor, a talent and evaluation platform, testing AI agents on economically valuable, real-world knowledge-work workflows. It's designed to complement broader agent benchmarks by focusing on tasks that mirror actual professional work. Kimi K3 scores 41.0% as of July 2026.

At a Glance

What it tests: An AI agent's ability to complete professional-grade knowledge-work tasks end-to-end, evaluated against Mercor's curated task pool.

Why it matters: Mercor's evaluation network draws on real hiring and workforce data, giving APEX-Agents a grounding in genuinely economically valuable work rather than synthetic agent tasks.

Known limitations: As a vendor-curated benchmark without a published academic paper, methodology details are less transparent than peer-reviewed benchmarks.

What APEX-Agents Measures

APEX-Agents evaluates AI agents on professional-grade tasks curated by Mercor, spanning knowledge-work domains that reflect real hiring and workforce assessment scenarios. Tasks are designed to test sustained, multi-step task completion rather than single-turn question answering, aligning with how agents would be deployed in real professional contexts.

Benchmark Specifications

FieldValue
Task categoryAgent / professional knowledge work
Metric% task success
SaturationLow
Created byMercor
Datasetmercor/apex-agents on HuggingFace

How APEX-Agents Is Scored

Agents attempt each professional-grade task, and outcomes are verified against Mercor's task-specific success criteria, producing an aggregate % task success rate.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K341.0%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

APEX-Agents on Benchgen

No Benchgen results yet — be the first to run APEX-Agents.

APEX-Agents vs Other Benchmarks

BenchmarkWhat it testsSaturation
APEX-AgentsProfessional-grade knowledge-work agent tasksLow
Agents' Last ExamEconomically valuable long-horizon agent tasksLow
JobBenchJob/workforce-oriented agent tasksLow
MCP AtlasMCP tool-use agentic workflowsLow

Run APEX-Agents on Your Model

Benchgen lets you run APEX-Agents against your own model and harness, tracking professional-task success rates over time to validate agent deployments before production rollout.

Frequently Asked Questions

What is APEX-Agents? APEX-Agents is a professional-grade agentic task benchmark from Mercor, testing AI agents on economically valuable, real-world knowledge-work workflows.
What does a good score look like on APEX-Agents? Kimi K3 reports 41.0% as of July 2026. Given the professional-grade difficulty of the task pool, scores in the 40%+ range represent frontier-level agentic performance.
Who created APEX-Agents? APEX-Agents is created and maintained by Mercor, a talent and workforce evaluation platform that curates professional-grade agentic benchmarks.