| Rank | Model | Score |
|---|---|---|
| 1 | grok-4-6 | 57.5 |
| 2 | kimi-k3 | 41 |
| 3 | solar-pro-4 | 18.7 |
1 phaseActive
Mercor's professional-grade agentic task benchmark, testing AI agents on economically valuable knowledge-work workflows. Metric: % task success.
Quick answer: APEX-Agents is a professional-grade agentic task benchmark from Mercor, a talent and evaluation platform, testing AI agents on economically valuable, real-world knowledge-work workflows. It's designed to complement broader agent benchmarks by focusing on tasks that mirror actual professional work. Kimi K3 scores 41.0% as of July 2026.
What it tests: An AI agent's ability to complete professional-grade knowledge-work tasks end-to-end, evaluated against Mercor's curated task pool.
Why it matters: Mercor's evaluation network draws on real hiring and workforce data, giving APEX-Agents a grounding in genuinely economically valuable work rather than synthetic agent tasks.
Known limitations: As a vendor-curated benchmark without a published academic paper, methodology details are less transparent than peer-reviewed benchmarks.
APEX-Agents evaluates AI agents on professional-grade tasks curated by Mercor, spanning knowledge-work domains that reflect real hiring and workforce assessment scenarios. Tasks are designed to test sustained, multi-step task completion rather than single-turn question answering, aligning with how agents would be deployed in real professional contexts.
| Field | Value |
|---|---|
| Task category | Agent / professional knowledge work |
| Metric | % task success |
| Saturation | Low |
| Created by | Mercor |
| Dataset | mercor/apex-agents on HuggingFace |
Agents attempt each professional-grade task, and outcomes are verified against Mercor's task-specific success criteria, producing an aggregate % task success rate.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 41.0% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run APEX-Agents.
| Benchmark | What it tests | Saturation |
|---|---|---|
| APEX-Agents | Professional-grade knowledge-work agent tasks | Low |
| Agents' Last Exam | Economically valuable long-horizon agent tasks | Low |
| JobBench | Job/workforce-oriented agent tasks | Low |
| MCP Atlas | MCP tool-use agentic workflows | Low |
Benchgen lets you run APEX-Agents against your own model and harness, tracking professional-task success rates over time to validate agent deployments before production rollout.