| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 95 |
| 2 | muse-glimmer | 74.6 |
1 phaseActive
Deep-research agent benchmark testing multi-hop web search and synthesis to answer complex questions. Metric: F1 score.
Quick answer: DeepSearchQA is a deep-research agent benchmark that tests an AI system's ability to conduct multi-hop web search, synthesize information across sources, and produce accurate, well-supported answers to complex questions. Results are reported as an F1 score. Kimi K3 scores 95.0 F1 as of July 2026.
What it tests: An agent's ability to plan and execute multi-step web research — issuing search queries, evaluating source reliability, and synthesizing a final accurate answer.
Why it matters: As "deep research" agent products become common, DeepSearchQA offers a task-specific way to measure real research quality rather than static knowledge recall.
Known limitations: As an emerging benchmark, methodology details (e.g., exact search tool access, source pool) are not yet independently documented outside its citation by Moonshot AI.
DeepSearchQA evaluates an AI agent's deep-research capability: given a complex question, the agent must plan and execute multi-hop web searches, evaluate and cross-reference retrieved sources, and synthesize a final answer supported by evidence. Unlike static QA benchmarks, this requires active tool use (web search) combined with reasoning over retrieved content.
| Field | Value |
|---|---|
| Task category | Agent / deep research |
| Metric | F1 score |
| Saturation | Low |
| Created by | Not yet independently documented |
Agent-produced answers are compared against reference answers, with correctness scored via F1 (balancing precision and recall of the answer's key facts), reflecting both accuracy and completeness of the synthesized response.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 95.0 F1 | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run DeepSearchQA.
| Benchmark | What it tests | Saturation |
|---|---|---|
| DeepSearchQA | Multi-hop web research & synthesis | Low |
| ResearchRubrics | Rubric-graded research report quality | Low |
| SimpleQA | Short-form factual QA | High |
| BrowseComp | Complex web-browsing agent tasks | Low |
Benchgen lets you run DeepSearchQA against your own model and search harness, tracking F1 performance over time to validate deep-research agent deployments.