Benchgen

DeepSearchQA — Results

RankModelScore
1kimi-k395
2muse-glimmer74.6
D

DeepSearchQA

1 phaseActive

Deep-research agent benchmark testing multi-hop web search and synthesis to answer complex questions. Metric: F1 score.

Overview

DeepSearchQA

Category Metric Saturation Created

Quick answer: DeepSearchQA is a deep-research agent benchmark that tests an AI system's ability to conduct multi-hop web search, synthesize information across sources, and produce accurate, well-supported answers to complex questions. Results are reported as an F1 score. Kimi K3 scores 95.0 F1 as of July 2026.

At a Glance

What it tests: An agent's ability to plan and execute multi-step web research — issuing search queries, evaluating source reliability, and synthesizing a final accurate answer.

Why it matters: As "deep research" agent products become common, DeepSearchQA offers a task-specific way to measure real research quality rather than static knowledge recall.

Known limitations: As an emerging benchmark, methodology details (e.g., exact search tool access, source pool) are not yet independently documented outside its citation by Moonshot AI.

What DeepSearchQA Measures

DeepSearchQA evaluates an AI agent's deep-research capability: given a complex question, the agent must plan and execute multi-hop web searches, evaluate and cross-reference retrieved sources, and synthesize a final answer supported by evidence. Unlike static QA benchmarks, this requires active tool use (web search) combined with reasoning over retrieved content.

Benchmark Specifications

FieldValue
Task categoryAgent / deep research
MetricF1 score
SaturationLow
Created byNot yet independently documented

How DeepSearchQA Is Scored

Agent-produced answers are compared against reference answers, with correctness scored via F1 (balancing precision and recall of the answer's key facts), reflecting both accuracy and completeness of the synthesized response.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K395.0 F1Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

DeepSearchQA on Benchgen

No Benchgen results yet — be the first to run DeepSearchQA.

DeepSearchQA vs Other Benchmarks

BenchmarkWhat it testsSaturation
DeepSearchQAMulti-hop web research & synthesisLow
ResearchRubricsRubric-graded research report qualityLow
SimpleQAShort-form factual QAHigh
BrowseCompComplex web-browsing agent tasksLow

Run DeepSearchQA on Your Model

Benchgen lets you run DeepSearchQA against your own model and search harness, tracking F1 performance over time to validate deep-research agent deployments.

Frequently Asked Questions

What is DeepSearchQA? DeepSearchQA is a deep-research agent benchmark testing multi-hop web search, source synthesis, and answer accuracy, scored via F1.
What does a good score look like on DeepSearchQA? Kimi K3 reports 95.0 F1 as of July 2026, a strong result indicating high precision and recall in synthesizing research answers.
Who created DeepSearchQA? DeepSearchQA's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.