Benchgen

PerceptionBench — Results

RankModelScore
1kimi-k358.5
P

PerceptionBench

1 phaseActive

Moonshot AI's fine-grained visual perception benchmark, testing detailed visual grounding and recognition. Metric: % accuracy.

Overview

PerceptionBench

Category Metric Saturation Created

Website

Quick answer: PerceptionBench is a fine-grained visual perception benchmark introduced by Moonshot AI, testing a model's ability to accurately recognize, localize, and describe detailed visual elements — going beyond high-level scene understanding to precise visual grounding. Kimi K3 scores 58.5% as of July 2026.

At a Glance

What it tests: Fine-grained visual perception — precise object recognition, spatial localization, and detailed visual attribute identification, rather than coarse scene-level description.

Why it matters: Many vision benchmarks test high-level scene understanding, but downstream applications (e.g., GUI agents, robotics, document parsing) require precise, low-level visual grounding. PerceptionBench targets this gap directly.

Known limitations: As a benchmark introduced by Moonshot AI itself, independent third-party validation and cross-lab adoption are still limited.

What PerceptionBench Measures

PerceptionBench evaluates a model's fine-grained visual perception capabilities — its ability to accurately identify, localize, and describe specific visual elements within an image, such as precise object boundaries, spatial relationships, and fine attribute details. This complements broader multimodal reasoning benchmarks by isolating the underlying visual grounding capability that many downstream agentic and reasoning tasks depend on.

Benchmark Specifications

FieldValue
Task categoryVision / fine-grained perception
Metric% accuracy
SaturationLow
Created byMoonshot AI
Announcementkimi.com/blog/perception-bench

How PerceptionBench Is Scored

Models are evaluated on their ability to correctly identify and localize fine-grained visual elements within images, scored as % accuracy against ground-truth annotations.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K358.5%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

PerceptionBench on Benchgen

No Benchgen results yet — be the first to run PerceptionBench.

PerceptionBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
PerceptionBenchFine-grained visual perceptionLow
WorldVQA ForceAnswerForced-answer visual question answeringLow
MMMU-ProMulti-discipline multimodal understandingLow
BabyVisionEarly-stage visual reasoningLow

Run PerceptionBench on Your Model

Benchgen lets you run PerceptionBench against your own multimodal model, tracking fine-grained visual perception accuracy to catch regressions in downstream vision-dependent applications.

Frequently Asked Questions

What is PerceptionBench? PerceptionBench is a fine-grained visual perception benchmark introduced by Moonshot AI, testing precise object recognition, localization, and visual attribute identification.
What does a good score look like on PerceptionBench? Kimi K3 reports 58.5% as of July 2026. Given the fine-grained difficulty of precise visual grounding tasks, this represents a solid frontier-level result.
Who created PerceptionBench? PerceptionBench was introduced by Moonshot AI, the developer of the Kimi model family, as part of their visual perception evaluation suite.