| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 58.5 |
1 phaseActive
Moonshot AI's fine-grained visual perception benchmark, testing detailed visual grounding and recognition. Metric: % accuracy.
Quick answer: PerceptionBench is a fine-grained visual perception benchmark introduced by Moonshot AI, testing a model's ability to accurately recognize, localize, and describe detailed visual elements — going beyond high-level scene understanding to precise visual grounding. Kimi K3 scores 58.5% as of July 2026.
What it tests: Fine-grained visual perception — precise object recognition, spatial localization, and detailed visual attribute identification, rather than coarse scene-level description.
Why it matters: Many vision benchmarks test high-level scene understanding, but downstream applications (e.g., GUI agents, robotics, document parsing) require precise, low-level visual grounding. PerceptionBench targets this gap directly.
Known limitations: As a benchmark introduced by Moonshot AI itself, independent third-party validation and cross-lab adoption are still limited.
PerceptionBench evaluates a model's fine-grained visual perception capabilities — its ability to accurately identify, localize, and describe specific visual elements within an image, such as precise object boundaries, spatial relationships, and fine attribute details. This complements broader multimodal reasoning benchmarks by isolating the underlying visual grounding capability that many downstream agentic and reasoning tasks depend on.
| Field | Value |
|---|---|
| Task category | Vision / fine-grained perception |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Moonshot AI |
| Announcement | kimi.com/blog/perception-bench |
Models are evaluated on their ability to correctly identify and localize fine-grained visual elements within images, scored as % accuracy against ground-truth annotations.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 58.5% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run PerceptionBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| PerceptionBench | Fine-grained visual perception | Low |
| WorldVQA ForceAnswer | Forced-answer visual question answering | Low |
| MMMU-Pro | Multi-discipline multimodal understanding | Low |
| BabyVision | Early-stage visual reasoning | Low |
Benchgen lets you run PerceptionBench against your own multimodal model, tracking fine-grained visual perception accuracy to catch regressions in downstream vision-dependent applications.