Benchgen

BabyVision — Results

RankModelScore
1kimi-k385.7
B

BabyVision

1 phaseActive

Early-stage visual reasoning benchmark testing foundational perception tasks that remain surprisingly hard for multimodal models. Metric: % accuracy.

Overview

BabyVision

Category Metric Saturation Created

Quick answer: BabyVision tests AI models on deceptively simple, "early-stage" visual reasoning tasks — the kind of basic perceptual judgments a young child could make — that nonetheless remain surprisingly difficult for current multimodal models to answer reliably without tool assistance. Kimi K3 scores 85.7% when allowed to use Python-based tool assistance, as of July 2026.

At a Glance

What it tests: Fundamental visual perception and reasoning — basic counting, spatial relationships, object comparison — designed to be intuitive for humans but revealing of subtle model weaknesses.

Why it matters: Frontier models can solve advanced graduate-level reasoning problems while still occasionally failing simple perceptual tasks. BabyVision surfaces this gap, which matters for real-world reliability.

Known limitations: As an emerging benchmark, exact task composition is not yet independently published outside its citation by Moonshot AI. Scores can vary significantly depending on whether tool assistance (e.g., Python for counting/measurement) is permitted.

What BabyVision Measures

BabyVision evaluates a model's performance on foundational, seemingly simple visual reasoning tasks — such as counting objects, judging relative size or position, and basic scene comparison — that remain a known weak point for multimodal models despite their strength on harder benchmarks. Results can be reported both with and without tool assistance (e.g., allowing a model to write and execute Python code to verify counts), since tool use often substantially improves accuracy on these tasks.

Benchmark Specifications

FieldValue
Task categoryReasoning / basic visual perception
Metric% accuracy
SaturationLow
Created byNot yet independently documented

How BabyVision Is Scored

Models answer basic visual perception questions, scored on % accuracy against ground truth. Scores are often reported both with and without permitted tool assistance (e.g., Python code execution) to isolate raw perceptual accuracy from tool-augmented accuracy.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K385.7% (w/ Python)Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

BabyVision on Benchgen

No Benchgen results yet — be the first to run BabyVision.

BabyVision vs Other Benchmarks

BenchmarkWhat it testsSaturation
BabyVisionFundamental/basic visual perceptionLow
PerceptionBenchFine-grained visual perceptionLow
WorldVQA ForceAnswerForced-answer real-world visual QALow
ZeroBenchExtremely hard visual reasoningVery low

Run BabyVision on Your Model

Benchgen lets you run BabyVision against your own multimodal model, tracking basic perceptual accuracy with and without tool assistance to catch subtle reliability gaps.

Frequently Asked Questions

What is BabyVision? BabyVision is a benchmark testing AI models on deceptively simple, foundational visual reasoning tasks that reveal subtle perceptual weaknesses even in frontier multimodal models.
What does a good score look like on BabyVision? Kimi K3 reports 85.7% with Python tool assistance as of July 2026, indicating that tool use substantially helps close the gap on basic perceptual tasks like counting.
Who created BabyVision? BabyVision's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.