| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 85.7 |
1 phaseActive
Early-stage visual reasoning benchmark testing foundational perception tasks that remain surprisingly hard for multimodal models. Metric: % accuracy.
Quick answer: BabyVision tests AI models on deceptively simple, "early-stage" visual reasoning tasks — the kind of basic perceptual judgments a young child could make — that nonetheless remain surprisingly difficult for current multimodal models to answer reliably without tool assistance. Kimi K3 scores 85.7% when allowed to use Python-based tool assistance, as of July 2026.
What it tests: Fundamental visual perception and reasoning — basic counting, spatial relationships, object comparison — designed to be intuitive for humans but revealing of subtle model weaknesses.
Why it matters: Frontier models can solve advanced graduate-level reasoning problems while still occasionally failing simple perceptual tasks. BabyVision surfaces this gap, which matters for real-world reliability.
Known limitations: As an emerging benchmark, exact task composition is not yet independently published outside its citation by Moonshot AI. Scores can vary significantly depending on whether tool assistance (e.g., Python for counting/measurement) is permitted.
BabyVision evaluates a model's performance on foundational, seemingly simple visual reasoning tasks — such as counting objects, judging relative size or position, and basic scene comparison — that remain a known weak point for multimodal models despite their strength on harder benchmarks. Results can be reported both with and without tool assistance (e.g., allowing a model to write and execute Python code to verify counts), since tool use often substantially improves accuracy on these tasks.
| Field | Value |
|---|---|
| Task category | Reasoning / basic visual perception |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models answer basic visual perception questions, scored on % accuracy against ground truth. Scores are often reported both with and without permitted tool assistance (e.g., Python code execution) to isolate raw perceptual accuracy from tool-augmented accuracy.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 85.7% (w/ Python) | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run BabyVision.
| Benchmark | What it tests | Saturation |
|---|---|---|
| BabyVision | Fundamental/basic visual perception | Low |
| PerceptionBench | Fine-grained visual perception | Low |
| WorldVQA ForceAnswer | Forced-answer real-world visual QA | Low |
| ZeroBench | Extremely hard visual reasoning | Very low |
Benchgen lets you run BabyVision against your own multimodal model, tracking basic perceptual accuracy with and without tool assistance to catch subtle reliability gaps.