| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 51 |
1 phaseActive
World-knowledge visual question answering benchmark with a forced-answer setting (no abstention allowed). Metric: % accuracy.
Quick answer: WorldVQA ForceAnswer tests AI models on real-world visual question answering under a "forced-answer" protocol, meaning the model cannot abstain or say "I don't know" — every question must receive a definitive answer, penalizing both incorrect guesses and evasive non-answers. Kimi K3 scores 51.0% as of July 2026.
What it tests: A model's ability to answer real-world visual questions accurately when abstention is not allowed, testing both visual perception and confident, committed reasoning.
Why it matters: Many benchmarks allow models to hedge or abstain on uncertain questions, inflating apparent reliability. The forced-answer setting removes this escape hatch, giving a more honest accuracy signal.
Known limitations: As an emerging benchmark, exact question sourcing and world-knowledge domain coverage are not yet independently published outside its citation by Moonshot AI.
WorldVQA ForceAnswer evaluates visual question answering grounded in general world knowledge, under a forced-answer protocol that disallows abstention. This design choice ensures the reported accuracy reflects genuine visual-plus-world-knowledge competence rather than a model's tendency to hedge on difficult questions.
| Field | Value |
|---|---|
| Task category | Reasoning / visual question answering |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models must provide a definitive answer to every visual question (no abstention permitted), scored on % accuracy against ground-truth answers.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 51.0% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run WorldVQA ForceAnswer.
| Benchmark | What it tests | Saturation |
|---|---|---|
| WorldVQA ForceAnswer | Forced-answer real-world visual QA | Low |
| PerceptionBench | Fine-grained visual perception | Low |
| SimpleQA | Short-form factual QA | High |
| BabyVision | Early-stage visual reasoning | Low |
Benchgen lets you run WorldVQA ForceAnswer against your own multimodal model, tracking forced-answer accuracy to measure genuine visual-plus-world-knowledge reliability.