| Rank | Model | Score |
|---|---|---|
| 1 | lfm2-5-vl-3b | 73.1 |
| 2 | north-micro-vision-instruct | 62.2 |
| 3 | qwen3-7-plus | 0.869 |
| 4 | seed-2-1-pro | 0.867 |
| 5 | seed-2-1-turbo | 0.863 |
| 6 | qwen3-6-plus | 0.854 |
| 7 | qwen3-6-35b-a3b | 0.853 |
| 8 | qwen3-5-122b-a10b | 0.851 |
| 9 | qwen3-5-35b-a3b | 0.841 |
| 10 | qwen3-6-27b | 0.841 |
| 11 | qwen3-5-27b | 0.837 |
| 12 | qwen3-vl-235b-a22b-thinking | 0.813 |
| 13 | qwen3-vl-235b-a22b-instruct | 0.793 |
| 14 | qwen3-vl-32b-instruct | 0.79 |
| 15 | qwen3-vl-32b-thinking | 0.784 |
| 16 | qwen3-vl-30b-a3b-thinking | 0.774 |
| 17 | qwen3-vl-30b-a3b-instruct | 0.737 |
| 18 | qwen3-vl-8b-thinking | 0.735 |
| 19 | qwen3-vl-4b-thinking | 0.732 |
| 20 | qwen3-vl-8b-instruct | 0.715 |
| 21 | qwen3-vl-4b-instruct | 0.709 |
| 22 | qwen2-5-omni-7b | 0.703 |
| 23 | grok-1-5 | 0.687 |
1 phaseActive
700+ anonymized real-world images from vehicles and everyday scenes — evaluates multimodal AI on spatial understanding and scene comprehension. Released by xAI. Metric: accuracy.
Quick answer: RealWorldQA is a spatial understanding benchmark released by xAI in 2024 alongside the Grok-1.5 Vision preview. It contains over 700 anonymized real-world images taken from vehicles and everyday scenes, each with a question and easily verifiable answer. Qwen3.8 Max leads with 88.0% across 26 evaluated models.
RealWorldQA evaluates AI models on basic real-world spatial understanding — the kind of visual reasoning needed for autonomous driving, robotics, and scene comprehension. Images come from vehicle dashcams and everyday environments, testing whether models can accurately interpret spatial relationships, distances, and scene elements.
| Focus area | Examples |
|---|---|
| Spatial relationships | Left/right, front/behind, above/below objects |
| Distance estimation | Relative distances between objects in a scene |
| Object identification | Recognizing objects in natural settings |
| Scene understanding | Interpreting the overall context of a scene |
Each image is paired with a multiple-choice question and a verifiable ground-truth answer. Accuracy is the fraction of questions answered correctly, normalized to 0–1.
| Property | Value |
|---|---|
| Released | 2024 |
| Images | 700+ |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | Qwen3.8 Max (0.880) |
| Models evaluated | 26 |
What is RealWorldQA? RealWorldQA is a benchmark testing multimodal AI on real-world spatial understanding using 700+ anonymized images from vehicles and everyday scenes, released by xAI as part of the Grok-1.5 Vision evaluation.
Who created RealWorldQA? RealWorldQA was released by xAI (Elon Musk's AI company) alongside the Grok-1.5 Vision model preview in 2024. There is no associated arXiv paper.
What types of images does RealWorldQA use? Images are primarily from vehicle dashcams and real-world settings, anonymized to remove identifying information. They cover everyday scenes requiring spatial and contextual understanding.
What score does the best model achieve on RealWorldQA? Qwen3.8 Max leads with 0.880 (88.0%), followed by Qwen3.7-Plus at 0.869 and Seed 2.1 Pro at 0.867.