Quick answer: Alpamayo 2 Super is NVIDIA's 34B-parameter foundation model for autonomous vehicle development, combining a 32B vision-language backbone (built on Cosmos 3 Super Reasoner) with a 2.3B diffusion action expert. It ranks first on LingoQA among nearly 40 evaluated models (79.2 Lingo-Judge score) and is openly licensed under OpenMDW-1.1 for commercial robotaxi and AV fleet deployment.
Where Alpamayo 2 Super leads
Where it lags
| Field | Value |
|---|---|
| Organization | NVIDIA |
| Parameters | 34B total (32B VLM backbone + 2.3B diffusion action expert) |
| Architecture | Vision-Language-Action (VLA); built on Cosmos 3 Super Reasoner |
| Modality | Image/video, text, egomotion history → text + trajectory |
| License | OpenMDW-1.1 (weights); Apache 2.0 (source code) |
| Release date | August 4, 2026 |
| Hugging Face | nvidia/Alpamayo2-Super |
Alpamayo 2 Super's own reported evaluations, per the NVIDIA model card:
| Benchmark | Score | Notes |
|---|---|---|
| LingoQA | 79.2 (Lingo-Judge) | #1 of ~40 models evaluated |
| AlpaSim (closed-loop, 910 scenarios) | 1.50 ± 0.13 | NVIDIA PhysicalAI-AV-NuRec dataset |
| Open-loop minADE₆ @ 6.4s | 0.911m | 937 challenging samples, PhysicalAI-AV dataset |
| Meta-Action IoU (Longitudinal) | 61.9% | vs. 35.8% for Qwen3-VL-32B-Instruct |
| Meta-Action IoU (Lateral) | 74.6% | vs. 47.5% for Qwen3-VL-32B-Instruct |
| Meta-Action IoU (Lane) | 73.5% | vs. 68.8% for Qwen3-VL-32B-Instruct |
| Auto-Labeling Accuracy | 65.2% | vs. 45.0% for Qwen3-VL-32B-Instruct |
| 2D Visual Grounding (IoU) | 71.0% | vs. 17.0% for Qwen3-VL-32B-Instruct |
| OOD Reasoning Score | 43.3 | vs. 39.6 for Qwen3-VL-32B-Instruct |
Scores sourced from NVIDIA's Alpamayo 2 Super model card and NVIDIA blog announcement, August 4, 2026.
This model isn’t on any benchmark leaderboard yet.