Quick answer: LLaVA-OneVision 7B is a compact open-weight multimodal model from ByteDance and NTU's S-Lab, and the top-scoring model among tracked models on the high-resolution MME-RealWorld benchmark.
Where LLaVA-OneVision 7B leads
Where it lags
Best for: efficient, self-hostable multimodal deployments needing strong high-resolution perception.
LLaVA-OneVision is designed as a single unified model handling single-image, multi-image, and video understanding tasks, built by ByteDance and NTU's S-Lab as an evolution of the LLaVA-NeXT family.
| Field | Value |
|---|---|
| Organization | ByteDance, NTU S-Lab |
| Modality | Multimodal (text + vision + video) |
| License | Open weights |
| Release date | September 2024 |