Benchgen
Models/bytedance-ntu-s-lab/

LLaVA-OneVision 72B

DraftPublic

Model Details

LLaVA-OneVision 72B

Organization Modality License Released

Quick answer: LLaVA-OneVision 72B is the flagship, largest model in ByteDance and NTU S-Lab's LLaVA-OneVision open-weight family, designed for unified image, multi-image, and video understanding.

At a Glance

Where LLaVA-OneVision 72B leads

  • Strongest model in the LLaVA-OneVision family; competitive open-weight performance on multimodal search/reasoning benchmarks like MMSearch

Where it lags

  • Trails top proprietary models (GPT-4o, Claude 3.5 Sonnet) on complex multimodal search tasks

Best for: self-hosted deployments needing strong unified image/video understanding at scale.

What LLaVA-OneVision 72B Is

LLaVA-OneVision 72B is the largest model in the LLaVA-OneVision family, sharing the same unified single-image/multi-image/video architecture as its smaller siblings, scaled up for stronger overall performance.

Specifications

FieldValue
OrganizationByteDance, NTU S-Lab
ModalityMultimodal (text + vision + video)
LicenseOpen weights
Release dateSeptember 2024