Benchgen
Models/bytedance-ntu-s-lab/

LLaVA-OneVision 7B

DraftPublic

Model Details

LLaVA-OneVision 7B

Organization Modality License Released

Quick answer: LLaVA-OneVision 7B is a compact open-weight multimodal model from ByteDance and NTU's S-Lab, and the top-scoring model among tracked models on the high-resolution MME-RealWorld benchmark.

At a Glance

Where LLaVA-OneVision 7B leads

  • Top MME-RealWorld score among tracked models despite a 7B parameter count
  • Designed for unified image, multi-image, and video understanding in one model

Where it lags

  • Smaller parameter count can limit performance on tasks requiring broader world knowledge

Best for: efficient, self-hostable multimodal deployments needing strong high-resolution perception.

What LLaVA-OneVision 7B Is

LLaVA-OneVision is designed as a single unified model handling single-image, multi-image, and video understanding tasks, built by ByteDance and NTU's S-Lab as an evolution of the LLaVA-NeXT family.

Specifications

FieldValue
OrganizationByteDance, NTU S-Lab
ModalityMultimodal (text + vision + video)
LicenseOpen weights
Release dateSeptember 2024