Benchgen
Models/opengvlab-shanghai-ai-laboratory/

InternVL2.5 78B

DraftPublic

Model Details

InternVL2.5 78B

Organization Modality License Released

Quick answer: InternVL2.5 78B is OpenGVLab's flagship open-weight multimodal model, part of the InternVL2.5 family built on a progressive scaling strategy, and the top-scoring open-source model on the MTVQA multilingual text-centric visual QA benchmark.

At a Glance

Where InternVL2.5 78B leads

  • Strong multilingual OCR and scene-text comprehension, outperforming larger proprietary models on MTVQA
  • Fully open weights, enabling self-hosted deployment and fine-tuning

Where it lags

  • Requires substantial GPU infrastructure to self-host at 78B parameters
  • Trails top proprietary models on some general reasoning benchmarks

Best for: teams needing strong open-weight multilingual visual document/OCR understanding with full deployment control.

What InternVL2.5 78B Is

InternVL2.5 78B is part of the InternVL2.5 series from OpenGVLab (Shanghai AI Laboratory), an open-weight multimodal LLM family built on a progressive scaling strategy that expands performance boundaries of open-source vision-language models. It combines a large vision encoder with an LLM backbone to handle high-resolution image and multi-image understanding tasks.

Specifications

FieldValue
OrganizationOpenGVLab (Shanghai AI Laboratory)
ModalityMultimodal (text + vision)
LicenseOpen weights (research license)
Release dateDecember 2024