Benchgen
Models/zhipu-ai/

CogVLM

DraftPublic

Model Details

CogVLM

Organization Modality License Released

Quick answer: CogVLM is Zhipu AI's 17B-parameter open-weight visual language model, notable for using a dedicated "visual expert" module within its attention layers rather than a simple projection, and a widely-cited early open baseline in MLLM benchmark research.

At a Glance

Where CogVLM leads

  • Fully open weights, early strong open baseline widely cited across MLLM benchmark papers

Where it lags

  • Superseded by newer open-weight families on most current benchmarks

Best for: open-weight comparison baseline in multimodal benchmark literature.

What CogVLM Is

CogVLM introduces a "visual expert" module — trainable visual-specific parameters added to each attention and feed-forward layer of a frozen LLM — allowing deep fusion of vision and language features rather than shallow projection, built on a Vicuna-7B backbone.

Specifications

FieldValue
OrganizationZhipu AI
ModalityMultimodal (text + vision)
LicenseOpen weights
Release dateNovember 2023