Benchgen
Models/meta/

Llama 3.2 90B Instruct

DraftPublic

Model Details

Llama 3.2 90B Vision Instruct

Organization Parameters Context License Weights Multimodal Released

Quick answer: Llama 3.2 90B Vision Instruct is Meta's September 2024 open-weight multimodal model, scoring 68.0% on MATH and 86.0% on MMLU. With a 128K context window and vision capability, it was the first large open-weight model to support image understanding at frontier text-model scale.

At a Glance

Where Llama 3.2 90B leads

  • First large open-weight multimodal model at frontier scale
  • 86.0% MMLU — strong general knowledge
  • 128K context window
  • Vision capability: understands images alongside text
  • Llama 3.2 Community License (free for most commercial uses)

Where it lags

  • 68.0% MATH — moderate mathematics
  • December 2023 knowledge cutoff — earlier than Llama 3.3 70B
  • Superseded by Llama 4 Maverick for multimodal tasks (1M context, better benchmarks)
  • 90B parameters — heavier than Llama 3.3 70B for text-only tasks

Best for: Open-weight vision-language tasks where Llama 4 Maverick is too recent for existing integrations; image analysis pipelines on Llama 3.2 stack.

What Llama 3.2 90B Vision Instruct Is

Llama 3.2 90B Vision Instruct (released September 25, 2024) was Meta's first large multimodal model — adding image understanding to the flagship Llama scale. Alongside the text-only Llama 3.2 3B and 1B models, it marked Meta's push into multimodal open-weight AI.

At 90B parameters, it provides text and vision capability at frontier scale — competitive with GPT-4V and Claude's vision capabilities at launch. The model handles image captioning, visual Q&A, document understanding, and chart analysis alongside standard text tasks.

For new multimodal open-weight deployments, Llama 4 Maverick (MoE, 1M context, better benchmarks) is the recommended choice. Llama 3.2 90B remains relevant for teams with existing 3.2 integrations.

Specifications

FieldValue
OrganizationMeta
Parameters90B (dense)
Context window128,000 tokens
LicenseLlama 3.2 Community License
HuggingFacemeta-llama/Llama-3.2-90B-Vision-Instruct
Release dateSeptember 25, 2024
Knowledge cutoffDecember 2023
ModalityText + Vision (multimodal)

Pricing

Open weights under Llama 3.2 Community License. Available via Meta AI API and third-party providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
MATH68.0%Benchgen evaluation2025-07
MMLU86.0%Benchgen evaluation2025-07

Llama 3.2 90B vs Alternatives

ModelMMLUVisionContextLicense
Llama 3.2 90B Vision Instruct86.0%Yes128KLlama 3.2
Llama 4 MaverickYes1MLlama 4
Gemma 3 27BYes128KGemma ToU

Llama 4 Maverick: better benchmarks, 1M context, multimodal — preferred for new deployments. Llama 3.2 90B for existing integrations or when Llama 3.2 stack is required.

Frequently Asked Questions

What is Llama 3.2 90B Vision Instruct? Meta's September 2024 open-weight multimodal 90B model scoring 68.0% MATH and 86.0% MMLU with 128K context. First large open-weight model with vision capability.
Does Llama 3.2 90B support images? Yes. Llama 3.2 90B Vision Instruct supports text and image inputs — vision-language tasks including image captioning, visual Q&A, and document understanding.

Specs from Meta's Llama 3.2 announcement (September 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.