Benchgen

MMAU — Results

RankModelScore
1gemini-3-1-pro82.5
2inkling77.2

MMAU

1 phaseActive

UMD GAMMA Lab's massive multitask audio understanding benchmark (Sakshi et al., 2024). 10,000 clips, 27 tasks across speech, sound, and music. Metric: % accuracy on test-mini. Apache 2.0.

Overview

MMAU

Category Metric Tasks License

Paper GitHub Dataset

Quick answer: MMAU (Massive Multi-Task Audio Understanding and Reasoning) is a benchmark from UMD GAMMA Lab (Sakshi et al., 2024) for evaluating large audio language models (LALMs) on expert-level audio understanding. It contains 10,000 carefully curated audio clips paired with natural-language questions spanning speech, environmental sounds, and music across 27 diverse tasks requiring both information retrieval and complex reasoning. Gemini 3.1 Pro leads the Inkling comparison set at 82.5% on the test-mini.

At a Glance

What it tests: A model's ability to understand and reason about audio content at an expert level — recognizing what is happening in audio clips, understanding speech nuance, identifying environmental sounds, and reasoning about musical content.

Why it matters: Most audio evaluation benchmarks test basic ASR or simple audio classification. MMAU requires domain expertise and complex reasoning about audio events, instruments, emotions, and speech content — more challenging than simple transcription or sound identification.

Task structure:

  • 12 information-retrieval tasks: What is being said? What sound is present? What instrument is playing?
  • 15 reasoning tasks: Why did this sound occur? What sequence of events happened? What emotion is expressed?

Benchmark Specifications

FieldValue
Task categoryAudio / multimodal understanding
Metric% accuracy (multiple-choice)
Number of tasks10,000 (1,000 test-mini + 9,000 test)
Audio domainsSpeech, environmental sounds, music
Task types27 (12 retrieval + 15 reasoning)
LicenseApache 2.0
VersionMMAU-v05.15.25 (latest)
Created byS Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, Dinesh Manocha
AffiliationGAMMA Lab, University of Maryland
PaperMMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark (arXiv 2410.19168)
GitHubSakshi113/MMAU
Dataset (test-mini)gamma-lab-umd/MMAU-test-mini on HuggingFace
Leaderboardsakshi113.github.io/mmau_homepage

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026). Only 2 models reported (native audio input required).

RankModelScoreWeights
1Gemini 3.1 Pro82.5%Closed
2Inkling77.2%Open

Last updated 2026-07-16.