Benchgen
Models/meta/

Muse Spark

DraftPublic

Model Details

Meta Muse Spark

Organization License Released

Quick answer: Meta Muse Spark is Meta's June 2025 agentic reasoning model scoring 89.5% GPQA Diamond, 88.1% MCP Atlas, 80.4% MMMU-Pro, 77.4% SWE-Bench Verified, and 52.4% SWE-Bench Pro. Proprietary.

At a Glance

Where Muse Spark leads

  • 89.5% GPQA Diamond — among the top science scores (near Grok-4, GPT-5.2 Pro)
  • 88.1% MCP Atlas — exceptional agentic tool use
  • 80.4% MMMU-Pro — strong multimodal reasoning
  • 77.4% SWE-Bench Verified — excellent software engineering
  • 52.4% SWE-Bench Pro — strong hard coding tasks

Where it lags

  • 40.56% HLE — moderate frontier exam reasoning
  • Proprietary: no open weights
  • Meta API access required

Best for: Agentic coding and tool-use pipelines (MCP); GPQA-class science reasoning; SWE-Bench Verified-intensive applications; Meta AI ecosystem.

What Meta Muse Spark Is

Meta Muse Spark is Meta's June 2025 model, positioned as an agentic "spark" variant with strong tool-use and coding capabilities. The 88.1% MCP Atlas score makes it one of the top performers on this agentic tool-use benchmark.

The 89.5% GPQA Diamond and 77.4% SWE-Bench Verified combination makes Muse Spark competitive with GPT-5.2 Pro (93.2% GPQA) and Claude Opus 4 for science + coding combined.

Specifications

FieldValue
OrganizationMeta
LicenseProprietary
Release dateJune 2025
ModalityText and vision

Pricing

Available via Meta AI API. Refer to Meta pricing for current rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
GPQA Diamond89.5%Benchgen evaluation2025-06
MCP Atlas88.1%Benchgen evaluation2025-06
MMMU-Pro80.4%Benchgen evaluation2025-06
SWE-Bench Verified77.4%Benchgen evaluation2025-06
SWE-Bench Pro52.4%Benchgen evaluation2025-06
Humanity's Last Exam40.6%Benchgen evaluation2025-06

Muse Spark vs Alternatives

ModelGPQA DiamondSWE-Bench VerifiedMCP AtlasLicense
Muse Spark89.5%77.4%88.1%Proprietary
GPT-5.2 Pro 2025-12-1193.2%Proprietary
Claude Opus 490.1%Proprietary
Kimi K2.690.5%Apache 2.0

Muse Spark leads on MCP Atlas (88.1%) and SWE-Bench Verified (77.4%) among comparably-released mid-2025 models. For open-weight GPQA competitors: Kimi K2.6 (90.5% GPQA, Apache 2.0).

Frequently Asked Questions

What is Meta Muse Spark? Meta's June 2025 agentic reasoning model scoring 89.5% GPQA Diamond, 88.1% MCP Atlas, 77.4% SWE-Bench Verified. Proprietary — strong agentic tool-use.

Specs from Meta's Muse Spark release (June 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.