Quick answer: Meta Muse Spark is Meta's June 2025 agentic reasoning model scoring 89.5% GPQA Diamond, 88.1% MCP Atlas, 80.4% MMMU-Pro, 77.4% SWE-Bench Verified, and 52.4% SWE-Bench Pro. Proprietary.
Where Muse Spark leads
Where it lags
Best for: Agentic coding and tool-use pipelines (MCP); GPQA-class science reasoning; SWE-Bench Verified-intensive applications; Meta AI ecosystem.
Meta Muse Spark is Meta's June 2025 model, positioned as an agentic "spark" variant with strong tool-use and coding capabilities. The 88.1% MCP Atlas score makes it one of the top performers on this agentic tool-use benchmark.
The 89.5% GPQA Diamond and 77.4% SWE-Bench Verified combination makes Muse Spark competitive with GPT-5.2 Pro (93.2% GPQA) and Claude Opus 4 for science + coding combined.
| Field | Value |
|---|---|
| Organization | Meta |
| License | Proprietary |
| Release date | June 2025 |
| Modality | Text and vision |
Available via Meta AI API. Refer to Meta pricing for current rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GPQA Diamond | 89.5% | Benchgen evaluation | 2025-06 |
| MCP Atlas | 88.1% | Benchgen evaluation | 2025-06 |
| MMMU-Pro | 80.4% | Benchgen evaluation | 2025-06 |
| SWE-Bench Verified | 77.4% | Benchgen evaluation | 2025-06 |
| SWE-Bench Pro | 52.4% | Benchgen evaluation | 2025-06 |
| Humanity's Last Exam | 40.6% | Benchgen evaluation | 2025-06 |
| Model | GPQA Diamond | SWE-Bench Verified | MCP Atlas | License |
|---|---|---|---|---|
| Muse Spark | 89.5% | 77.4% | 88.1% | Proprietary |
| GPT-5.2 Pro 2025-12-11 | 93.2% | — | — | Proprietary |
| Claude Opus 4 | 90.1% | — | — | Proprietary |
| Kimi K2.6 | 90.5% | — | — | Apache 2.0 |
Muse Spark leads on MCP Atlas (88.1%) and SWE-Bench Verified (77.4%) among comparably-released mid-2025 models. For open-weight GPQA competitors: Kimi K2.6 (90.5% GPQA, Apache 2.0).
Specs from Meta's Muse Spark release (June 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.