Quick answer: Longcat Flash Thinking is Meituan's July 2025 thinking model scoring 79.4% LiveCodeBench, 82.6% MMLU-Pro, 56.6% BrowseComp, 74.4% BFCL-v3, and 50.3% ARC-AGI. Proprietary.
Where Longcat Flash Thinking leads
Where it lags
Best for: LiveCodeBench-heavy coding pipelines; function-calling (BFCL) heavy workloads; Meituan cloud integrations.
Longcat Flash Thinking is Meituan's July 2025 "thinking" (chain-of-thought) model from Meituan AI Lab. Meituan is primarily a Chinese delivery/commerce company but has invested significantly in LLM research. The "Flash" variant is optimised for speed while maintaining reasoning depth.
The 79.4% LiveCodeBench and 74.4% BFCL-v3 combination is the model's key strength — competitive coding with strong API/tool calling.
| Field | Value |
|---|---|
| Organization | Meituan |
| License | Proprietary |
| Release date | July 2025 |
| Modality | Text only |
Available via Meituan AI API. Refer to Meituan pricing.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| LiveCodeBench | 79.4% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 82.6% | Benchgen evaluation | 2025-07 |
| BrowseComp | 56.6% | Benchgen evaluation | 2025-07 |
| BFCL-v3 | 74.4% | Benchgen evaluation | 2025-07 |
| ARC-AGI | 50.3% | Benchgen evaluation | 2025-07 |
| Model | LiveCodeBench | MMLU-Pro | BrowseComp | License |
|---|---|---|---|---|
| Longcat Flash Thinking | 79.4% | 82.6% | 56.6% | Proprietary |
| MiniMax M2 | 83.0% | 82% | 44.0% | MiniMax Commercial |
| Nemotron 3 Super 120B | 81.2% | 83.7% | — | NVIDIA Open |
Longcat Flash Thinking leads on BrowseComp (56.6% vs 44.0% MiniMax M2) but trails on LiveCodeBench. For balanced web research + coding with function calling (BFCL-v3): Longcat Flash Thinking.
Specs from Meituan's Longcat Flash Thinking release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.