Quick answer: Nex-N2.5-Max is Nex-AGI's flagship model in the N2.5 family — a 1.6-trillion-parameter, text-only Mixture-of-Experts model and the lab's first complete post-training effort at trillion-parameter scale. It scores 86.1% on Terminal-Bench 2.1, 65.7% on SWE-Bench Pro, and 65.6% on DeepSWE v1.1, positioning it as Nex-AGI's top text-agent model, released alongside the smaller multimodal Nex-N2.5-Pro and Nex-N2.5-mini.
Where Nex N2.5 Max leads
Where it lags
Best for: teams wanting Nex-AGI's strongest open-weight text-agent model for coding/tool-use workloads where multimodal input isn't required.
Nex-N2.5-Max is the largest model in Nex-AGI's second-generation N2.5 release, built on a 1.6-trillion-parameter, text-only Mixture-of-Experts foundation model. Nex-AGI describes it as their first complete post-training effort carried out at full trillion-parameter scale, following broader expansion of agent training environments, task types, and productivity scenarios versus the prior N2 generation.
Unlike its siblings Nex-N2.5-Pro and Nex-N2.5-mini — which continue the multimodal, computer-use-focused line — Max is text-only, positioned as the flagship for long-horizon coding and tool-use agent tasks rather than visual grounding. It supports the same reasoning_effort control (none/medium/high) as the rest of the family and uses a deepseek-r1-style reasoning parser in its reference SGLang deployment (contrasted with the qwen3 parser used for Pro/mini), suggesting a distinct post-training reasoning-trace format at this scale.
| Field | Value |
|---|---|
| Organization | Nex-AGI |
| Parameters | 1.6T total (MoE); active-parameter count not disclosed |
| Context window | Not disclosed in model card |
| Architecture | Text-only Mixture-of-Experts; NexAU harness for coding evaluation |
| License | Open weights (Hugging Face + ModelScope) |
| Release date | September 2026 |
| Modality | Text only |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
Open-weight model; self-hosting cost only. Multi-node (2×8×H200) deployment recipe published by Nex-AGI; no hosted-API listing found at time of writing.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Terminal-Bench 2.1 | 86.1% | Nex-N2.5-Max model card | 2026-09 |
| SWE-Bench Pro | 65.7% | Nex-N2.5-Max model card | 2026-09 |
| DeepSWE | 65.6% (v1.1) | Nex-N2.5-Max model card | 2026-09 |
| AutomationBench | 50.2% (v1.0.6) | Nex-N2.5-Max model card | 2026-09 |
| Toolathlon Verified | 74.7% | Nex-N2.5-Max model card | 2026-09 |
| GDPval-AA v2 | 1713 | Nex-N2.5-Max model card | 2026-09 |
| Job Bench | 53.6% | Nex-N2.5-Max model card | 2026-09 |
| BrowseComp | 92.6% | Nex-N2.5-Max model card | 2026-09 |
Scores above are reported by Nex-AGI and shown for context. They are not Benchgen measurements. Coding tasks use Nex-AGI's NexAU harness; sampling uses temperature=0.7, top_p=0.95, top_k=40. Nex-N2.5-Max is text-only and is not evaluated on the multimodal benchmark table (OSWorld, WebArena, Vision2Web, etc.) published for Nex-N2.5-Pro/mini.
| Model | Terminal-Bench 2.1 | SWE-Bench Pro | Job Bench | Price (in/out per 1M) |
|---|---|---|---|---|
| Nex N2.5 Max | 86.1% | 65.7% | 53.6% | Open weights |
| Claude Opus 5 | 89.1% | 79.2% | 65.7% | — |
| Nex N2.5 Pro | 82.7% | 61.2% | 41.4% | Open weights |
| DeepSeek V4 Pro | 87.9% | 67.7% | 54.1% | $0.435/$— |
Nex-N2.5-Max closes much of the gap to closed frontier models (Claude Opus 5, GPT-5.6 Sol) on agentic coding benchmarks versus its own smaller Pro sibling, but still trails Claude Opus 5 and DeepSeek V4 Pro on SWE-Bench Pro and Job Bench specifically — the trade-off is open weights and self-hostability rather than category-leading raw scores.
As Nex-AGI's first full trillion-parameter-scale post-training run, Nex-N2.5-Max's benchmark profile shows broad, consistent gains over Nex-N2.5-Pro across coding and tool-use evals (Terminal-Bench 2.1, SWE-Bench Pro, DeepSWE v1.1, Toolathlon Verified, Job Bench) rather than a narrow specialization — evidence the scale-up primarily improved general agentic reliability rather than one task category. Because it's text-only, teams needing computer-use/vision-grounded agent behavior should evaluate Nex-N2.5-Pro instead; Max is the better fit for pure coding/tool-orchestration agents at self-hosted scale.
from openai import OpenAI # self-host via the published SGLang multi-node recipe, or check for OpenRouter/provider listings
client = OpenAI(api_key="YOUR_API_KEY", base_url="http://localhost:8000/v1")
response = client.chat.completions.create(
model="nex-agi/nex-n2.5-max",
messages=[{"role": "user", "content": "Refactor this repository's test suite..."}],
)
print(response.choices[0].message.content)Specs and scores sourced from Nex-AGI's official Hugging Face model card. Last updated 2026-09-15.