Benchgen
Models/nex-agi/

Nex-N2.5-Max

DraftPublic

Model Details

Nex N2.5 Max

Organization License Modality Released

Quick answer: Nex-N2.5-Max is Nex-AGI's flagship model in the N2.5 family — a 1.6-trillion-parameter, text-only Mixture-of-Experts model and the lab's first complete post-training effort at trillion-parameter scale. It scores 86.1% on Terminal-Bench 2.1, 65.7% on SWE-Bench Pro, and 65.6% on DeepSWE v1.1, positioning it as Nex-AGI's top text-agent model, released alongside the smaller multimodal Nex-N2.5-Pro and Nex-N2.5-mini.

At a Glance

Where Nex N2.5 Max leads

  • 86.1% Terminal-Bench 2.1 — ahead of GPT-5.6 Sol (88.8%)/Kimi-K3 (88.3%) is not quite matched, but close to frontier tier and best in the Nex family
  • 65.7% SWE-Bench Pro and 65.6% DeepSWE v1.1 — clear step up from Nex-N2.5-Pro (61.2% / 55.8%)
  • 74.7% Toolathlon Verified, 53.6% Job Bench — leads its own family on agentic/tool-use tasks
  • First trillion-parameter-scale post-training run completed by Nex-AGI — signals infra/training maturity at that scale

Where it lags

  • Text-only — no native multimodal input, unlike Nex-N2.5-Pro
  • 92.6% BrowseComp trails GPT-5.6 Sol (90.4%)/Kimi-K3(91.2%) comparably but still behind Claude Opus 5 (90.8%) is close; overall still behind top closed frontier models on several agentic evals (e.g. AutomationBench v1.0.6 at 50.2% vs GLM-5.3's 48.2% — competitive but not leading)
  • Undisclosed active-parameter count (backbone is 1.6T total, active portion not stated in the model card)

Best for: teams wanting Nex-AGI's strongest open-weight text-agent model for coding/tool-use workloads where multimodal input isn't required.

What Nex N2.5 Max Is

Nex-N2.5-Max is the largest model in Nex-AGI's second-generation N2.5 release, built on a 1.6-trillion-parameter, text-only Mixture-of-Experts foundation model. Nex-AGI describes it as their first complete post-training effort carried out at full trillion-parameter scale, following broader expansion of agent training environments, task types, and productivity scenarios versus the prior N2 generation.

Unlike its siblings Nex-N2.5-Pro and Nex-N2.5-mini — which continue the multimodal, computer-use-focused line — Max is text-only, positioned as the flagship for long-horizon coding and tool-use agent tasks rather than visual grounding. It supports the same reasoning_effort control (none/medium/high) as the rest of the family and uses a deepseek-r1-style reasoning parser in its reference SGLang deployment (contrasted with the qwen3 parser used for Pro/mini), suggesting a distinct post-training reasoning-trace format at this scale.

Specifications

FieldValue
OrganizationNex-AGI
Parameters1.6T total (MoE); active-parameter count not disclosed
Context windowNot disclosed in model card
ArchitectureText-only Mixture-of-Experts; NexAU harness for coding evaluation
LicenseOpen weights (Hugging Face + ModelScope)
Release dateSeptember 2026
ModalityText only

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Open weights——

Open-weight model; self-hosting cost only. Multi-node (2×8×H200) deployment recipe published by Nex-AGI; no hosted-API listing found at time of writing.

Public Benchmark Scores

Scores above are reported by Nex-AGI and shown for context. They are not Benchgen measurements. Coding tasks use Nex-AGI's NexAU harness; sampling uses temperature=0.7, top_p=0.95, top_k=40. Nex-N2.5-Max is text-only and is not evaluated on the multimodal benchmark table (OSWorld, WebArena, Vision2Web, etc.) published for Nex-N2.5-Pro/mini.

Nex N2.5 Max vs Alternatives

ModelTerminal-Bench 2.1SWE-Bench ProJob BenchPrice (in/out per 1M)
Nex N2.5 Max86.1%65.7%53.6%Open weights
Claude Opus 589.1%79.2%65.7%—
Nex N2.5 Pro82.7%61.2%41.4%Open weights
DeepSeek V4 Pro87.9%67.7%54.1%$0.435/$—

Nex-N2.5-Max closes much of the gap to closed frontier models (Claude Opus 5, GPT-5.6 Sol) on agentic coding benchmarks versus its own smaller Pro sibling, but still trails Claude Opus 5 and DeepSeek V4 Pro on SWE-Bench Pro and Job Bench specifically — the trade-off is open weights and self-hostability rather than category-leading raw scores.

How Nex N2.5 Max Performs on Real Agent Tasks

As Nex-AGI's first full trillion-parameter-scale post-training run, Nex-N2.5-Max's benchmark profile shows broad, consistent gains over Nex-N2.5-Pro across coding and tool-use evals (Terminal-Bench 2.1, SWE-Bench Pro, DeepSWE v1.1, Toolathlon Verified, Job Bench) rather than a narrow specialization — evidence the scale-up primarily improved general agentic reliability rather than one task category. Because it's text-only, teams needing computer-use/vision-grounded agent behavior should evaluate Nex-N2.5-Pro instead; Max is the better fit for pure coding/tool-orchestration agents at self-hosted scale.

Use Nex N2.5 Max via API

from openai import OpenAI  # self-host via the published SGLang multi-node recipe, or check for OpenRouter/provider listings

client = OpenAI(api_key="YOUR_API_KEY", base_url="http://localhost:8000/v1")

response = client.chat.completions.create(
    model="nex-agi/nex-n2.5-max",
    messages=[{"role": "user", "content": "Refactor this repository's test suite..."}],
)
print(response.choices[0].message.content)

Frequently Asked Questions

What is Nex N2.5 Max? Nex-N2.5-Max is Nex-AGI's flagship model — a 1.6-trillion-parameter, text-only Mixture-of-Experts model and the lab's first complete post-training run at trillion-parameter scale.
What is Nex N2.5 Max's context window? Not disclosed in the official model card at time of writing.
How much does Nex N2.5 Max cost? It is open weights (free to self-host); no confirmed hosted-API pricing was found at time of writing.
Is Nex N2.5 Max open source? Yes — released as open weights on Hugging Face and ModelScope.
What is Nex N2.5 Max's knowledge cutoff? Not disclosed in the official model card.

Specs and scores sourced from Nex-AGI's official Hugging Face model card. Last updated 2026-09-15.