Quick answer: Nex-N2.5-Pro is Nex-AGI's mid-tier open-weight agentic model in the Nex-N2.5 family (mini/Pro/Max), built for long-horizon computer-use, web-browsing, and coding agent tasks. It scores 82.7% on Terminal-Bench 2.1, 61.2% on SWE-Bench Pro, and 82.2% on OSWorld-Verified, and is available both as open weights on Hugging Face and hosted via OpenRouter.
Where Nex N2.5 Pro leads
Where it lags
Best for: teams wanting an open-weight, self-hostable agent model with strong computer-use/web-browsing grounding rather than best-in-class raw coding accuracy.
Nex-N2.5-Pro is the mid-size model in Nex-AGI's second-generation N2.5 agentic family, sitting between Nex-N2.5-mini and the 1.6T-parameter Nex-N2.5-Max. It builds on the multimodal foundations of the earlier Nex-N2 line, with targeted improvements in computer use, web browsing, and visually-grounded agentic capability — treating vision not just as an input modality but as the feedback channel an agent uses to verify its own actions in a browser or desktop environment.
Coding evaluations use Nex-AGI's own NexAU harness; computer-use and browser-use benchmarks (OSWorld, WebTest, WebArena) run through a NexCUA harness with grounding coordinates normalized to a 0–1000 scale (soon to be open-sourced). The model supports a reasoning_effort parameter (none/medium/high) to trade inference cost for deliberation, and ships with SGLang deployment recipes for single-node 8×H100 serving.
| Field | Value |
|---|---|
| Organization | Nex-AGI |
| Parameters | Undisclosed (mid-tier of the N2.5 family; Max is 1.6T MoE) |
| Context window | Not disclosed in model card |
| Architecture | Multimodal agentic transformer; NexAU (coding) / NexCUA (computer-use) harnesses |
| License | Open weights (Hugging Face + ModelScope) |
| Release date | September 2026 |
| Modality | Multimodal (computer-use / web-browsing / vision-grounded agent) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Open weights | — | — |
| OpenRouter | Varies by provider listing | Varies by provider listing |
Open-weight model; self-hosting cost only, or use via OpenRouter at provider-set rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Terminal-Bench 2.1 | 82.7% | Nex-N2.5-Pro model card | 2026-09 |
| SWE-Bench Pro | 61.2% | Nex-N2.5-Pro model card | 2026-09 |
| DeepSWE | 55.8% (v1.1) | Nex-N2.5-Pro model card | 2026-09 |
| AutomationBench | 44.2% (v1.0.6) | Nex-N2.5-Pro model card | 2026-09 |
| Toolathlon Verified | 68.5% | Nex-N2.5-Pro model card | 2026-09 |
| GDPval-AA v2 | 1628 | Nex-N2.5-Pro model card | 2026-09 |
| Job Bench | 41.4% | Nex-N2.5-Pro model card | 2026-09 |
| BrowseComp | 89.7% | Nex-N2.5-Pro model card | 2026-09 |
| OSWorld-Verified | 82.2% | Nex-N2.5-Pro model card | 2026-09 |
| OSWorld-2 | 56.4% | Nex-N2.5-Pro model card | 2026-09 |
| OmniDocBench | 92.2% | Nex-N2.5-Pro model card | 2026-09 |
Scores above are reported by Nex-AGI and shown for context. They are not Benchgen measurements. Coding tasks use the NexAU harness; computer-use/browser tasks use the NexCUA harness. WebTest, WebArena-Verified, OSWorld-G, Vision2Web, and SWE-MM scores are documented in the accompanying bucket-B benchmark configs added alongside this model rather than repeated here to avoid duplicate unsourced rows before those pages exist; see each benchmark's own leaderboard entry for this model's score.
| Model | Terminal-Bench 2.1 | SWE-Bench Pro | OSWorld-Verified | Price (in/out per 1M) |
|---|---|---|---|---|
| Nex N2.5 Pro | 82.7% | 61.2% | 82.2% | Open weights |
| Claude Opus 5 | 89.1% | 79.2% | 83.4% | — |
| DeepSeek V4 Pro | 87.9% | 67.7%* | — | $0.435/$— |
*DeepSeek-V4-Pro-0813 score as reported on the Nex-N2.5-Pro card; see DeepSeek's own card for a possibly different harness/date.
Nex-N2.5-Pro trades top-tier coding accuracy for stronger computer-use and browsing grounding at open-weight cost — a fit for teams building browser/desktop agents who want to self-host rather than pay closed-API prices, at the cost of some ground on pure SWE-bench-style coding tasks versus Claude Opus 5 or DeepSeek V4 Pro.
Nex-N2.5-Pro's benchmark spread signals a model tuned specifically for the "verify with vision" loop of computer-use agents: it leads or nearly matches frontier closed models on OSWorld-G (87.4%, ahead of Claude Opus 5's 76.8%) and stays competitive on OSWorld-Verified, while trailing further on text-only coding benchmarks like SWE-Bench Pro and AutomationBench. That pattern is consistent with a lab explicitly optimizing for browser/desktop agent reliability over raw code-generation accuracy — useful context for anyone picking a model based on a single averaged score rather than the task-specific breakdown.
from openai import OpenAI # OpenRouter exposes an OpenAI-compatible API
client = OpenAI(api_key="YOUR_OPENROUTER_KEY", base_url="https://openrouter.ai/api/v1")
response = client.chat.completions.create(
model="nex-agi/nex-n2.5-pro",
messages=[{"role": "user", "content": "Book a flight using this browser session..."}],
)
print(response.choices[0].message.content)Specs and scores sourced from Nex-AGI's official Hugging Face model card. Last updated 2026-09-15.