Benchgen
Models/nex-agi/

Nex-N2.5-Pro

DraftPublic

Model Details

Nex N2.5 Pro

Organization License Modality Released

Quick answer: Nex-N2.5-Pro is Nex-AGI's mid-tier open-weight agentic model in the Nex-N2.5 family (mini/Pro/Max), built for long-horizon computer-use, web-browsing, and coding agent tasks. It scores 82.7% on Terminal-Bench 2.1, 61.2% on SWE-Bench Pro, and 82.2% on OSWorld-Verified, and is available both as open weights on Hugging Face and hosted via OpenRouter.

At a Glance

Where Nex N2.5 Pro leads

  • 82.2% OSWorld-Verified and 56.4% OSWorld-2 — strong computer-use grounding for an open-weight model
  • 87.4% OSWorld-G — leads several closed frontier rivals shown on its own card (e.g. Claude Opus 5 at 76.8%)
  • 68.2% Vision2Web — competitive web-GUI-agent grounding
  • Open weights + hosted OpenRouter access — flexible deployment

Where it lags

  • 61.2% SWE-Bench Pro — behind Claude Opus 5 (79.2%) and DeepSeek-V4-Pro-0813 (67.7%)
  • 44.2% AutomationBench v1.0.6 — trails most frontier rivals on its own comparison table
  • 38.2% SWE-MM — weakest shown among compared models on multimodal SWE tasks

Best for: teams wanting an open-weight, self-hostable agent model with strong computer-use/web-browsing grounding rather than best-in-class raw coding accuracy.

What Nex N2.5 Pro Is

Nex-N2.5-Pro is the mid-size model in Nex-AGI's second-generation N2.5 agentic family, sitting between Nex-N2.5-mini and the 1.6T-parameter Nex-N2.5-Max. It builds on the multimodal foundations of the earlier Nex-N2 line, with targeted improvements in computer use, web browsing, and visually-grounded agentic capability — treating vision not just as an input modality but as the feedback channel an agent uses to verify its own actions in a browser or desktop environment.

Coding evaluations use Nex-AGI's own NexAU harness; computer-use and browser-use benchmarks (OSWorld, WebTest, WebArena) run through a NexCUA harness with grounding coordinates normalized to a 0–1000 scale (soon to be open-sourced). The model supports a reasoning_effort parameter (none/medium/high) to trade inference cost for deliberation, and ships with SGLang deployment recipes for single-node 8×H100 serving.

Specifications

FieldValue
OrganizationNex-AGI
ParametersUndisclosed (mid-tier of the N2.5 family; Max is 1.6T MoE)
Context windowNot disclosed in model card
ArchitectureMultimodal agentic transformer; NexAU (coding) / NexCUA (computer-use) harnesses
LicenseOpen weights (Hugging Face + ModelScope)
Release dateSeptember 2026
ModalityMultimodal (computer-use / web-browsing / vision-grounded agent)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Open weights——
OpenRouterVaries by provider listingVaries by provider listing

Open-weight model; self-hosting cost only, or use via OpenRouter at provider-set rates.

Public Benchmark Scores

Scores above are reported by Nex-AGI and shown for context. They are not Benchgen measurements. Coding tasks use the NexAU harness; computer-use/browser tasks use the NexCUA harness. WebTest, WebArena-Verified, OSWorld-G, Vision2Web, and SWE-MM scores are documented in the accompanying bucket-B benchmark configs added alongside this model rather than repeated here to avoid duplicate unsourced rows before those pages exist; see each benchmark's own leaderboard entry for this model's score.

Nex N2.5 Pro vs Alternatives

ModelTerminal-Bench 2.1SWE-Bench ProOSWorld-VerifiedPrice (in/out per 1M)
Nex N2.5 Pro82.7%61.2%82.2%Open weights
Claude Opus 589.1%79.2%83.4%—
DeepSeek V4 Pro87.9%67.7%*—$0.435/$—

*DeepSeek-V4-Pro-0813 score as reported on the Nex-N2.5-Pro card; see DeepSeek's own card for a possibly different harness/date.

Nex-N2.5-Pro trades top-tier coding accuracy for stronger computer-use and browsing grounding at open-weight cost — a fit for teams building browser/desktop agents who want to self-host rather than pay closed-API prices, at the cost of some ground on pure SWE-bench-style coding tasks versus Claude Opus 5 or DeepSeek V4 Pro.

How Nex N2.5 Pro Performs on Real Agent Tasks

Nex-N2.5-Pro's benchmark spread signals a model tuned specifically for the "verify with vision" loop of computer-use agents: it leads or nearly matches frontier closed models on OSWorld-G (87.4%, ahead of Claude Opus 5's 76.8%) and stays competitive on OSWorld-Verified, while trailing further on text-only coding benchmarks like SWE-Bench Pro and AutomationBench. That pattern is consistent with a lab explicitly optimizing for browser/desktop agent reliability over raw code-generation accuracy — useful context for anyone picking a model based on a single averaged score rather than the task-specific breakdown.

Use Nex N2.5 Pro via API

from openai import OpenAI  # OpenRouter exposes an OpenAI-compatible API

client = OpenAI(api_key="YOUR_OPENROUTER_KEY", base_url="https://openrouter.ai/api/v1")

response = client.chat.completions.create(
    model="nex-agi/nex-n2.5-pro",
    messages=[{"role": "user", "content": "Book a flight using this browser session..."}],
)
print(response.choices[0].message.content)

Frequently Asked Questions

What is Nex N2.5 Pro? Nex-N2.5-Pro is an open-weight multimodal agentic model from Nex-AGI, the mid-tier model in the Nex-N2.5 family, focused on computer-use, web-browsing, and coding agent tasks.
What is Nex N2.5 Pro's context window? Not disclosed in the official model card at time of writing.
How much does Nex N2.5 Pro cost? It is open weights (free to self-host); it's also available hosted via OpenRouter at provider-set per-token rates.
Is Nex N2.5 Pro open source? Yes — released as open weights on Hugging Face and ModelScope.
What is Nex N2.5 Pro's knowledge cutoff? Not disclosed in the official model card.

Specs and scores sourced from Nex-AGI's official Hugging Face model card. Last updated 2026-09-15.