Benchgen
Models/ornith-deepreinforce/

Ornith-1.5-35B-A3B

DraftPublic

Model Details

Ornith-1.5-35B-A3B

Organization Params License Modality Released

Quick answer: Ornith-1.5-35B-A3B is the mid-size mixture-of-experts member of Ornith's (DeepReinforce's) Ornith-1.5 family — a 35B-total, ~3B-activated MoE model released under MIT license. It scores 79% on SWE-bench Verified, 67.8% on Terminal Bench 2.1, and 89.2% on GPQA Diamond, significantly outperforming its similarly-sized peer Qwen 3.6-35B and dense models like Gemma 4-31B and Meta's Muse Glimmer-30B despite activating only ~3B parameters per token.

At a Glance

Where Ornith-1.5-35B-A3B leads

  • 79% SWE-bench Verified, 68.5% Terminal Bench 2.1 (Claude Code harness) — outperforms Gemma 4-31B (52.0% SWE-bench Verified) and Muse Glimmer-30B (76.0%) by wide margins on agentic coding despite activating only ~3B parameters
  • 89.2% GPQA Diamond — strong graduate-level science reasoning for its size class
  • 72.5% ClawEval, 67.6% BrowseComp — solid agentic tool-use and web-browsing performance
  • MIT license — fully open weights, permissive for commercial use
  • Efficient MoE design (~3B active) makes it deployable on 2× 80GB GPUs

Where it lags

  • 22% DeepSWE — noticeably weaker than its flagship 397B sibling (56%) on this harder coding benchmark
  • 33.4% HLE (with tools) / 25.6% (no tools) — modest on frontier-exam reasoning at this scale
  • 5.1% Frontier-Bench v0.1 — reflects the benchmark's extreme difficulty across all but flagship-tier models

Best for: cost-efficient self-hosted coding agents and tool-use workflows where MoE efficiency (only ~3B active parameters) needs to significantly outperform similarly-sized dense models.

What Ornith-1.5-35B-A3B Is

Ornith-1.5-35B-A3B is the mid-size mixture-of-experts model in Ornith's (built by the DeepReinforce team) Ornith-1.5 family, sitting between the 9B dense edge model and the 397B flagship. Like its siblings, it inherits Ornith-1.5's self-improvement training loop — the model jointly generates its own training tasks, task-specific scaffolds, and solution rollouts, optimized end-to-end via reinforcement learning (GRPO), rather than relying on a fixed human-curated training distribution.

Despite activating only ~3B parameters per token, Ornith-1.5-35B-A3B "significantly outperforms its similar-sized peer Qwen 3.6-35B across all coding and agentic benchmarks," per the official model card, and beats larger dense models — Gemma 4-31B and Meta's Muse Glimmer-30B — by wide margins on agentic coding (68.5% vs. 43.4% and 51.7% on Terminal Bench 2.1's Claude Code harness; 79.0% vs. 52.0% and 76.0% on SWE-bench Verified). It supports a 262,144-token context window, extensible to roughly 1,000,000 tokens via YaRN RoPE scaling, and serves comfortably on 2× 80GB GPUs.

Specs

FieldValue
OrganizationOrnith (DeepReinforce)
ArchitectureMixture-of-Experts (Qwen3.5 MoE foundation)
Total parameters35B (36B per HF card)
Activated parameters~3B
Context window262,144 tokens (extensible to ~1M via YaRN)
ModalityText (reasoning model — <think> traces by default)
LicenseMIT
Release dateAugust 2026

Pricing

Ornith-1.5-35B-A3B is released as open weights (self-hosted); Ornith/DeepReinforce has not announced a hosted API pricing tier for this model.

Public Benchmark Scores

BenchmarkScoreSourceDate
Terminal Bench 2.167.8%Ornith-1.5-35B-A3B model card (HF Eval Results, Terminus-2 harness)2026-08
SWE-bench Verified79%Ornith-1.5-35B-A3B model card (HF Eval Results)2026-08
SWE-bench Pro59.6%Ornith-1.5-35B-A3B model card (HF Eval Results)2026-08
DeepSWE22%Ornith-1.5-35B-A3B model card (HF Eval Results)2026-08
GPQA Diamond89.2%Ornith-1.5-35B-A3B model card (HF Eval Results)2026-08
Humanity's Last Exam33.4%Ornith-1.5-35B-A3B model card (HF Eval Results, with tools)2026-08
MCP-Atlas70.2%Ornith-1.5-35B-A3B model card (benchmark appendix)2026-08
Toolathlon-Verified48.7%Ornith-1.5-35B-A3B model card (benchmark appendix)2026-08
BrowseComp67.6%Ornith-1.5-35B-A3B model card (benchmark appendix)2026-08

Scores are self-reported by the Ornith team on the official Hugging Face model card (ornith-ai/Ornith-1.5-35B-A3B) and the accompanying technical blog post; the top 6 rows are additionally confirmed via the card's structured HF "Evaluation Results" metadata. All results are averaged over 5 independent runs. HLE score uses the "with tools" evaluation condition (25.6% without tools).

Notable Results Not Yet Tracked as Benchgen Benchmarks

  • SWE-bench Multilingual (71.4%), Frontier-Bench v0.1 (5.1%), NL2Repo (46.2%), SWE Atlas – QnA (39.8%), WideSearch (67.8%), and ClawEval (72.5%) are all real, methodology-documented benchmarks per the model card's footnotes, but don't have existing Benchgen benchmark pages yet — flagged as follow-ups (same set flagged on the sibling Ornith-1.5-397B page).

Ornith-1.5-35B-A3B vs Alternatives

ModelTerminal Bench 2.1 (Claude Code)SWE-bench VerifiedActive Params
Ornith-1.5-35B-A3B68.5%79.0%~3B
Gemma 4-31B43.4%52.0%31B (dense)
Muse Glimmer-30B51.7%76.0%30B (dense)

Despite activating far fewer parameters than either dense competitor, Ornith-1.5-35B-A3B outperforms both on agentic coding benchmarks — a notable efficiency result for its MoE design.

FAQ

Is Ornith-1.5-35B-A3B open source? Yes — MIT license, with weights available on Hugging Face (ornith-ai/Ornith-1.5-35B-A3B), plus FP8, NVFP4, and GGUF quantized variants.

Who makes Ornith-1.5-35B-A3B? Ornith, built by the DeepReinforce team.

How big is Ornith-1.5-35B-A3B? 35B total parameters (Mixture-of-Experts), with only ~3B activated per forward pass.

What is Ornith-1.5-35B-A3B's context window? 262,144 tokens natively, extensible to roughly 1,000,000 tokens via YaRN RoPE scaling.

How does Ornith-1.5-35B-A3B compare to similarly-sized models? It significantly outperforms Qwen 3.6-35B across coding and agentic benchmarks, and beats larger dense models (Gemma 4-31B, Muse Glimmer-30B) on agentic coding despite activating far fewer parameters per token.

Where can I access Ornith-1.5-35B-A3B? Self-hosted from the open weights at huggingface.co/ornith-ai/Ornith-1.5-35B-A3B (BF16, FP8, or NVFP4), via vLLM or SGLang on 2× 80GB GPUs.


Benchmark scores sourced from the official Hugging Face model card for ornith-ai/Ornith-1.5-35B-A3B and the Ornith team's technical blog (Aug 2026), cross-checked against the card's structured HF Eval Results metadata where available.