Quick answer: Gemini 4 Argon is Google's frontier agentic coding and reasoning model, announced September 30, 2026 in a limited release through Google's Fairwind Program rather than broad general availability. It scores 77.9% on DeepSWE, 57.4% on Terminal-Bench 4.0, and 51.3% on AutomationBench, with a 1M output token context window. Google's own comparison table shows Claude Opus 5.5 ahead on several agentic coding benchmarks (e.g. Terminal-Bench 4.0: 66.4% vs. Argon's 57.4%), while GPT-6 Astra leads on FrontierSWE v2 and Terminal-Bench Science 0.1.
Where it leads: Strong agentic software-engineering performance (DeepSWE 77.9%, Finance Agent v2 65.4%) and a 1M output token context window, among the largest disclosed for a frontier model.
Where it lags: Trails Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs. 66.4%) and trails GPT-6 Astra by roughly 10.5 points on both FrontierSWE v2 and Terminal-Bench Science 0.1, per Google's own comparison table.
Availability: Limited release via Google's Fairwind Program at launch — not yet broadly available through the standard Gemini API or consumer surfaces.
Gemini 4 Argon is Google's latest frontier model, positioned around agentic coding and long-horizon reasoning tasks. Announced September 30, 2026, it ships initially through Google's Fairwind Program — a limited, invitation-based access tier — rather than immediate general availability, suggesting Google is still gathering real-world deployment feedback before a wider rollout.
Google's own announcement includes a multi-model comparison table citing scores for Claude Opus 5.5 and GPT-6 Astra alongside Argon's own results. On that table, Argon's largest disclosed strengths are in real-world software-engineering agent tasks (DeepSWE, Finance Agent v2, Terminal-Bench Science 0.1), while Claude Opus 5.5 holds a clear lead on Terminal-Bench 4.0 and GPT-6 Astra leads on FrontierSWE v2 and Terminal-Bench Science 0.1 by about 10.5 percentage points in both cases.
A standout spec is Argon's 1M output token context window, among the largest disclosed by any frontier lab to date, aimed at long-horizon agentic workflows that generate large amounts of code or text in a single pass.
| Field | Value |
|---|---|
| Organization | |
| License | Proprietary (API only) |
| Modality | Multimodal (text + vision) |
| Release date | 2026-09-30 |
| Output context window | 1,000,000 tokens |
| Availability | Limited — Google Fairwind Program |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.4% | Google: Gemini 4 Argon | 2026-09 |
| Terminal-Bench Science 0.1 | 57.6% | Google: Gemini 4 Argon | 2026-09 |
| AutomationBench | 51.3% | Google: Gemini 4 Argon | 2026-09 |
| DeepSWE (v1.1) | 77.9% | Google: Gemini 4 Argon | 2026-09 |
| Finance Agent v2 | 65.4% | Google: Gemini 4 Argon | 2026-09 |
| PostTrainBench | 45.3% | Google: Gemini 4 Argon | 2026-09 |
Scores above are reported by Google's own Gemini 4 Argon announcement and shown for context; not independent Benchgen measurements. The same announcement reports Claude Opus 5.5 scoring 66.4% on Terminal-Bench 4.0 and GPT-6 Astra scoring 65.5% on FrontierSWE v2 and 68.1% on Terminal-Bench Science 0.1 — both ahead of Argon on those specific benchmarks; Harvey Legal Agent Benchmark and FrontierSWE v2 figures from the same table were not added here due to scale/version mismatches with Benchgen's existing pages for those benchmarks.