Quick answer: GPT-5 Codex is OpenAI's September 2025 coding-specialised model, scoring 74.5% on SWE-Bench Verified. It represents OpenAI's revival of the "Codex" brand for code-focused models within the GPT-5 generation, targeting real-world software engineering tasks.
Where GPT-5 Codex leads
Where it lags
Best for: Software engineering agentic tasks using the Codex interface; production coding pipelines; automated code review and PR fixing.
GPT-5 Codex (released September 2025) is OpenAI's code-specialised model in the GPT-5 generation — a revival of the Codex brand last seen in the original Codex/Copilot era. It is distinct from the general-purpose GPT-5 models, with training and fine-tuning specifically targeting software engineering.
The 74.5% SWE-Bench Verified score places it competitively with DeepSeek-V3.2 Speciale (73.1%) and just above the threshold that makes models practical for real-world repository-level coding tasks. GPT-5.1 Codex (February 2026) provides a modest improvement to 73.7% — suggesting the generation gap between Codex versions was relatively small.
| Field | Value |
|---|---|
| Organization | OpenAI |
| License | Proprietary (API only) |
| Release date | September 2025 |
| Knowledge cutoff | September 2024 |
| Modality | Text / Code |
| Focus | Software engineering (SWE-Bench) |
Available via OpenAI API — refer to the OpenAI pricing page for current Codex model rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| SWE-Bench Verified | 74.5% | Benchgen evaluation | 2025-09 |
| Model | SWE-Bench Verified | License |
|---|---|---|
| GPT-5 Codex | 74.5% | Proprietary |
| GPT-5.1 Codex | 73.7% | Proprietary |
| DeepSeek-V3.2 Speciale | 73.1% | MIT |
| DeepSeek-V4 Flash Max | 79% | Proprietary |
GPT-5 Codex vs DeepSeek-V3.2 Speciale: +1.4pp SWE-Bench, proprietary vs MIT. For open-weight SWE, V3.2 Speciale is preferred. For maximum SWE-Bench: DeepSeek-V4 Flash Max (79%).
Specs from OpenAI's GPT-5 Codex release (September 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.