Benchgen
Models/openai/

GPT-6 Astra Ultrafast

DraftPublic

Model Details

GPT-6 Astra Ultrafast

Organization Pricing License Modality Speed

Quick answer: GPT-6 Astra Ultrafast is OpenAI's low-latency inference tier of GPT-6 Astra, generating output at roughly 300 tokens/second — about 6-8x faster than Standard GPT-6 Astra — at 6x the Standard price ($60 input / $300 output per 1M tokens). It shares Astra's underlying intelligence profile and benchmark scores; the difference is purely inference speed and cost, not capability.

At a Glance

What it is: A speed-optimized serving tier of GPT-6 Astra, aimed at latency-sensitive agentic and interactive workloads where raw throughput matters more than per-token cost.

Why it matters: Lets developers trade cost for roughly an order-of-magnitude latency reduction on the same underlying model, without retraining or switching to a smaller/weaker model.

Known limitations: At 6x Standard pricing, it's materially more expensive per token; OpenAI has not published a separate benchmark suite for the Ultrafast tier since it uses the same weights as Standard GPT-6 Astra.

What GPT-6 Astra Ultrafast Is

GPT-6 Astra Ultrafast is a dedicated low-latency serving configuration of OpenAI's GPT-6 Astra model (released September 3, 2026). Rather than a distinct model checkpoint, it is an inference-infrastructure tier optimized for generation speed — reported at approximately 300 tokens/second, a 6-8x improvement over Astra's Standard serving tier — intended for latency-sensitive use cases such as real-time coding agents, interactive computer-use workflows, and customer-facing applications where response time is a hard constraint.

Because it runs the same underlying weights as Standard GPT-6 Astra, Ultrafast inherits Astra's full benchmark profile (FrontierMath, GPQA Diamond, ARC-AGI-3, OSWorld 2.0, ExploitBench, and the rest of Astra's published results) — see the GPT-6 Astra page for those scores. OpenAI prices the speed tier at 6x Standard: $60 per million input tokens and $300 per million output tokens, versus Astra Standard's $10/$50.

Availability mirrors Astra's rollout: the OpenAI API (as a distinct Ultrafast model name), plus ChatGPT Work, Codex for Pro/500, and Enterprise tiers where OpenAI has enabled the faster serving path.

Specifications

FieldValue
OrganizationOpenAI
Base modelGPT-6 Astra (same weights)
LicenseProprietary (API only)
ModalityMultimodal (text + vision)
Release date2026-09 (Ultrafast tier)
Generation speed~300 tokens/second (6-8x Standard Astra)
AvailabilityOpenAI API, ChatGPT Work, Codex (Pro/500), Enterprise

Pricing

Input (per 1M tokens)Output (per 1M tokens)
GPT-6 Astra Standard$10.00$50.00
GPT-6 Astra Ultrafast$60.00$300.00

Ultrafast pricing is 6x Standard GPT-6 Astra pricing, reflecting the dedicated low-latency serving infrastructure rather than a different model.

Public Benchmark Scores

GPT-6 Astra Ultrafast runs the same underlying weights as Standard GPT-6 Astra, so its benchmark profile is identical — see GPT-6 Astra's benchmark table for full results (FrontierMath 97.6%, ARC-AGI-3 99.9%, OSWorld 2.0 72.6%, ExploitBench 100.0%, and more). OpenAI has not published a separate benchmark suite specific to the Ultrafast serving tier.

Frequently Asked Questions

What is GPT-6 Astra Ultrafast? GPT-6 Astra Ultrafast is a low-latency inference tier of OpenAI's GPT-6 Astra model, generating output at roughly 300 tokens/second — about 6-8x faster than Standard Astra — at 6x the price.
How much does GPT-6 Astra Ultrafast cost? $60 per million input tokens and $300 per million output tokens — 6x Standard GPT-6 Astra's $10/$50 pricing.
Does GPT-6 Astra Ultrafast score differently on benchmarks than Standard Astra? No — Ultrafast runs the same underlying model weights as Standard GPT-6 Astra, so published benchmark scores are identical. The difference is purely inference speed and cost.
Is GPT-6 Astra Ultrafast available via API? Yes, via the OpenAI API as a distinct model tier, as well as ChatGPT Work, Codex (Pro/500), and Enterprise where OpenAI has enabled the faster serving path.