Quick answer: GPT-6 Astra Ultrafast is OpenAI's low-latency inference tier of GPT-6 Astra, generating output at roughly 300 tokens/second — about 6-8x faster than Standard GPT-6 Astra — at 6x the Standard price ($60 input / $300 output per 1M tokens). It shares Astra's underlying intelligence profile and benchmark scores; the difference is purely inference speed and cost, not capability.
What it is: A speed-optimized serving tier of GPT-6 Astra, aimed at latency-sensitive agentic and interactive workloads where raw throughput matters more than per-token cost.
Why it matters: Lets developers trade cost for roughly an order-of-magnitude latency reduction on the same underlying model, without retraining or switching to a smaller/weaker model.
Known limitations: At 6x Standard pricing, it's materially more expensive per token; OpenAI has not published a separate benchmark suite for the Ultrafast tier since it uses the same weights as Standard GPT-6 Astra.
GPT-6 Astra Ultrafast is a dedicated low-latency serving configuration of OpenAI's GPT-6 Astra model (released September 3, 2026). Rather than a distinct model checkpoint, it is an inference-infrastructure tier optimized for generation speed — reported at approximately 300 tokens/second, a 6-8x improvement over Astra's Standard serving tier — intended for latency-sensitive use cases such as real-time coding agents, interactive computer-use workflows, and customer-facing applications where response time is a hard constraint.
Because it runs the same underlying weights as Standard GPT-6 Astra, Ultrafast inherits Astra's full benchmark profile (FrontierMath, GPQA Diamond, ARC-AGI-3, OSWorld 2.0, ExploitBench, and the rest of Astra's published results) — see the GPT-6 Astra page for those scores. OpenAI prices the speed tier at 6x Standard: $60 per million input tokens and $300 per million output tokens, versus Astra Standard's $10/$50.
Availability mirrors Astra's rollout: the OpenAI API (as a distinct Ultrafast model name), plus ChatGPT Work, Codex for Pro/500, and Enterprise tiers where OpenAI has enabled the faster serving path.
| Field | Value |
|---|---|
| Organization | OpenAI |
| Base model | GPT-6 Astra (same weights) |
| License | Proprietary (API only) |
| Modality | Multimodal (text + vision) |
| Release date | 2026-09 (Ultrafast tier) |
| Generation speed | ~300 tokens/second (6-8x Standard Astra) |
| Availability | OpenAI API, ChatGPT Work, Codex (Pro/500), Enterprise |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| GPT-6 Astra Standard | $10.00 | $50.00 |
| GPT-6 Astra Ultrafast | $60.00 | $300.00 |
Ultrafast pricing is 6x Standard GPT-6 Astra pricing, reflecting the dedicated low-latency serving infrastructure rather than a different model.
GPT-6 Astra Ultrafast runs the same underlying weights as Standard GPT-6 Astra, so its benchmark profile is identical — see GPT-6 Astra's benchmark table for full results (FrontierMath 97.6%, ARC-AGI-3 99.9%, OSWorld 2.0 72.6%, ExploitBench 100.0%, and more). OpenAI has not published a separate benchmark suite specific to the Ultrafast serving tier.