
What Shipped on September 29
Ultrafast is OpenAI's premium speed tier for workloads where speed matters most. Per the DevDay 2026 recap, GPT-6 Astra Ultrafast became available on September 29, 2026 in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans, while GPT-6.1 Sol Ultrafast is coming soon.
Alongside it, OpenAI introduced Pro 500, described as offering “our highest usage allowance at 25 times the ChatGPT Plus allowance” and including access to Ultrafast, and reopened Pro 200 subscriptions with continued access to frontier models including GPT-6.1 Sol.
This is Ultrafast. Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API.
— @OpenAI September 29, 2026

What Ultrafast Actually Measures
OpenAI's product documentation is unusually precise about the metric, and the precision matters: Ultrafast “generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex.” OpenAI then adds the caveat directly — this comparison measures token generation speed, not billing rates or overall task completion time.
So the honest reading of “up to 8x faster” is: up to 8x on token generation throughput, up to 300 tokens per second in Codex, and up to 6x in the API. It is not a claim that a task finishes 8x sooner. Agentic work spends time on tool calls, filesystem operations, network round trips, and model reasoning as well as token emission. A tier that doubles output speed can change wall-clock time for a long-generation task far more than for a task dominated by tool latency.
That is also why OpenAI recommends WebSockets for this workload. Its API documentation states that without a persistent connection, network overhead can reduce the latency gains, and strongly recommends WebSockets especially for agentic applications that make many tool calls in quick succession.
Pro 100, Pro 200 and Pro 500
OpenAI's Pro tiers help article now lists three monthly Pro plans. Prices and Ultrafast inclusion are stated in a table there.
| Plan | Monthly price | Ultrafast | Notes |
|---|---|---|---|
| Pro 100 | $100/month | Not included | Baseline Pro tier. Buying credits on Pro 100 does not unlock Ultrafast at launch. |
| Pro 200 | $200/month | Not included | More included usage than Pro 100. Reopened to new subscriptions; new non-grandfathered subscriptions include a lower usage allowance. |
| Pro 500 | $500/month | Included | Highest included usage of the three Pro plans and the only Pro tier with Ultrafast. Uses included usage first, then credits. |
Prices as published in OpenAI's “About ChatGPT Pro tiers” help article on September 29, 2026. Check current pricing before subscribing; plan structures change.
Two structural details are easy to miss. First, Pro 500 is the only Pro tier that includes Ultrafast, buying credits on Pro 100 or Pro 200 does not unlock Ultrafast at launch. Second, on Pro 500 Ultrafast uses your included usage first and then draws from your credit balance once that allowance is used. Included usage is consumed before credits on other eligible features too.

Ultrafast is available today for GPT-6 Astra in Codex, ChatGPT Work, and the API, with GPT-6.1 Sol coming soon. To access it in Codex and ChatGPT Work, we're introducing Pro 500—a new plan with our highest usage limits.
— @OpenAI September 29, 2026
Billing Multipliers, Not Speed
This is the single most misread part of the launch. OpenAI publishes two different numbers per tier, a speed figure and a billing multiplier, and the documentation says outright that the billing multipliers do not describe speed increases.
| Tier | Speed | Billed against included subscription limits | Purchased credits / Enterprise pay-as-you-go |
|---|---|---|---|
| Fast mode | 1.5x for GPT-5.6 and GPT-5.5; model-dependent for other supported models | 2.5x the Standard rate | 2x the Standard rate |
| Ultrafast (GPT-6 Astra) | Up to 8x faster token generation than Standard in Codex | 8x the Standard rate | 6x the Standard rate |
So for GPT-6 Astra, Ultrafast consumes included subscription limits at 8x the Standard rate, and purchased credits or Enterprise pay-as-you-go usage are billed at 6x the Standard rate. The 8x billing multiplier coincidentally matches the headline 8x speed figure, which is exactly why the distinction gets lost: they are separate numbers for separate reasons. Enterprise billing remains subject to the workspace's agreement.
The practical consequence for planning is simple arithmetic: Ultrafast buys you speed at a proportional cost in allowance. If your subscription allowance is the constraint, enabling Ultrafast is closer to a budget decision than a performance setting. That is also why OpenAI places it behind Pro 500 and Enterprise rather than making it a toggle for everyone.
Which Plans and Workspaces Get Access
The eligibility rules are more restrictive than the headline suggests, and several are stated only in the docs.
| Surface | Requirement |
|---|---|
| Codex and ChatGPT Work | Pro $500, plus eligible Enterprise and Edu plans. Enterprise workspaces have Ultrafast off by default and a workspace owner must enable it. |
| OpenAI API | Broadly available for GPT-6 Astra to all API users at low rate limits. Set model to gpt-6-astra and service_tier to ultrafast. |
| Other self-serve plans | No Ultrafast access at launch, even with purchased credits. |
| Inference residency | Ultrafast is not available to workspaces that require inference residency outside the United States. A workspace's location alone does not determine eligibility. |
Three further details for enterprise buyers. Eligible Enterprise workspaces use credit-based or USD usage-based agreements, and usage is billed according to the workspace's agreement. Eligible Edu plans use credits. And legacy Enterprise plans that rely on rate limits instead of usage-based billing are not supported. Existing per-user spend controls apply to eligible Ultrafast usage.
ChatGPT Work and Codex also share usage: both use the same pricing, credits, and usage limits. If you enable Ultrafast in one, you are drawing from the same pool in the other.
Pro 200 Reopening and Grandfathering
OpenAI reopened Pro 200 to new subscriptions, and the conditions are specific enough to quote.
- New Pro 200 subscriptions are available from the pricing page at $200 per month.
- New subscriptions that are not eligible for grandfathering include a lower usage allowance than previously offered with Pro 200, which OpenAI describes as reflecting its increasingly efficient models. The monthly price stays $200.
- If your Pro 200 subscription was active at the eligibility cutoff or during the seven days before it, you keep your previous included usage allowance through October 29, 2026 while the subscription is active. After that date the subscription moves to the lower included usage allowance, still at $200/month.
- If your Pro 200 subscription lapsed during those seven days, you can still receive the previous allowance through October 29, 2026 by subscribing again.
- Keeping your current allowance does not upgrade your plan or add Ultrafast. It remains Pro 200.
OpenAI states that only users affected by the allowance change will receive an email with additional details. If you are planning a purchase, read the allowance terms rather than only the price: the gap between a grandfathered and a new Pro 200 allowance is a real difference in usable capacity at the same $200 headline.
We're also reopening Pro 200 subscriptions, with continued access to frontier models like Astra, including our new GPT-6.1 Sol model which brings near-Astra capabilities to a model you can use every day.
— @OpenAI September 29, 2026
One thing OpenAI's public material does not state is a dollar value for usage beyond the plan prices and the “25× the ChatGPT Plus allowance” framing for Pro 500. The recap links chatgpt.com/pricing for the feature comparison. We did not use any credit denomination that is not published in a first-party source.
Using Ultrafast in the API
On the API side, OpenAI offers Ultrafast mode as the fastest service tier, broadly available for GPT-6 Astra and available to all API users at low rate limits. If your organization works with an OpenAI account team, OpenAI says to contact them to request higher rate limits.
The configuration is two parameters. Set model to gpt-6-astra and service_tier to ultrafast in each request.
# Python (per OpenAI's Ultrafast mode docs)
# Install: pip install --upgrade "openai[realtime]"
# Set OPENAI_API_KEY in your environment.
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-astra",
input="Explain why the sky is blue in one sentence.",
service_tier="ultrafast",
)
print(response.output_text)For agentic loops with many tool calls, OpenAI's documented pattern reuses one WebSocket connection and passes each previous response ID as previous_response_id so the second and later turns continue from the prior response instead of resending the whole exchange. OpenAI's example streams two responses over the same connection and reuses it for later turns and tool results.
There is an HTTP path as well, but the documentation is explicit that persistent WebSockets are the recommended route when tool calls are frequent, precisely because per-request connection overhead can eat the latency gains.
One important billing note: with an API key, Codex uses API token pricing instead, and ChatGPT credit multipliers do not apply. The 8x and 6x allowance multipliers described above are a ChatGPT-plan concept; API usage is metered in tokens at published API rates.
Fast Mode and Ultrafast Are Different
OpenAI also has a lower tier called Fast mode, and mixing them up is a common error. Fast mode speeds up supported models. OpenAI lists GPT-6.1 Sol, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna as supported where available, and for GPT-5.6 and GPT-5.5 the documented speed increase is 1.5x. It uses included subscription limits at 2.5x the Standard rate, with purchased credits and Enterprise pay-as-you-go billed at 2x the Standard rate.
Feature coverage is the other difference. Fast mode is usable in the ChatGPT desktop app, the Codex CLI, and the IDE extension when you sign in with ChatGPT, toggled with /fast in the CLI, with /statusline to show it in the footer, or persisted through service_tier = "fast" plus [features].fast_mode = true in config.toml. Ultrafast, by contrast, is restricted to Pro 500 and eligible Enterprise/Edu workspaces.
OpenAI also notes in its model documentation that GPT-6.1 Sol supports Standard and Fast where available, which is why Sol appearing in Fast mode today while Sol Ultrafast is “coming soon” is consistent rather than contradictory.

The GPT-5.5 Retirement Date
A concrete deadline sits in the same documentation set: GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex on all plans on October 14, 2026. The OpenAI API is not affected. If you rely on GPT-5.5 in a ChatGPT or Codex workflow, that is a near-term migration, not a background concern.
Choosing a Tier
Four options now sit side by side, and they are not interchangeable. This table maps the decision to the constraint rather than to the headline number.
| If your constraint is… | Consider | Why |
|---|---|---|
| Cost per completed task | GPT-6.1 Sol at Standard speed | $2 input, $10 output, and $0.10 cached input per million tokens, with vendor-reported near-Astra performance on agentic coding and computer use. |
| Peak output throughput | GPT-6 Astra Ultrafast | Up to 8x token generation in Codex (300 tokens per second). Requires Pro 500 or eligible Enterprise/Edu. |
| Maximum reasoning quality | GPT-6 Astra at Standard effort | Astra still holds the highest tested score on OpenAI's hardest scientific benchmark, and the lowest measured misaligned-outcome rate. |
| Whole-task latency in an agentic loop | Fix the transport first | No tier fixes tool latency. A persistent WebSocket connection is the documented prerequisite for seeing Ultrafast's gains. |
Implementation Notes for Ultrafast
If you do adopt Ultrafast, four details from OpenAI's documentation will save time.
- Reuse one connection across turns and tool results. OpenAI's example streams two responses over the same WebSocket and passes the first response's ID as
previous_response_idfor the second. The same pattern extends to tool results: keep the connection open rather than reconnecting per call. - Handle the failure event types. OpenAI's sample code treats
response.failedandresponse.incompleteas error conditions, and separately raises if the connection closes before the response completes. In an agent loop, a silent disconnect is the failure mode most likely to look like a model problem. - Do not assume HTTP is equivalent. Ultrafast supports HTTP through the SDK, but OpenAI states that for agentic applications with frequent tool calls you should use a persistent WebSocket to reduce overhead between requests. If you benchmark Ultrafast over stateless HTTP and see only a modest gain, that is the expected result, not a bug.
- Reconcile two billing systems if you use both surfaces. With an API key, Codex uses API token pricing and ChatGPT credit multipliers do not apply. Teams that use ChatGPT Work during the day and an API key in CI are billed by two different mechanisms for the same model, which complicates internal cost attribution unless you decide upfront who pays for which workload.
There is also a capacity question the documentation answers indirectly: Ultrafast for GPT-6 Astra is available to all API users at low rate limits, and OpenAI says to contact your account team for higher limits. If Ultrafast becomes load- bearing for a production feature, that conversation is the scaling path rather than a self-serve setting.
Pay For It If
Pay for Pro 500 if your work is generation-throughput-bound rather than tool-latency-bound, and you already use most of a lower tier's allowance. Pro 500 is the only Pro plan with Ultrafast, so it is the only way to evaluate the tier as an individual rather than through an enterprise agreement. The 8x included-usage multiplier means it also consumes your allowance about as fast as it produces tokens, so go in expecting to manage headroom.
Use Ultrafast in the API if you are building a latency-sensitive application and can restructure your client around a persistent WebSocket connection. Without that change, OpenAI's own docs warn that network overhead can reduce the gains. It is broadly available for GPT-6 Astra to all API users at low rate limits, which makes it a cheaper way to evaluate the tier than a $500 subscription.
Stay where you are if your bottleneck is reasoning quality or cost per completed task rather than tokens per second. GPT-6.1 Sol at $2 input and $10 output per million tokens with $0.10 cached input is the cost-efficiency lever, and Sol Ultrafast is not generally available yet. Also stay put if your Pro 200 subscription is in the grandfathering window and you mainly want to keep your existing allowance, and OpenAI states that keeping the allowance does not upgrade you or add Ultrafast, so the decision is separate.
FAQ
How fast is Ultrafast?
OpenAI reports up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API, relative to GPT-6 Astra in Standard mode. The documentation states this measures token generation speed, not billing rates or overall task completion time.
Which ChatGPT plans include Ultrafast?
Among Pro plans, only Pro 500 at $500 per month. Ultrafast is also available in Codex and ChatGPT Work on eligible Enterprise and Edu plans, where it is off by default and a workspace owner must enable it. Other self-serve plans have no access at launch, even with purchased credits.
Does Ultrafast cost more on top of the $500 plan?
Ultrafast uses your included usage first, then draws from your credit balance after that allowance is used. Included subscription limits are consumed at 8x the Standard rate for GPT-6 Astra; purchased credits and Enterprise pay-as-you-go usage are billed at 6x the Standard rate.
Can I get Ultrafast on Pro 100 or Pro 200 by buying credits?
No. OpenAI states that at launch, buying credits on Pro 100 or Pro 200 does not unlock Ultrafast.
Is Pro 200 available again?
Yes, at $200 per month. New subscriptions that are not eligible for grandfathering include a lower usage allowance than previously offered. Subscriptions active at the cutoff or within the seven days before it keep the previous allowance through October 29, 2026.
When does GPT-6.1 Sol get Ultrafast?
OpenAI lists GPT-6.1 Sol Ultrafast as coming soon, and the GPT-6.1 Sol launch page says it will be offered in the coming days with up to 8x faster token generation compared with its standard speed in Codex. No fixed date was published.
Why do I need WebSockets for Ultrafast?
OpenAI strongly recommends WebSockets for Ultrafast, especially for agentic applications making many tool calls in quick succession, because without a persistent connection network overhead can reduce the latency gains.
Does Ultrafast apply to every model?
No. It is broadly available for GPT-6 Astra, with preview access for GPT-5.6 Sol, and OpenAI says GPT-6.1 Sol Ultrafast is coming in the coming days. Separately, GPT-5.5 is set to retire from ChatGPT, ChatGPT Work, and Codex on all plans on October 14, 2026, while the API is unaffected.
Can I turn Ultrafast off?
Yes, and on Enterprise and Edu it is off by default until a workspace owner enables it. On Pro 500 you select it in the model picker, so it is a per-conversation choice rather than a permanent plan setting. The reason to be deliberate is the multiplier: included usage is consumed at 8x the Standard rate, so a session left on Ultrafast burns the allowance faster than the token count alone suggests.
How was this article verified?
Every figure here comes from OpenAI's own material, checked on September 29, 2026: the DevDay 2026 recap, the Ultrafast API guide, the speed documentation in the ChatGPT and Codex docs, the Pro tiers help-center article, the GPT-6.1 Sol launch post, and the @OpenAI announcement posts. No third-party benchmarks, screenshots from other outlets, or unreleased pricing are used. Where OpenAI publishes no number — absolute credit balances, model-level Ultrafast limits, or a date for Sol Ultrafast, this article says so rather than estimating.
Sources
- OpenAI — DevDay 2026 Recap (Ultrafast availability, Pro 500, announcement list)
- OpenAI — Speed docs (Ultrafast and Fast mode, billing multipliers, eligibility, residency rule)
- OpenAI — Ultrafast mode in the API (service_tier, WebSocket guidance, rate limits, code samples)
- OpenAI Help Center — About ChatGPT Pro tiers (Pro 100, Pro 200, Pro 500 prices and grandfathering)
- OpenAI — WebSocket mode (incremental inputs for Ultrafast loops)
- OpenAI — Introducing GPT-6.1 Sol (Sol Ultrafast coming soon, token prices)
- @OpenAI on X — Ultrafast announcement (8x Codex, 6x API, 300 tokens/second)
- @OpenAI on X — Pro 500 introduction and Ultrafast availability
- @OpenAI on X — Pro 200 reopening
Related on Agentpedia: GPT-6.1 Sol: Benchmarks, Pricing and API Guide, DevDay 2026: Everything OpenAI Announced, Antigravity Credits and Pricing Explained, and Cheapest AI Coding Models in 2026.
Get the latest on AI, LLMs & developer tools
New MCP servers, model updates, and guides like this one — delivered weekly.