Plans & Speed

ChatGPT Ultrafast, Pro 500 and Pro 200: Plans Guide (2026)

OpenAI launched Ultrafast on September 29, 2026: a premium speed tier that generates tokens up to 8x faster (300 tokens per second) in Codex and up to 6x faster in the API. It arrives with a new Pro 500 plan at $500 per month (the only Pro tier that includes Ultrafast) and the reopening of Pro 200 to new subscriptions with a lower usage allowance. This guide separates the speed claim from the billing multiplier, maps exactly which plans and workspaces get access, and shows the API request shape.

Ultrafast and Pro 500 hero art — a speed tier sitting above standard processing
Ultrafast raises token generation speed; the plan tier decides whether you can use it. Diagram: Agentpedia.

What Shipped on September 29

Ultrafast is OpenAI's premium speed tier for workloads where speed matters most. Per the DevDay 2026 recap, GPT-6 Astra Ultrafast became available on September 29, 2026 in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans, while GPT-6.1 Sol Ultrafast is coming soon.

Alongside it, OpenAI introduced Pro 500, described as offering “our highest usage allowance at 25 times the ChatGPT Plus allowance” and including access to Ultrafast, and reopened Pro 200 subscriptions with continued access to frontier models including GPT-6.1 Sol.

This is Ultrafast. Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API.

— @OpenAI September 29, 2026
Official Astra Ultrafast title art over a starfield, showing availability in ChatGPT, Codex, and the API
Official Ultrafast art naming the three surfaces: ChatGPT, Codex, and the API. (Image: OpenAI)

What Ultrafast Actually Measures

OpenAI's product documentation is unusually precise about the metric, and the precision matters: Ultrafast “generates tokens up to 8x faster than GPT-6 Astra in Standard mode in Codex.” OpenAI then adds the caveat directly — this comparison measures token generation speed, not billing rates or overall task completion time.

So the honest reading of “up to 8x faster” is: up to 8x on token generation throughput, up to 300 tokens per second in Codex, and up to 6x in the API. It is not a claim that a task finishes 8x sooner. Agentic work spends time on tool calls, filesystem operations, network round trips, and model reasoning as well as token emission. A tier that doubles output speed can change wall-clock time for a long-generation task far more than for a task dominated by tool latency.

That is also why OpenAI recommends WebSockets for this workload. Its API documentation states that without a persistent connection, network overhead can reduce the latency gains, and strongly recommends WebSockets especially for agentic applications that make many tool calls in quick succession.

Pro 100, Pro 200 and Pro 500

OpenAI's Pro tiers help article now lists three monthly Pro plans. Prices and Ultrafast inclusion are stated in a table there.

PlanMonthly priceUltrafastNotes
Pro 100$100/monthNot includedBaseline Pro tier. Buying credits on Pro 100 does not unlock Ultrafast at launch.
Pro 200$200/monthNot includedMore included usage than Pro 100. Reopened to new subscriptions; new non-grandfathered subscriptions include a lower usage allowance.
Pro 500$500/monthIncludedHighest included usage of the three Pro plans and the only Pro tier with Ultrafast. Uses included usage first, then credits.

Prices as published in OpenAI's “About ChatGPT Pro tiers” help article on September 29, 2026. Check current pricing before subscribing; plan structures change.

Two structural details are easy to miss. First, Pro 500 is the only Pro tier that includes Ultrafast, buying credits on Pro 100 or Pro 200 does not unlock Ultrafast at launch. Second, on Pro 500 Ultrafast uses your included usage first and then draws from your credit balance once that allowance is used. Included usage is consumed before credits on other eligible features too.

Official Pro 500 artwork with Ultrafast lettering on a starfield
Official Pro 500 art from the DevDay 2026 recap. (Image: OpenAI)

Ultrafast is available today for GPT-6 Astra in Codex, ChatGPT Work, and the API, with GPT-6.1 Sol coming soon. To access it in Codex and ChatGPT Work, we're introducing Pro 500—a new plan with our highest usage limits.

— @OpenAI September 29, 2026

Billing Multipliers, Not Speed

This is the single most misread part of the launch. OpenAI publishes two different numbers per tier, a speed figure and a billing multiplier, and the documentation says outright that the billing multipliers do not describe speed increases.

TierSpeedBilled against included subscription limitsPurchased credits / Enterprise pay-as-you-go
Fast mode1.5x for GPT-5.6 and GPT-5.5; model-dependent for other supported models2.5x the Standard rate2x the Standard rate
Ultrafast (GPT-6 Astra)Up to 8x faster token generation than Standard in Codex8x the Standard rate6x the Standard rate

So for GPT-6 Astra, Ultrafast consumes included subscription limits at 8x the Standard rate, and purchased credits or Enterprise pay-as-you-go usage are billed at 6x the Standard rate. The 8x billing multiplier coincidentally matches the headline 8x speed figure, which is exactly why the distinction gets lost: they are separate numbers for separate reasons. Enterprise billing remains subject to the workspace's agreement.

The practical consequence for planning is simple arithmetic: Ultrafast buys you speed at a proportional cost in allowance. If your subscription allowance is the constraint, enabling Ultrafast is closer to a budget decision than a performance setting. That is also why OpenAI places it behind Pro 500 and Enterprise rather than making it a toggle for everyone.

Which Plans and Workspaces Get Access

The eligibility rules are more restrictive than the headline suggests, and several are stated only in the docs.

SurfaceRequirement
Codex and ChatGPT WorkPro $500, plus eligible Enterprise and Edu plans. Enterprise workspaces have Ultrafast off by default and a workspace owner must enable it.
OpenAI APIBroadly available for GPT-6 Astra to all API users at low rate limits. Set model to gpt-6-astra and service_tier to ultrafast.
Other self-serve plansNo Ultrafast access at launch, even with purchased credits.
Inference residencyUltrafast is not available to workspaces that require inference residency outside the United States. A workspace's location alone does not determine eligibility.

Three further details for enterprise buyers. Eligible Enterprise workspaces use credit-based or USD usage-based agreements, and usage is billed according to the workspace's agreement. Eligible Edu plans use credits. And legacy Enterprise plans that rely on rate limits instead of usage-based billing are not supported. Existing per-user spend controls apply to eligible Ultrafast usage.

ChatGPT Work and Codex also share usage: both use the same pricing, credits, and usage limits. If you enable Ultrafast in one, you are drawing from the same pool in the other.

Pro 200 Reopening and Grandfathering

OpenAI reopened Pro 200 to new subscriptions, and the conditions are specific enough to quote.

  • New Pro 200 subscriptions are available from the pricing page at $200 per month.
  • New subscriptions that are not eligible for grandfathering include a lower usage allowance than previously offered with Pro 200, which OpenAI describes as reflecting its increasingly efficient models. The monthly price stays $200.
  • If your Pro 200 subscription was active at the eligibility cutoff or during the seven days before it, you keep your previous included usage allowance through October 29, 2026 while the subscription is active. After that date the subscription moves to the lower included usage allowance, still at $200/month.
  • If your Pro 200 subscription lapsed during those seven days, you can still receive the previous allowance through October 29, 2026 by subscribing again.
  • Keeping your current allowance does not upgrade your plan or add Ultrafast. It remains Pro 200.

OpenAI states that only users affected by the allowance change will receive an email with additional details. If you are planning a purchase, read the allowance terms rather than only the price: the gap between a grandfathered and a new Pro 200 allowance is a real difference in usable capacity at the same $200 headline.

We're also reopening Pro 200 subscriptions, with continued access to frontier models like Astra, including our new GPT-6.1 Sol model which brings near-Astra capabilities to a model you can use every day.

— @OpenAI September 29, 2026

One thing OpenAI's public material does not state is a dollar value for usage beyond the plan prices and the “25× the ChatGPT Plus allowance” framing for Pro 500. The recap links chatgpt.com/pricing for the feature comparison. We did not use any credit denomination that is not published in a first-party source.

Using Ultrafast in the API

On the API side, OpenAI offers Ultrafast mode as the fastest service tier, broadly available for GPT-6 Astra and available to all API users at low rate limits. If your organization works with an OpenAI account team, OpenAI says to contact them to request higher rate limits.

The configuration is two parameters. Set model to gpt-6-astra and service_tier to ultrafast in each request.

# Python (per OpenAI's Ultrafast mode docs)
# Install: pip install --upgrade "openai[realtime]"
# Set OPENAI_API_KEY in your environment.

from openai import OpenAI

client = OpenAI()
response = client.responses.create(
    model="gpt-6-astra",
    input="Explain why the sky is blue in one sentence.",
    service_tier="ultrafast",
)

print(response.output_text)

For agentic loops with many tool calls, OpenAI's documented pattern reuses one WebSocket connection and passes each previous response ID as previous_response_id so the second and later turns continue from the prior response instead of resending the whole exchange. OpenAI's example streams two responses over the same connection and reuses it for later turns and tool results.

There is an HTTP path as well, but the documentation is explicit that persistent WebSockets are the recommended route when tool calls are frequent, precisely because per-request connection overhead can eat the latency gains.

One important billing note: with an API key, Codex uses API token pricing instead, and ChatGPT credit multipliers do not apply. The 8x and 6x allowance multipliers described above are a ChatGPT-plan concept; API usage is metered in tokens at published API rates.

Fast Mode and Ultrafast Are Different

OpenAI also has a lower tier called Fast mode, and mixing them up is a common error. Fast mode speeds up supported models. OpenAI lists GPT-6.1 Sol, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna as supported where available, and for GPT-5.6 and GPT-5.5 the documented speed increase is 1.5x. It uses included subscription limits at 2.5x the Standard rate, with purchased credits and Enterprise pay-as-you-go billed at 2x the Standard rate.

Feature coverage is the other difference. Fast mode is usable in the ChatGPT desktop app, the Codex CLI, and the IDE extension when you sign in with ChatGPT, toggled with /fast in the CLI, with /statusline to show it in the footer, or persisted through service_tier = "fast" plus [features].fast_mode = true in config.toml. Ultrafast, by contrast, is restricted to Pro 500 and eligible Enterprise/Edu workspaces.

OpenAI also notes in its model documentation that GPT-6.1 Sol supports Standard and Fast where available, which is why Sol appearing in Fast mode today while Sol Ultrafast is “coming soon” is consistent rather than contradictory.

Official OpenAI model cards for GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna with input, output, and cached input prices
Which model you pair with Ultrafast matters more than the tier itself: Astra is the Ultrafast-capable model today, and Sol costs one-fifth as much per token. (Image: OpenAI)

The GPT-5.5 Retirement Date

A concrete deadline sits in the same documentation set: GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex on all plans on October 14, 2026. The OpenAI API is not affected. If you rely on GPT-5.5 in a ChatGPT or Codex workflow, that is a near-term migration, not a background concern.

Choosing a Tier

Four options now sit side by side, and they are not interchangeable. This table maps the decision to the constraint rather than to the headline number.

If your constraint is…ConsiderWhy
Cost per completed taskGPT-6.1 Sol at Standard speed$2 input, $10 output, and $0.10 cached input per million tokens, with vendor-reported near-Astra performance on agentic coding and computer use.
Peak output throughputGPT-6 Astra UltrafastUp to 8x token generation in Codex (300 tokens per second). Requires Pro 500 or eligible Enterprise/Edu.
Maximum reasoning qualityGPT-6 Astra at Standard effortAstra still holds the highest tested score on OpenAI's hardest scientific benchmark, and the lowest measured misaligned-outcome rate.
Whole-task latency in an agentic loopFix the transport firstNo tier fixes tool latency. A persistent WebSocket connection is the documented prerequisite for seeing Ultrafast's gains.

Implementation Notes for Ultrafast

If you do adopt Ultrafast, four details from OpenAI's documentation will save time.

  • Reuse one connection across turns and tool results. OpenAI's example streams two responses over the same WebSocket and passes the first response's ID as previous_response_id for the second. The same pattern extends to tool results: keep the connection open rather than reconnecting per call.
  • Handle the failure event types. OpenAI's sample code treats response.failed and response.incomplete as error conditions, and separately raises if the connection closes before the response completes. In an agent loop, a silent disconnect is the failure mode most likely to look like a model problem.
  • Do not assume HTTP is equivalent. Ultrafast supports HTTP through the SDK, but OpenAI states that for agentic applications with frequent tool calls you should use a persistent WebSocket to reduce overhead between requests. If you benchmark Ultrafast over stateless HTTP and see only a modest gain, that is the expected result, not a bug.
  • Reconcile two billing systems if you use both surfaces. With an API key, Codex uses API token pricing and ChatGPT credit multipliers do not apply. Teams that use ChatGPT Work during the day and an API key in CI are billed by two different mechanisms for the same model, which complicates internal cost attribution unless you decide upfront who pays for which workload.

There is also a capacity question the documentation answers indirectly: Ultrafast for GPT-6 Astra is available to all API users at low rate limits, and OpenAI says to contact your account team for higher limits. If Ultrafast becomes load- bearing for a production feature, that conversation is the scaling path rather than a self-serve setting.

Pay For It If

Pay for Pro 500 if your work is generation-throughput-bound rather than tool-latency-bound, and you already use most of a lower tier's allowance. Pro 500 is the only Pro plan with Ultrafast, so it is the only way to evaluate the tier as an individual rather than through an enterprise agreement. The 8x included-usage multiplier means it also consumes your allowance about as fast as it produces tokens, so go in expecting to manage headroom.

Use Ultrafast in the API if you are building a latency-sensitive application and can restructure your client around a persistent WebSocket connection. Without that change, OpenAI's own docs warn that network overhead can reduce the gains. It is broadly available for GPT-6 Astra to all API users at low rate limits, which makes it a cheaper way to evaluate the tier than a $500 subscription.

Stay where you are if your bottleneck is reasoning quality or cost per completed task rather than tokens per second. GPT-6.1 Sol at $2 input and $10 output per million tokens with $0.10 cached input is the cost-efficiency lever, and Sol Ultrafast is not generally available yet. Also stay put if your Pro 200 subscription is in the grandfathering window and you mainly want to keep your existing allowance, and OpenAI states that keeping the allowance does not upgrade you or add Ultrafast, so the decision is separate.

FAQ

How fast is Ultrafast?

OpenAI reports up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API, relative to GPT-6 Astra in Standard mode. The documentation states this measures token generation speed, not billing rates or overall task completion time.

Which ChatGPT plans include Ultrafast?

Among Pro plans, only Pro 500 at $500 per month. Ultrafast is also available in Codex and ChatGPT Work on eligible Enterprise and Edu plans, where it is off by default and a workspace owner must enable it. Other self-serve plans have no access at launch, even with purchased credits.

Does Ultrafast cost more on top of the $500 plan?

Ultrafast uses your included usage first, then draws from your credit balance after that allowance is used. Included subscription limits are consumed at 8x the Standard rate for GPT-6 Astra; purchased credits and Enterprise pay-as-you-go usage are billed at 6x the Standard rate.

Can I get Ultrafast on Pro 100 or Pro 200 by buying credits?

No. OpenAI states that at launch, buying credits on Pro 100 or Pro 200 does not unlock Ultrafast.

Is Pro 200 available again?

Yes, at $200 per month. New subscriptions that are not eligible for grandfathering include a lower usage allowance than previously offered. Subscriptions active at the cutoff or within the seven days before it keep the previous allowance through October 29, 2026.

When does GPT-6.1 Sol get Ultrafast?

OpenAI lists GPT-6.1 Sol Ultrafast as coming soon, and the GPT-6.1 Sol launch page says it will be offered in the coming days with up to 8x faster token generation compared with its standard speed in Codex. No fixed date was published.

Why do I need WebSockets for Ultrafast?

OpenAI strongly recommends WebSockets for Ultrafast, especially for agentic applications making many tool calls in quick succession, because without a persistent connection network overhead can reduce the latency gains.

Does Ultrafast apply to every model?

No. It is broadly available for GPT-6 Astra, with preview access for GPT-5.6 Sol, and OpenAI says GPT-6.1 Sol Ultrafast is coming in the coming days. Separately, GPT-5.5 is set to retire from ChatGPT, ChatGPT Work, and Codex on all plans on October 14, 2026, while the API is unaffected.

Can I turn Ultrafast off?

Yes, and on Enterprise and Edu it is off by default until a workspace owner enables it. On Pro 500 you select it in the model picker, so it is a per-conversation choice rather than a permanent plan setting. The reason to be deliberate is the multiplier: included usage is consumed at 8x the Standard rate, so a session left on Ultrafast burns the allowance faster than the token count alone suggests.

How was this article verified?

Every figure here comes from OpenAI's own material, checked on September 29, 2026: the DevDay 2026 recap, the Ultrafast API guide, the speed documentation in the ChatGPT and Codex docs, the Pro tiers help-center article, the GPT-6.1 Sol launch post, and the @OpenAI announcement posts. No third-party benchmarks, screenshots from other outlets, or unreleased pricing are used. Where OpenAI publishes no number — absolute credit balances, model-level Ultrafast limits, or a date for Sol Ultrafast, this article says so rather than estimating.

Sources

Related on Agentpedia: GPT-6.1 Sol: Benchmarks, Pricing and API Guide, DevDay 2026: Everything OpenAI Announced, Antigravity Credits and Pricing Explained, and Cheapest AI Coding Models in 2026.

Get the latest on AI, LLMs & developer tools

New MCP servers, model updates, and guides like this one — delivered weekly.