
What Shipped on September 29
GPT-6.1 Sol is an upgrade to GPT-6 Sol, not a new generation. OpenAI's launch page frames it as a model that “nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work” at one-fifth of Astra's standard input and output token prices. The developer identifier is gpt-6.1-sol.
The single most consequential number for anyone running agents is the cached input price: $0.10 per million tokens, which OpenAI describes as 95% less than standard input pricing and 50% less than GPT-6 Sol's cached input price. Long multi-step agent runs re-send the same system prompt, tool definitions, and file context on every turn. A 90% cut on the cached portion of those requests changes the economics of a workload far more than a headline output price does.
OpenAI announced the model inside its DevDay 2026 keynote block; the recap page lists it alongside dots, Ultrafast, and Codex updates under the heading “A whole new way to work with AI.”
GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It's the most cost-efficient model for its performance available today.
— @OpenAI September 29, 2026

API Pricing and the Cache Discount
OpenAI published the pricing on both the launch page and the official model-card graphic. Prices below are per million tokens, as of September 29, 2026.
| Model | Input | Cached input | Output | Positioning |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | Highest capability tier |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | Near-Astra at one-fifth of Astra's token price |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | High-volume everyday work |
Per million tokens. Source: OpenAI model-card graphic published September 29, 2026. Prices are volatile; confirm current rates before committing a budget.

The practical read: if your workload sends a large, stable prefix on every request (repository context, tool schemas, a long system prompt), so the cached input line item is where Sol's discount compounds. Workloads that send fresh context every call see the 5x input and output discount instead, which is still substantial but a different calculation.
Agentic Coding: DeepSWE v1.1
DeepSWE v1.1 evaluates complex software-engineering tasks inside real codebases. OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost, while beating GPT-6 Sol's best score by 6.4 percentage points at a lower reasoning effort and cost. The companion developer post puts specific numbers on it: 75.2% at high reasoning effort against GPT-6 Sol's best of 68.8% at maximum effort, at approximately 76% lower cost per task.
Read the effort settings carefully. The two figures are not measured at the same reasoning level: 75.2% is at high effort and 68.8% is at maximum effort. That is the point OpenAI is making: Sol reaches a higher score without spending the maximum reasoning budget, which is where the cost-per-task saving comes from.
On DeepSWE v1.1, GPT-6.1 Sol achieves 75.2% at high reasoning effort, surpassing GPT-6 Sol's best score of 68.8% at maximum effort, at approximately 76% lower cost per task.
— @OpenAIDevs September 29, 2026

The chart shows curves, not single points, and the curves cross. Sol is not uniformly better than Astra on this benchmark; it dominates on the low-cost end of the axis, which is the region most production agents operate in.
Professional Work: GDP.pdf and AutomationBench
Two benchmarks cover document-heavy and workflow-heavy professional tasks.
GDP.pdf measures how accurately models answer professional questions using complex PDF documents including tables, charts, diagrams, and fine-print details. OpenAI reports GPT-6.1 Sol scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings, and approaches GPT-6 Astra's performance at roughly one-fifth the cost per task. The benchmark draws prompts from professional workflows in finance, healthcare, legal, and seven other domains.
AutomationBench 1.0.6 tests whether agents complete multi-step business workflows using 47 tools across sales, marketing, operations, support, finance, and HR. OpenAI reports 31.7% at medium reasoning effort, 2.2 points above Opus 5.5 at roughly a third of the cost, and up 4.8 points from GPT-6 Sol at the same setting.
One honest note sits in OpenAI's own footnote: the Claude Fable 5.1 datapoint on AutomationBench understates that model's actual cost, because it omits the cost of fallbacks, which occurred on roughly 40% of tasks. If you compare vendor charts across labs, that omission matters.

On AutomationBench, GPT-6.1 Sol scores 31.7% at medium reasoning effort, up 4.8 percentage points from GPT-6 Sol at the same setting. On OSWorld 2.0's offline set, it scores 71.4% versus Astra's 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra's cost per task.
— @OpenAIDevs September 29, 2026
Computer Use: OSWorld 2.0
OSWorld 2.0's offline set evaluates agents on long-horizon computer-use workflows spanning everyday and professional tasks. OpenAI reports GPT-6.1 Sol at 71.4% versus Astra's 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra's cost per task. The launch page adds that Sol outperforms GPT-6 Sol by seven percentage points at maximum effort at less than half the cost.
The configuration detail is in the footnote: these are partial rewards on the offline set from the v2026.08.08 release of the benchmark. Partial reward means tasks that complete some but not all steps still score. Do not read 71.4% as “71.4% of workflows completed end-to-end.”

Scientific Research: Terminal-Bench Science
On Terminal-Bench Science 0.1, which covers scientific workflows including data analysis, simulation, and theorem proving, OpenAI reports GPT-6.1 Sol more than doubles GPT-6 Sol's score at maximum reasoning effort at less than half the cost per task. At maximum effort, Sol costs an average of $5.47 per task, compared with $23.21 for Opus 5.5 and $23.80 for Astra.
Critically, OpenAI states that GPT-6 Astra still achieves the highest score of the models tested at 68.1% and should be used for the most difficult scientific research tasks. This is the clearest case in the launch material where the cheaper model is explicitly not the recommendation.
Factuality on Difficult Prompts
OpenAI's factuality evaluation measures the share of answers containing at least one factual error. On deliberately difficult prompts, GPT-6.1 Sol reduces that share from 11.4% to 7.7% at low reasoning effort (roughly a 32% reduction) and stays within 1.9 percentage points of Astra across tested settings at less than one-fifth the cost per task.
The methodology caveat is significant and OpenAI states it twice: the prompts come from de-identified ChatGPT conversations where a user had already flagged a factual error from a prior model. These are error-inducing conversations chosen because a model failed, so the absolute error rates are far higher than typical usage. The value of the number is the relative improvement, not the absolute rate.

Safety and Alignment Evaluations
OpenAI reports substantial improvement over GPT-6 Sol in its alignment evaluations, bringing GPT-6.1 Sol closer to Astra. The launch page makes three specific claims: Sol is more transparent about its limitations, more reliable at respecting user intent and safety constraints, and shows lower failure rates on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks.
On the automated safety review evaluation, OpenAI observed no attempts to bypass the safety reviewer, matching GPT-6 Astra and GPT-6 Sol. Full details are in the GPT-6.1 Sol system card addendum, a PDF published alongside the launch.
| Model | Misaligned outcome rate (lower is better) | Fails to disclose broken search tool |
|---|---|---|
| GPT-6 Astra | 2.4% | 1.5% |
| GPT-6.1 Sol | 4.3% | 2.1% |
| GPT-6 Luna | 13.7% | 28.7% |
| GPT-6 Sol | 17.4% | 4.9% |
Left column: OpenAI's computer-use safety stress test, deliberately adversarial, effort set to maximum. Right column: the broken-search-tool disclosure eval from the launch page. Both sets of evaluations are designed to elicit failures and do not represent typical-use failure rates.


Availability and Model Choice
GPT-6.1 Sol is available starting September 29, 2026 to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. The launch page is explicit that GPT-6.1 Sol is not yet available in Chat — a distinction the DevDay recap page blurs by saying “available to all API, Plus, Pro, Business, Enterprise, and Edu users.” Where the two disagree, follow the product page: Work and Codex now, Chat later.
Developers reach it through the OpenAI API as gpt-6.1-sol. OpenAI states that GPT-6.1 Sol Ultrafast is coming in the coming days, with up to 8x faster token generation compared with its standard speed in Codex.
OpenAI's own routing guidance is worth repeating verbatim in substance: choose Astra when maximizing quality matters most, and choose GPT-6.1 Sol for complex work you want to run more often at substantially lower cost.
| Benchmark | GPT-6.1 Sol result | Comparison | Configuration |
|---|---|---|---|
| DeepSWE v1.1 | 75.2% at high reasoning effort | Above GPT-6 Sol's best score of 68.8% at maximum effort | Agentic software engineering in real codebases; roughly 76% lower cost per task than the Sol figure |
| AutomationBench 1.0.6 | 31.7% at medium reasoning effort | +4.8 points over GPT-6 Sol at the same setting; 2.2 points above Opus 5.5 at about a third of the cost | End-to-end business workflows across 47 tools in sales, marketing, operations, support, finance, and HR |
| OSWorld 2.0 (offline set) | 71.4% at maximum reasoning effort | 2.1 points behind Astra's 73.5% at roughly one-seventh of Astra's cost per task; seven points above GPT-6 Sol at under half the cost | Partial reward on the offline set from the v2026.08.08 release |
| Terminal-Bench Science 0.1 | More than doubles GPT-6 Sol's score at maximum effort | $5.47 average cost per task versus $23.21 for Opus 5.5 and $23.80 for Astra; Astra remains highest at 68.1% | Scientific workflows including data analysis, simulation, and theorem proving |
| Factuality on difficult prompts | Error rate 11.4% to 7.7% at low reasoning effort | About a 32% reduction; within 1.9 points of Astra across tested settings at under one-fifth the cost per task | De-identified conversations where a user flagged an earlier model's factual error; not representative of typical usage |
All figures are OpenAI-reported as of September 29, 2026. Benchmark configurations are preserved exactly as stated in the launch page and developer posts, including differing reasoning-effort settings and partial-reward scoring.
Migration Checklist
If you are moving work from GPT-6 Sol or Astra to GPT-6.1 Sol, the change is mostly a configuration and prompt-caching decision rather than an API break.
- Confirm the identifier. Use
gpt-6.1-solin your model configuration and check that your provider account can see it before changing production traffic. - Measure your cached-prefix ratio. Log the share of input tokens served from cache on your current workload. Workloads with a large stable prefix benefit most from the $0.10 cached input rate; a low cache-hit workload sees a smaller real saving.
- Re-check reasoning effort. OpenAI's coding comparison pits high effort (Sol) against maximum effort (GPT-6 Sol). If your pipelines pin a reasoning level, re-run your own eval at both settings before assuming the vendor delta carries over.
- Keep an Astra route for peak difficulty. OpenAI still recommends Astra for the hardest scientific research tasks, where it holds the highest tested score at 68.1%. A two-model router is the intended pattern.
- Re-run your agentic safety tests. The alignment numbers improve substantially over GPT-6 Sol, but Sol still shows a 4.3% misaligned outcome rate on the adversarial computer-use stress test versus Astra's 2.4%. If you gate on that metric, re-measure rather than assume parity.
- Watch the transparency eval if you use search. Sol fails to disclose a broken search tool in 2.1% of cases, better than GPT-6 Sol's 4.9% but slightly worse than Astra's 1.5%. If silent guessing is a failure mode you care about, test it.
How to Read OpenAI's Cost Curves
OpenAI published three of these results as cost-versus-score curves rather than headline numbers, which is more informative than a single score but also easier to misread. A few rules make them useful.
- Do not read a crossing point as a verdict. The DeepSWE v1.1, GDP.pdf, and OSWorld 2.0 curves all cross. On DeepSWE the GPT-6.1 Sol series reaches roughly the mid-70s percent band at a small fraction of the cost axis where Astra's series arrives, but Astra's series extends further right because it is being run at higher reasoning effort. When two curves cross, the question is which region your workload lives in, not which line is above the other.
- Curves include the reasoning-effort sweep. Each point on a series is a different effort setting, so the horizontal spread is partly the cost of thinking longer. A workload pinned to maximum effort will not see the left-hand end of the Sol curve.
- Check the y-axis range before quoting a delta. The GDP.pdf chart plots scores in a 20–34% band. A four-point difference on that axis is a large relative difference; the same four points on a 0–80% axis would be noise. OpenAI does not hide this, but the visual scale is where over-claiming starts.
- Treat per-task cost as the unit that matters. Every OpenAI figure here is cost per task, which already folds in reasoning tokens, retries, and output length. That is the right unit for budgeting an agent, and a better predictor of your bill than the per-million-token rates alone.
The practical implication is that you should reproduce the axes rather than the score. Take a sample of your own tasks, run them under two or three reasoning settings on both models, and plot your own cost-versus-success curve. The vendor chart tells you where to look; it does not tell you where your workload sits.
What OpenAI Did Not Publish
Knowing the boundaries of the evidence is part of using it well. As of September 29, 2026, the launch material leaves several things unstated.
- No context window or max output figure. The GPT-6.1 Sol launch page does not state a context window or maximum output length, and unlike the GPT-6 Astra launch there is no linked model page in this material. If your architecture depends on either number, confirm it against the API model documentation before migrating.
- No knowledge cutoff. Not stated on the launch page.
- No independent reproduction. Every benchmark in this article is OpenAI-reported. The benchmark owners (Surge HQ for GDP.pdf, Zapier for AutomationBench, the OSWorld team) publish methodology, which is why the configuration column above can be precise, but OpenAI selected the reported settings.
- No latency figures. The launch page gives cost deltas, not time-to-first-token or throughput. Sol Ultrafast is the announced answer to latency, and it is listed as coming soon.
- No deprecation notice for GPT-6 Sol. OpenAI frames GPT-6.1 Sol as an upgrade to GPT-6 Sol rather than a replacement, and no retirement date for GPT-6 Sol appears in the launch material. Note that a separate retirement — GPT-5.5 from ChatGPT, ChatGPT Work, and Codex on October 14, 2026 — is documented elsewhere in OpenAI's model documentation, and the API is not affected.
None of these gaps are unusual for a same-day launch post. They are listed so you can decide what to verify yourself before moving production traffic rather than discovering it after.
Prompting and Integration Notes
A few practical consequences follow directly from the published numbers.
Structure prompts for cache hits. The largest single lever is the $0.10 cached input rate. That requires a stable prefix: put system instructions, tool schemas, and long-lived repository context first, and keep them byte-identical across requests. Reordering or interleaving changing content into the prefix is the common way teams accidentally pay full input rates on every call.
Let the model spend effort where the task is hard. OpenAI's best Sol coding result came at high effort rather than maximum, and its cost advantage over GPT-6 Sol's best result came from not running at the top of the effort range. If you currently hardcode maximum reasoning effort for every coding task, the benchmark set suggests that is a cost decision worth revisiting per task class.
Watch factuality at low effort specifically. The largest factuality gain OpenAI reports is at low reasoning effort, from 11.4% to 7.7% on error-inducing prompts. If your pipeline deliberately runs cheap, low-effort calls for classification or extraction, that is the setting where the upgrade is most visible, and also the setting where absolute error rates are highest.
Keep a fallback path to Astra. OpenAI's own benchmark set contains a case — Terminal-Bench Science 0.1, where Astra holds the highest tested score at 68.1% — where the more expensive model is still the recommendation. A router with an Astra escape hatch for peak-difficulty requests matches how OpenAI describes the two models rather than treating Sol as a drop-in replacement.
Use It If
Use GPT-6.1 Sol if you run agentic coding, computer-use, or document-heavy professional workflows at volume and your bottleneck is cost per completed task, not peak capability. The cached-input rate makes long-running agents with stable prefixes materially cheaper, and the benchmark set OpenAI published is exactly the set those workloads care about.
Stay on or route to GPT-6 Astra if you need the highest score on the hardest scientific research tasks (Astra 68.1% on Terminal-Bench Science 0.1 against a Sol cost of $5.47 per task versus Astra's $23.80, you are buying difficulty headroom, not throughput), or if your safety gate requires the lowest measured misaligned-outcome rate.
Skip the migration for now if your product depends on Sol appearing inside the main Chat surface (the launch page states it is not yet there), or if you need Ultrafast on Sol, which OpenAI lists as coming soon rather than available today.
FAQ
What is the GPT-6.1 Sol model ID?
gpt-6.1-sol in the OpenAI API.
How much does GPT-6.1 Sol cost?
$2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens, as of September 29, 2026. That is one-fifth of GPT-6 Astra's $10/$50 standard input and output rates; Astra's cached input is $1.00.
Is GPT-6.1 Sol better than GPT-6 Astra?
No, and OpenAI does not claim it is. Sol approaches Astra's performance on several benchmarks at much lower cost per task, but Astra retains the highest tested score on Terminal-Bench Science 0.1 and a lower misaligned-outcome rate on the computer-use stress test. OpenAI's guidance is to pick Astra for maximum quality and Sol for work you want to run more often.
Is GPT-6.1 Sol available in ChatGPT?
In ChatGPT Work and in Codex, yes, for Plus, Pro, Business, Enterprise, and Edu users as of September 29, 2026. The launch page states Sol is not yet available in Chat itself.
Where is the GPT-6.1 Sol system card?
OpenAI published a system card addendum as a PDF alongside the launch, linked from the launch page. It covers the alignment and safety evaluations summarized above.
Sources
- OpenAI — Introducing GPT-6.1 Sol (benchmarks, pricing, availability, alignment)
- OpenAI — DevDay 2026 Recap (announcement manifest, availability lines, official media)
- OpenAI — GPT-6.1 Sol system card addendum (PDF)
- @OpenAIDevs on X — DeepSWE v1.1 figure (75.2% high effort, 68.8% Sol best, ~76% lower cost per task)
- @OpenAIDevs on X — AutomationBench and OSWorld 2.0 figures
- @OpenAIDevs on X — factuality on deliberately difficult prompts
- Surge HQ — GDP.pdf benchmark (referenced by OpenAI for methodology)
- Zapier — AutomationBench 1.0.6 (referenced by OpenAI for methodology)
- OSWorld 2.0 (referenced by OpenAI for methodology)
- OpenAI — GPT-6 Astra launch page (comparison baseline)
Related on Agentpedia: GPT-6 Astra: Complete Guide to OpenAI's Frontier Release, GPT-5.6 Sol, Terra & Luna: Pricing and Fast Mode, and DevDay 2026: Everything OpenAI Announced.
Get the latest on AI, LLMs & developer tools
New MCP servers, model updates, and guides like this one — delivered weekly.