Model Release

GPT-6.1 Sol: Benchmarks, Pricing and API Guide (2026)

OpenAI launched GPT-6.1 Sol on September 29, 2026 at DevDay 2026: a mid-tier upgrade that approaches GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra's standard input and output token prices. Standard API pricing is $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. This guide covers what is confirmed, every OpenAI-reported benchmark with its exact settings, the alignment numbers, and where Sol stops being the right model. All benchmark figures below are vendor-reported.

GPT-6.1 Sol hero art — a cost-versus-capability curve sitting between GPT-6 Sol and GPT-6 Astra
GPT-6.1 Sol sits between GPT-6 Sol and GPT-6 Astra on the cost curve. Diagram: Agentpedia.

What Shipped on September 29

GPT-6.1 Sol is an upgrade to GPT-6 Sol, not a new generation. OpenAI's launch page frames it as a model that “nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work” at one-fifth of Astra's standard input and output token prices. The developer identifier is gpt-6.1-sol.

The single most consequential number for anyone running agents is the cached input price: $0.10 per million tokens, which OpenAI describes as 95% less than standard input pricing and 50% less than GPT-6 Sol's cached input price. Long multi-step agent runs re-send the same system prompt, tool definitions, and file context on every turn. A 90% cut on the cached portion of those requests changes the economics of a workload far more than a headline output price does.

OpenAI announced the model inside its DevDay 2026 keynote block; the recap page lists it alongside dots, Ultrafast, and Codex updates under the heading “A whole new way to work with AI.”

GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It's the most cost-efficient model for its performance available today.

— @OpenAI September 29, 2026
Official GPT-6.1 Sol title art on a dark starry background
Official GPT-6.1 Sol art from the DevDay 2026 recap page. (Image: OpenAI)

API Pricing and the Cache Discount

OpenAI published the pricing on both the launch page and the official model-card graphic. Prices below are per million tokens, as of September 29, 2026.

ModelInputCached inputOutputPositioning
GPT-6 Astra$10.00$1.00$50.00Highest capability tier
GPT-6.1 Sol$2.00$0.10$10.00Near-Astra at one-fifth of Astra's token price
GPT-6 Luna$0.10$0.01$0.50High-volume everyday work

Per million tokens. Source: OpenAI model-card graphic published September 29, 2026. Prices are volatile; confirm current rates before committing a budget.

Official OpenAI model cards for GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna with input, output, and cached input prices
The official three-tier pricing card. Astra $10/$50 with $1.00 cached input; GPT-6.1 Sol $2/$10 with $0.10 cached input; Luna $0.10/$0.50 with $0.01 cached input. (Image: OpenAI)

The practical read: if your workload sends a large, stable prefix on every request (repository context, tool schemas, a long system prompt), so the cached input line item is where Sol's discount compounds. Workloads that send fresh context every call see the 5x input and output discount instead, which is still substantial but a different calculation.

Agentic Coding: DeepSWE v1.1

DeepSWE v1.1 evaluates complex software-engineering tasks inside real codebases. OpenAI reports that GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost, while beating GPT-6 Sol's best score by 6.4 percentage points at a lower reasoning effort and cost. The companion developer post puts specific numbers on it: 75.2% at high reasoning effort against GPT-6 Sol's best of 68.8% at maximum effort, at approximately 76% lower cost per task.

Read the effort settings carefully. The two figures are not measured at the same reasoning level: 75.2% is at high effort and 68.8% is at maximum effort. That is the point OpenAI is making: Sol reaches a higher score without spending the maximum reasoning budget, which is where the cost-per-task saving comes from.

On DeepSWE v1.1, GPT-6.1 Sol achieves 75.2% at high reasoning effort, surpassing GPT-6 Sol's best score of 68.8% at maximum effort, at approximately 76% lower cost per task.

— @OpenAIDevs September 29, 2026
Official DeepSWE chart plotting score against cost per task for GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Astra
Official DeepSWE chart: score against cost per task. The pale-yellow GPT-6.1 Sol series reaches a higher score at lower cost than the orange GPT-6 Sol series, while Astra's blue series sits further right on the cost axis. (Image: @OpenAIDevs)

The chart shows curves, not single points, and the curves cross. Sol is not uniformly better than Astra on this benchmark; it dominates on the low-cost end of the axis, which is the region most production agents operate in.

Professional Work: GDP.pdf and AutomationBench

Two benchmarks cover document-heavy and workflow-heavy professional tasks.

GDP.pdf measures how accurately models answer professional questions using complex PDF documents including tables, charts, diagrams, and fine-print details. OpenAI reports GPT-6.1 Sol scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings, and approaches GPT-6 Astra's performance at roughly one-fifth the cost per task. The benchmark draws prompts from professional workflows in finance, healthcare, legal, and seven other domains.

AutomationBench 1.0.6 tests whether agents complete multi-step business workflows using 47 tools across sales, marketing, operations, support, finance, and HR. OpenAI reports 31.7% at medium reasoning effort, 2.2 points above Opus 5.5 at roughly a third of the cost, and up 4.8 points from GPT-6 Sol at the same setting.

One honest note sits in OpenAI's own footnote: the Claude Fable 5.1 datapoint on AutomationBench understates that model's actual cost, because it omits the cost of fallbacks, which occurred on roughly 40% of tasks. If you compare vendor charts across labs, that omission matters.

Official GDP.pdf chart plotting score against cost per task for GPT-6.1 Sol, GPT-6 Sol, GPT-6 Astra, and Opus 5.5 with fallbacks
Official GDP.pdf chart. The 1490x1310 source image plots scores in the 20-34% band against cost per task up to $2. (Image: @OpenAIDevs)

On AutomationBench, GPT-6.1 Sol scores 31.7% at medium reasoning effort, up 4.8 percentage points from GPT-6 Sol at the same setting. On OSWorld 2.0's offline set, it scores 71.4% versus Astra's 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra's cost per task.

— @OpenAIDevs September 29, 2026

Computer Use: OSWorld 2.0

OSWorld 2.0's offline set evaluates agents on long-horizon computer-use workflows spanning everyday and professional tasks. OpenAI reports GPT-6.1 Sol at 71.4% versus Astra's 73.5%, both at maximum reasoning effort, at roughly one-seventh of Astra's cost per task. The launch page adds that Sol outperforms GPT-6 Sol by seven percentage points at maximum effort at less than half the cost.

The configuration detail is in the footnote: these are partial rewards on the offline set from the v2026.08.08 release of the benchmark. Partial reward means tasks that complete some but not all steps still score. Do not read 71.4% as “71.4% of workflows completed end-to-end.”

Official OSWorld 2.0 offline set chart plotting score against cost per task for GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Astra
Official OSWorld 2.0 offline chart. Scores run 40-75% against cost per task up to $10, so the horizontal distance between Sol and Astra is the cost argument. (Image: @OpenAIDevs)

Scientific Research: Terminal-Bench Science

On Terminal-Bench Science 0.1, which covers scientific workflows including data analysis, simulation, and theorem proving, OpenAI reports GPT-6.1 Sol more than doubles GPT-6 Sol's score at maximum reasoning effort at less than half the cost per task. At maximum effort, Sol costs an average of $5.47 per task, compared with $23.21 for Opus 5.5 and $23.80 for Astra.

Critically, OpenAI states that GPT-6 Astra still achieves the highest score of the models tested at 68.1% and should be used for the most difficult scientific research tasks. This is the clearest case in the launch material where the cheaper model is explicitly not the recommendation.

Factuality on Difficult Prompts

OpenAI's factuality evaluation measures the share of answers containing at least one factual error. On deliberately difficult prompts, GPT-6.1 Sol reduces that share from 11.4% to 7.7% at low reasoning effort (roughly a 32% reduction) and stays within 1.9 percentage points of Astra across tested settings at less than one-fifth the cost per task.

The methodology caveat is significant and OpenAI states it twice: the prompts come from de-identified ChatGPT conversations where a user had already flagged a factual error from a prior model. These are error-inducing conversations chosen because a model failed, so the absolute error rates are far higher than typical usage. The value of the number is the relative improvement, not the absolute rate.

Official factuality chart titled factual error rate on difficult prompts, lower is better, plotting answers with any factual error against cost per task
Official factuality chart. The Sol series sits left of Astra on the cost axis with overlapping error rates. (Image: @OpenAIDevs)

Safety and Alignment Evaluations

OpenAI reports substantial improvement over GPT-6 Sol in its alignment evaluations, bringing GPT-6.1 Sol closer to Astra. The launch page makes three specific claims: Sol is more transparent about its limitations, more reliable at respecting user intent and safety constraints, and shows lower failure rates on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks.

On the automated safety review evaluation, OpenAI observed no attempts to bypass the safety reviewer, matching GPT-6 Astra and GPT-6 Sol. Full details are in the GPT-6.1 Sol system card addendum, a PDF published alongside the launch.

ModelMisaligned outcome rate (lower is better)Fails to disclose broken search tool
GPT-6 Astra2.4%1.5%
GPT-6.1 Sol4.3%2.1%
GPT-6 Luna13.7%28.7%
GPT-6 Sol17.4%4.9%

Left column: OpenAI's computer-use safety stress test, deliberately adversarial, effort set to maximum. Right column: the broken-search-tool disclosure eval from the launch page. Both sets of evaluations are designed to elicit failures and do not represent typical-use failure rates.

Official computer-use safety stress test bar chart showing misaligned outcome rate for GPT-6 Astra 2.4%, GPT-6.1 Sol 4.3%, GPT-6 Luna 13.7%, and GPT-6 Sol 17.4%
Official computer-use safety stress test. Lower is better; these are adversarial evaluations, not typical-use rates. (Image: @OpenAI)
Official API pricing card for GPT-6.1 Sol showing $2 per million input tokens and $10 per million output tokens
OpenAI's developer-facing pricing card for GPT-6.1 Sol: $2 per million input tokens and $10 per million output tokens, with cached input at $0.10. (Image: @OpenAIDevs)

Availability and Model Choice

GPT-6.1 Sol is available starting September 29, 2026 to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. The launch page is explicit that GPT-6.1 Sol is not yet available in Chat — a distinction the DevDay recap page blurs by saying “available to all API, Plus, Pro, Business, Enterprise, and Edu users.” Where the two disagree, follow the product page: Work and Codex now, Chat later.

Developers reach it through the OpenAI API as gpt-6.1-sol. OpenAI states that GPT-6.1 Sol Ultrafast is coming in the coming days, with up to 8x faster token generation compared with its standard speed in Codex.

OpenAI's own routing guidance is worth repeating verbatim in substance: choose Astra when maximizing quality matters most, and choose GPT-6.1 Sol for complex work you want to run more often at substantially lower cost.

BenchmarkGPT-6.1 Sol resultComparisonConfiguration
DeepSWE v1.175.2% at high reasoning effortAbove GPT-6 Sol's best score of 68.8% at maximum effortAgentic software engineering in real codebases; roughly 76% lower cost per task than the Sol figure
AutomationBench 1.0.631.7% at medium reasoning effort+4.8 points over GPT-6 Sol at the same setting; 2.2 points above Opus 5.5 at about a third of the costEnd-to-end business workflows across 47 tools in sales, marketing, operations, support, finance, and HR
OSWorld 2.0 (offline set)71.4% at maximum reasoning effort2.1 points behind Astra's 73.5% at roughly one-seventh of Astra's cost per task; seven points above GPT-6 Sol at under half the costPartial reward on the offline set from the v2026.08.08 release
Terminal-Bench Science 0.1More than doubles GPT-6 Sol's score at maximum effort$5.47 average cost per task versus $23.21 for Opus 5.5 and $23.80 for Astra; Astra remains highest at 68.1%Scientific workflows including data analysis, simulation, and theorem proving
Factuality on difficult promptsError rate 11.4% to 7.7% at low reasoning effortAbout a 32% reduction; within 1.9 points of Astra across tested settings at under one-fifth the cost per taskDe-identified conversations where a user flagged an earlier model's factual error; not representative of typical usage

All figures are OpenAI-reported as of September 29, 2026. Benchmark configurations are preserved exactly as stated in the launch page and developer posts, including differing reasoning-effort settings and partial-reward scoring.

Migration Checklist

If you are moving work from GPT-6 Sol or Astra to GPT-6.1 Sol, the change is mostly a configuration and prompt-caching decision rather than an API break.

  1. Confirm the identifier. Use gpt-6.1-sol in your model configuration and check that your provider account can see it before changing production traffic.
  2. Measure your cached-prefix ratio. Log the share of input tokens served from cache on your current workload. Workloads with a large stable prefix benefit most from the $0.10 cached input rate; a low cache-hit workload sees a smaller real saving.
  3. Re-check reasoning effort. OpenAI's coding comparison pits high effort (Sol) against maximum effort (GPT-6 Sol). If your pipelines pin a reasoning level, re-run your own eval at both settings before assuming the vendor delta carries over.
  4. Keep an Astra route for peak difficulty. OpenAI still recommends Astra for the hardest scientific research tasks, where it holds the highest tested score at 68.1%. A two-model router is the intended pattern.
  5. Re-run your agentic safety tests. The alignment numbers improve substantially over GPT-6 Sol, but Sol still shows a 4.3% misaligned outcome rate on the adversarial computer-use stress test versus Astra's 2.4%. If you gate on that metric, re-measure rather than assume parity.
  6. Watch the transparency eval if you use search. Sol fails to disclose a broken search tool in 2.1% of cases, better than GPT-6 Sol's 4.9% but slightly worse than Astra's 1.5%. If silent guessing is a failure mode you care about, test it.

How to Read OpenAI's Cost Curves

OpenAI published three of these results as cost-versus-score curves rather than headline numbers, which is more informative than a single score but also easier to misread. A few rules make them useful.

  • Do not read a crossing point as a verdict. The DeepSWE v1.1, GDP.pdf, and OSWorld 2.0 curves all cross. On DeepSWE the GPT-6.1 Sol series reaches roughly the mid-70s percent band at a small fraction of the cost axis where Astra's series arrives, but Astra's series extends further right because it is being run at higher reasoning effort. When two curves cross, the question is which region your workload lives in, not which line is above the other.
  • Curves include the reasoning-effort sweep. Each point on a series is a different effort setting, so the horizontal spread is partly the cost of thinking longer. A workload pinned to maximum effort will not see the left-hand end of the Sol curve.
  • Check the y-axis range before quoting a delta. The GDP.pdf chart plots scores in a 20–34% band. A four-point difference on that axis is a large relative difference; the same four points on a 0–80% axis would be noise. OpenAI does not hide this, but the visual scale is where over-claiming starts.
  • Treat per-task cost as the unit that matters. Every OpenAI figure here is cost per task, which already folds in reasoning tokens, retries, and output length. That is the right unit for budgeting an agent, and a better predictor of your bill than the per-million-token rates alone.

The practical implication is that you should reproduce the axes rather than the score. Take a sample of your own tasks, run them under two or three reasoning settings on both models, and plot your own cost-versus-success curve. The vendor chart tells you where to look; it does not tell you where your workload sits.

What OpenAI Did Not Publish

Knowing the boundaries of the evidence is part of using it well. As of September 29, 2026, the launch material leaves several things unstated.

  • No context window or max output figure. The GPT-6.1 Sol launch page does not state a context window or maximum output length, and unlike the GPT-6 Astra launch there is no linked model page in this material. If your architecture depends on either number, confirm it against the API model documentation before migrating.
  • No knowledge cutoff. Not stated on the launch page.
  • No independent reproduction. Every benchmark in this article is OpenAI-reported. The benchmark owners (Surge HQ for GDP.pdf, Zapier for AutomationBench, the OSWorld team) publish methodology, which is why the configuration column above can be precise, but OpenAI selected the reported settings.
  • No latency figures. The launch page gives cost deltas, not time-to-first-token or throughput. Sol Ultrafast is the announced answer to latency, and it is listed as coming soon.
  • No deprecation notice for GPT-6 Sol. OpenAI frames GPT-6.1 Sol as an upgrade to GPT-6 Sol rather than a replacement, and no retirement date for GPT-6 Sol appears in the launch material. Note that a separate retirement — GPT-5.5 from ChatGPT, ChatGPT Work, and Codex on October 14, 2026 — is documented elsewhere in OpenAI's model documentation, and the API is not affected.

None of these gaps are unusual for a same-day launch post. They are listed so you can decide what to verify yourself before moving production traffic rather than discovering it after.

Prompting and Integration Notes

A few practical consequences follow directly from the published numbers.

Structure prompts for cache hits. The largest single lever is the $0.10 cached input rate. That requires a stable prefix: put system instructions, tool schemas, and long-lived repository context first, and keep them byte-identical across requests. Reordering or interleaving changing content into the prefix is the common way teams accidentally pay full input rates on every call.

Let the model spend effort where the task is hard. OpenAI's best Sol coding result came at high effort rather than maximum, and its cost advantage over GPT-6 Sol's best result came from not running at the top of the effort range. If you currently hardcode maximum reasoning effort for every coding task, the benchmark set suggests that is a cost decision worth revisiting per task class.

Watch factuality at low effort specifically. The largest factuality gain OpenAI reports is at low reasoning effort, from 11.4% to 7.7% on error-inducing prompts. If your pipeline deliberately runs cheap, low-effort calls for classification or extraction, that is the setting where the upgrade is most visible, and also the setting where absolute error rates are highest.

Keep a fallback path to Astra. OpenAI's own benchmark set contains a case — Terminal-Bench Science 0.1, where Astra holds the highest tested score at 68.1% — where the more expensive model is still the recommendation. A router with an Astra escape hatch for peak-difficulty requests matches how OpenAI describes the two models rather than treating Sol as a drop-in replacement.

Use It If

Use GPT-6.1 Sol if you run agentic coding, computer-use, or document-heavy professional workflows at volume and your bottleneck is cost per completed task, not peak capability. The cached-input rate makes long-running agents with stable prefixes materially cheaper, and the benchmark set OpenAI published is exactly the set those workloads care about.

Stay on or route to GPT-6 Astra if you need the highest score on the hardest scientific research tasks (Astra 68.1% on Terminal-Bench Science 0.1 against a Sol cost of $5.47 per task versus Astra's $23.80, you are buying difficulty headroom, not throughput), or if your safety gate requires the lowest measured misaligned-outcome rate.

Skip the migration for now if your product depends on Sol appearing inside the main Chat surface (the launch page states it is not yet there), or if you need Ultrafast on Sol, which OpenAI lists as coming soon rather than available today.

FAQ

What is the GPT-6.1 Sol model ID?

gpt-6.1-sol in the OpenAI API.

How much does GPT-6.1 Sol cost?

$2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens, as of September 29, 2026. That is one-fifth of GPT-6 Astra's $10/$50 standard input and output rates; Astra's cached input is $1.00.

Is GPT-6.1 Sol better than GPT-6 Astra?

No, and OpenAI does not claim it is. Sol approaches Astra's performance on several benchmarks at much lower cost per task, but Astra retains the highest tested score on Terminal-Bench Science 0.1 and a lower misaligned-outcome rate on the computer-use stress test. OpenAI's guidance is to pick Astra for maximum quality and Sol for work you want to run more often.

Is GPT-6.1 Sol available in ChatGPT?

In ChatGPT Work and in Codex, yes, for Plus, Pro, Business, Enterprise, and Edu users as of September 29, 2026. The launch page states Sol is not yet available in Chat itself.

Where is the GPT-6.1 Sol system card?

OpenAI published a system card addendum as a PDF alongside the launch, linked from the launch page. It covers the alignment and safety evaluations summarized above.

Sources

Related on Agentpedia: GPT-6 Astra: Complete Guide to OpenAI's Frontier Release, GPT-5.6 Sol, Terra & Luna: Pricing and Fast Mode, and DevDay 2026: Everything OpenAI Announced.

Get the latest on AI, LLMs & developer tools

New MCP servers, model updates, and guides like this one — delivered weekly.