Model Release

Gemini 4 Argon: Complete Guide to Benchmarks, Pricing and Access

Google DeepMind’s new frontier model ships with a 1M-token output limit, a cyber-defense focus, and access that starts with trusted testers rather than the public API. Here is every official benchmark table, the introductory pricing, the internal Google case studies, and the parts of the developer surface that are still missing.

Editorial illustration of a glowing model core emitting a very long ribbon of tokens that overflows a wide open output window, with a translucent shield scanning a stack of code repository bars.
Illustration: a frontier model core with an output budget large enough to overflow a single window, plus the cyber-defense scanning motif Google leads with.

Google DeepMind announced Gemini 4 Argon on September 30, 2026 — a frontier model positioned for complex workflows in coding, enterprise knowledge work, and cybersecurity defense, rolling out to a limited set of trusted testers through Google’s Fairwind Program. The headline specification is an output limit of 1 million tokens, which Google describes as industry-leading and which is a large increase over the 64K output ceiling that governed earlier Gemini releases. Google is simultaneously publishing introductory per-token pricing, but has not yet added Argon to its public Gemini API model list or pricing table.

This guide is built from first-party sources only: the official Google announcement (authored by Koray Kavukcuoglu), the DeepMind Fairwind Program page, the earlier Fairwind launch post from September 2, 2026, the live Gemini API models documentation and pricing page, the DeepMind model card index and Frontier Safety Framework pages, and the launch posts from Google’s own accounts on X. Every figure below is checked on September 30, 2026, the day of the announcement.

What Google announced

The lead announcement comes from the Google DeepMind account and frames Argon as a frontier model rather than a refresh of an existing tier. The stated target workloads — complex workflows across coding, enterprise knowledge work, and cybersecurity defense — describe long-running agentic work rather than chat. The distribution method is the part that differs most from a normal Gemini release: a set of trusted testers, reached through the Fairwind Program, on day one.

Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

— @GoogleDeepMind September 30, 2026

A follow-up from the same account addresses the output ceiling directly, and a longer thread from Sundar Pichai gives the framing Google wants attached to the release: broad internal use at Google already, with feedback from teams spanning coding through quantum computing.

With a 1M token output limit, Argon adds a deeper level of reasoning to tackle…

— @GoogleDeepMind September 30, 2026

Google’s corporate account repeats the three workload areas and states the output limit as a differentiator. Read the two texts side by side and the emphasis is consistent: the same three domains, the same 1M number, the same “frontier” label.

Today we’re introducing Gemini 4 Argon. It delivers frontier performance in complex workflows across real-world software engineering, knowledge work, and cybersecurity defense with an industry-leading 1M token output limit.

— @Google September 30, 2026

Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

— @sundarpichai September 30, 2026

That last post is worth reading carefully because it sets the evidence standard for everything that follows. Google is describing its own internal usage and its own benchmark runs. Pichai’s post says “here’s a look at the benchmarks” and attaches charts, which makes those charts primary evidence for the claims in this article — but they remain vendor-run measurements, not independent ones.

Official Google key art for Gemini 4 Argon: a stylized blue number four rendered in front of a deep blue gradient.
Official key art published with the announcement. Source: Google blog

The 1M token output limit

Output limits are the constraint that quietly shapes agent architecture. When a model can only emit 64K tokens in one response, long-horizon work has to be chopped into steps, state has to be handed between calls, and the orchestration layer carries the burden of stitching a result back together. Raising the ceiling to 1 million tokens changes where that burden sits — a single generation can carry work that previously required a loop.

The official numbers for comparison, stated on Google’s own channels:

SpecificationGemini 4 ArgonPrior Gemini generation
Output token limit1,000,000 tokens64K tokens (earlier Gemini releases)
Google’s framing“Industry-leading 1M token output limit”—
Stated reasoning effect“adds a deeper level of reasoning to tackle” complex tasks—
Public API model IDNot published in the Gemini API models list—
Independent verificationNone found in first-party sources—

Three qualifications belong next to that table. First, no first-party source checked for this article defines the exact semantics of the 1M ceiling — whether streaming, multi-turn continuation, or per-attempt accounting applies. Second, Google has not published a rate limit, latency, or time-to-first-token figure alongside the limit. Third, because Argon has no model ID in the public API documentation as of September 30, there is currently no way for an outside developer to test the claim directly.

For teams doing capacity planning, the practical reading is that 1M output tokens is an announced capability under limited access, not a purchasable tier. The right move is to design against the ceiling you can actually call today and treat the larger budget as upside when access arrives, rather than re-architecting around a number you cannot yet measure.

Benchmark snapshot

Google’s announcement leads with benchmark gains across coding, general agentic capability, and long-context video understanding. Every number in this section is vendor-reported: Google selected the comparison models, ran or commissioned the evaluations, and published the charts. No independent reproduction was found in first-party sources on announcement day, and no harness configuration or run log was published alongside these figures.

EvaluationArgon resultGoogle’s framingCaveat
DeepSWE v1.1 (software engineering)77.9%New state of the artVendor-run; harness and configuration not fully published
Vals Index (economic-impact capability)Described as leadingFrontier position on the indexVendor-reported; no score reproduced independently
AutomationBench (Zapier)51.3%Ranked firstVendor-reported
LVBench (long-video understanding)91.7%State of the artVendor-reported
CWE-bench v1 (security patching)68%Tied for firstThree-way tie at 68%; see chart below
Gray Swan IPI (indirect prompt injection)Leading robustness profileMost robust of the comparison setLower attack success is better; vendor-reported

One row deserves explicit framing. CWE-bench v1 is not a clean win: the published chart shows a three-way tie at 68% involving Argon, Grok 4.7, and GPT-6 Astra. Google lists Argon as tied for first, which is accurate and also means the security-patching result does not separate Argon from the field. If a decision depends on that specific benchmark, the honest conclusion is a draw, not a lead.

The same idea applies to the benchmark set as a whole. The eval mix Google chose emphasizes software engineering, automation, video understanding, and security. That is a reasonable proxy for the three workload areas in the announcement, but it is not a general capability index, and a model can look strong on this particular selection while being unremarkable on reasoning or multilingual work.

Cyber defense evaluations

The cybersecurity results are where the announcement spends the most chart space, which matches the fact that the first access cohort is cyber defenders. Two of the charts cover public benchmarks; two cover Google-internal or partner-run security work.

Official Google chart of CWE-bench v1 results across eleven models, showing Gemini 4 Argon tied at 68 percent with Grok 4.7 and GPT-6 Astra.
CWE-bench v1 across the comparison set. Vendor-reported; note the three-way tie at 68%. Source: Google blog

CWE-bench measures whether a model can patch real vulnerability classes. A 68% result is genuine progress on a hard benchmark, and the chart makes the tie visible rather than burying it — which is the useful part for a reader making a decision.

Official Google chart of Gray Swan indirect prompt injection robustness, showing Gemini 4 Argon with the lowest attack success rate at 0.7 percent.
Gray Swan indirect prompt injection robustness (lower attack success is better). Vendor-reported. Source: Google blog

The Gray Swan chart reads in the opposite direction from the others: the metric is attack success rate, so a lower number is a better result. Google reports Argon at a 0.7% attack success rate at k=15 attempts, the strongest profile in the comparison set. For anyone deploying agents that read untrusted content — web pages, issue trackers, email — indirect prompt injection robustness is one of the few security properties with a directly measurable number, so this chart is more decision-relevant than the coding scores.

Official Google chart comparing Gemini 4 Argon at 85.8 percent against Gemini 3.8 Flash Cyber at 71.0 percent on real-world vulnerability discovery, and 70.9 against 58.2 percent on a Wiz penetration-testing benchmark.
Real-world vulnerability discovery and a partner-run penetration-testing benchmark, Argon against the previous cyber-focused model. Vendor-reported; Wiz is the partner named on the chart. Source: Google blog
Security evaluationGemini 4 ArgonGemini 3.8 Flash CyberStatus
Real-world vulnerability discovery85.8%71.0%Google-internal benchmark
Wiz penetration-testing benchmark70.9%58.2%Partner-run, reported by Google
CWE-bench v168% (tied first)Listed on chartPublic benchmark, vendor-reported run
Gray Swan IPI ASR @ k=150.7% (lowest)Listed on chartPublic benchmark, vendor-reported run

The comparison target here is instructive. Google benchmarks Argon against 3.8 Flash Cyber, its own cyber-tuned model, rather than only against competitors. Beating an internal specialist model by roughly 15 points on vulnerability discovery is a meaningful internal progression. It is also an internal comparison, evaluated on internal infrastructure, with no published test set, and the reader should treat the absolute percentages as directional. The Wiz row is different in kind: Wiz is an external security vendor and the benchmark is theirs, but Google is still the party reporting the result, and no Wiz write-up of the same run was located in first-party sources.

What Google ran internally

The most distinctive part of the announcement is a set of internal case studies, because Google is unusually specific about the work rather than just claiming productivity gains. These are internal results with no external test harness, so they illustrate capability rather than prove it — but they are concrete enough to reason about.

Internal workReported resultWhat it suggests
Language migration (C/C++ to Rust)32K lines of SIMD code replaced; result reported as 2.7× faster than the earlier Rust port and closer to optimized C++Large mechanical refactors of performance-critical code, with the caveat that “closer to” is qualitative
Fuchsia Zircon kernel migration800K+ lines migratedThe 1M output budget applied to a single very large codebase task
Memory efficiency work300 TiB freed, with an estimated 500 TiB to 1 PiB addressableInfrastructure-scale optimization passes
Quantum optimizationBeat a published baseline by 40%Research-domain work; no methodology published
Broad internal usageTeams using Argon “extensively,” from coding to quantum computing, per PichaiLong-running dogfooding before external access

The migration numbers connect most directly to the output limit. Replacing 32K lines of SIMD code, or moving 800K+ lines of kernel code, is the class of task that a 64K output ceiling makes awkward and a 1M ceiling makes feasible in fewer passes. That relationship is a plausible mechanism rather than proof, because Google has not published the prompt structure, number of turns, or review process behind either migration.

Treat the qualitative wording carefully. “2.7× faster than the Rust port” is a relative claim between two internal artifacts, and “closer to optimized C++” is Google’s own characterization rather than a benchmark result. The quantum-optimization figure has no published baseline reference or evaluation protocol, so it is best read as a signal of where Google is applying the model rather than as a performance claim you can compare against anything.

Fairwind Program and rollout

Access is the constraint that matters most today, and it runs through the Fairwind Program rather than the Gemini API. Fairwind was introduced on September 2, 2026 as a defender-focused program; the DeepMind program page describes it as working with a global network of partners, and reports 650+ partners worldwide. That figure covers the program as a whole, not Argon access specifically.

The Fairwind program page also connects earlier DeepMind security work to this release, naming Gemini 3.8 Flash Cyber and CodeMender, the code-security agent DeepMind published in October 2025. Argon is the third step in that line: a cyber-tuned model, an automated patching agent, and now a frontier model whose first cohort is defenders.

StageWho gets accessChannelDate published
1 (live)Trusted testers, described by Google as an initial cohort of cyber defendersFairwind ProgramSeptember 30, 2026
2 (announced)Paid API customers and Google AI Ultra subscribersGemini API, consumer subscriptionNot dated
3 (announced)Developers, enterprises, and consumers broadlyFull surfaceNot dated

No dates are attached to stages two and three. That matters for planning: if you are waiting for Argon on the general Gemini API, there is currently no published timeline to plan against. Teams with a near-term need should verify what is available on the API today rather than build a schedule around an unannounced date.

Pricing

Google published introductory pricing with the announcement. The structure is a promotional rate that steps up afterward, with a large cached-input discount, and no separate premium for the cyber-defense capability.

Line itemIntroductory rateStandard rate after the introductory period
Input$2.00 / 1M tokens$4.00 / 1M tokens
Output$10.00 / 1M tokens$20.00 / 1M tokens
Cached input95% discount on the cached portionDiscount structure not published for the standard period
Cyber-defense capabilityNo separate line itemNo separate line item
Public pricing-page rowNot present as of September 30, 2026Not present as of September 30, 2026

Three observations. The introductory input and output rates are exactly double after the promotion ends, so any cost model should be built on $4/$20 and treated as a discount rather than a price. The cached-input discount is the largest lever available: for agent workloads that resend a large stable prefix, a 95% reduction on cached input dominates everything else in the table. And because there is no row on the public Gemini API pricing page (last updated September 8, 2026), these rates exist in the announcement text and the launch charts but not yet in the billing documentation a developer would consult.

For a worked sense of scale, output tokens dominate cost in long-generation work: at $10 per million output tokens, a single 100K-token generated artifact costs about a dollar during the introductory period and about two dollars afterward. That is the arithmetic that makes the 1M output budget interesting and also the arithmetic that makes it expensive — a full 1M-token generation is roughly $10 at the promotional rate and $20 at the standard rate, before any input cost.

API status today

This is the section where the announcement and the developer documentation disagree, and it is the most important practical gap in the release.

SurfaceWhat first-party sources show
Gemini API models listNo Argon model ID; the page was last updated September 4, 2026 — before the announcement
Gemini API pricing pageNo Argon pricing row; last updated September 8, 2026
DeepMind model cards indexNo Gemini 4 Argon model card published
Safety documentationNo Argon-specific evaluation report located
Official demo videoNone located on Google or DeepMind channels
Access channel in the announcementFairwind Program trusted testers

The pattern is consistent: the marketing surfaces moved on announcement day and the reference surfaces have not caught up. That is normal for a limited-access launch, but it has a concrete consequence — there is no public model ID, so there is nothing for an external developer to call, and no published rate limit, context-window breakdown, or regional availability to plan around.

For agents and skills that wrap model calls, this is the practical scenario worth designing for: a new frontier model that is announced with pricing but not yet addressable. The Gemini 3.8 Flash guide and the Gemini 3.8 Live guide both cover models that already have stable IDs, which makes them the safer base for anything shipping this month.

Safeguards

The rollout framing is explicitly cautious. Access begins with a narrow cohort of defenders, which is itself a risk-limiting decision: security work has clear legitimate uses, measurable outcomes, and partners who can report misuse. Google’s Frontier Safety Framework pages describe the company’s staged evaluation approach for frontier models, and the announcement points to that framework as the governance context.

What is missing is the Argon-specific evidence. Google has not published a model card, a capability threshold assessment, or a misuse evaluation for this model as of September 30, 2026. A reader cannot currently check Argon against the framework’s own thresholds. The honest summary is that Google describes a responsible staged rollout and that the document that would let outsiders audit that claim has not been published yet.

What this changes for agent workflows

Setting the access question aside, the capabilities Google emphasizes point at four workflow patterns that get easier if the claims hold.

Long single-pass generation. A 1M output budget makes it reasonable to ask for a complete artifact in one call — a full migration patch set, a large refactor across files, a generated test suite — instead of driving a loop of small edits. The tradeoff is cost concentration and review difficulty: a single enormous response is harder to review incrementally than ten smaller ones, so the workflow needs a verification step rather than more prompting.

Security review with measurable robustness. The indirect prompt injection result is the most actionable number in the release for agent builders. If an agent ingests web content, tickets, or email, injection robustness is a live threat model, and a 0.7% attack success rate is a figure you can reason about — while remembering it is a vendor-reported run at a fixed attempt budget.

Migration and large-codebase work. The C/C++ to Rust and Fuchsia figures describe long-horizon code transformation, which is the workload class where output limits historically forced the most orchestration overhead. Teams doing this work today typically split it across many calls; the Argon framing suggests consolidating passes.

Vulnerability triage. Combining a higher vulnerability-discovery rate with a patching benchmark (CWE-bench) describes a triage-to-fix loop. The caveat is unchanged: both numbers are vendor-reported, and the patching result is a three-way tie.

A practical adoption sequence, given the access constraints: keep your current model IDs in production; prototype against the models with stable public IDs today; track the Gemini API models and pricing pages for an Argon row; and if you are a defender organization, evaluate the Fairwind Program as the entry path rather than waiting on general availability. If your work depends on the 1M output budget, instrument the cost of a long generation before committing to it, because that is the dimension the pricing table rewards least.

Caveats

Everything measured here is vendor-reported. Google ran or commissioned the benchmarks, picked the comparison models, and published the charts. No independent replication of any Argon number was found in first-party sources on announcement day, and several results — the vulnerability-discovery and quantum-optimization figures — describe internal work with no published methodology.

The security benchmark lead is not always a lead. CWE-bench v1 is a three-way tie at 68%. Google’s “tied for first” language is accurate; treating that row as a win would not be.

The API surface is not published. No model ID, no pricing row, no model card, no rate limits, no latency figures, and no regional availability. The 1M output limit in particular has no published semantics for streaming or multi-turn use.

Two of the four launch posts in the embeds above are quoted in part. The DeepMind output-limit post and Google’s derived posts continue beyond the opening clause shown; the full text is available at the linked post. Nothing here paraphrases text that could not be read directly.

Rollout dates are unannounced. The path from trusted testers to paid API customers to general availability has no published timeline, so any schedule built on Argon arriving on the public API is an assumption.

Internal results are illustrative. The memory, migration, and quantum figures show where Google is applying the model. They are not comparable to external benchmarks and, in the case of the 2.7× Rust comparison, the baseline is another internal artifact.

Use it if

Use Argon if you are a defender organization or a partner with Fairwind access and your work involves vulnerability discovery, security patching at scale, or long-horizon code migration. The combination of a 1M output budget and the cyber-defense evaluation profile is aimed squarely at that work, and at the introductory rate it is priced competitively against frontier alternatives.

Wait if you need a stable public model ID, published rate limits, a model card, or a documented safety evaluation. None of those exist for Argon today. For production code this month, the models with stable IDs are the safer choice, and the 3.8 Flash tier covers most high-volume workloads at a fraction of the cost.

Watch it if your architecture is shaped by output limits. The 1M budget is the specification most likely to change how long-running agents are structured, and it is worth tracking when the API documentation catches up — at which point the practical questions become rate limits, latency, and whether cached-input pricing holds at the standard tier.

FAQ

What is Gemini 4 Argon?

Google DeepMind’s frontier model announced September 30, 2026, built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense. Google describes it as delivering frontier performance with an industry-leading 1M token output limit.

Can I use Gemini 4 Argon through the Gemini API today?

Not as a documented public model. The Gemini API models page (last updated September 4, 2026) has no Argon entry and the pricing page (last updated September 8, 2026) has no Argon row. Access currently runs through the Fairwind Program for trusted testers.

What is the 1M token output limit and why does it matter?

It is the maximum number of tokens Argon can generate in a response, an increase over the 64K ceiling governing earlier Gemini releases. Larger output budgets mean long-horizon work can be produced in fewer passes instead of being split across many calls. Google has not published the limit’s exact semantics for streaming or multi-turn continuation, nor any rate limit or latency figure.

How much does Gemini 4 Argon cost?

Introductory pricing is $2.00 per million input tokens and $10.00 per million output tokens, with a 95% discount on cached input. After the introductory period those rates double to $4.00 and $20.00 per million tokens. None of it appears on the public pricing page yet.

What benchmarks did Google publish?

DeepSWE v1.1 at 77.9% (claimed state of the art), AutomationBench at 51.3% (claimed first), LVBench at 91.7%, a leading position on the Vals Index, CWE-bench v1 at 68% tied for first, and a leading robustness profile on Gray Swan indirect prompt injection. All are vendor-reported with no independent replication.

Is the CWE-bench result a win?

No — it is a three-way tie at 68% between Argon, Grok 4.7, and GPT-6 Astra. Google lists Argon as tied for first, which is accurate, but the benchmark does not separate Argon from the field.

What is the Fairwind Program?

Google’s defender-focused program, introduced September 2, 2026, that is the access path for Argon’s first cohort. The DeepMind program page reports 650+ partners worldwide and names earlier DeepMind security work including Gemini 3.8 Flash Cyber and the CodeMender patching agent.

When will Argon be generally available?

Google has announced the rollout order — trusted testers, then paid API customers and Google AI Ultra subscribers, then developers, enterprises, and consumers — but has published no dates for the later stages.

What did Google use Argon for internally?

Reported internal work includes a C/C++ to Rust migration that replaced 32K lines of SIMD code, an 800K+ line Fuchsia Zircon kernel migration, memory work that freed 300 TiB with an estimated 500 TiB to 1 PiB addressable, and a quantum optimization result reported at 40% better than a published baseline. All are internal results without published methodology.

Is there a Gemini 4 Argon model card or system card?

Not as of September 30, 2026. The DeepMind model card index has no Argon entry, and no Argon-specific capability or misuse evaluation was located in first-party sources.

Should I migrate production workloads to Argon now?

No. There is no public model ID to call, no published rate limits, and no model card. Keep current stable model IDs in production, prototype against documented models, and revisit when the API documentation lists Argon.

How does Argon relate to Gemini 3.8?

They are different generations. Gemini 3.8 Flash and 3.8 Live are documented models with stable IDs and published pricing; Argon is a separately announced frontier model whose public API surface is not published. The 3.8 Flash Cyber model appears as the comparison baseline on Google’s Argon security charts.

Sources

Sources checked September 30, 2026:

Related reading: Gemini 3.8 Flash complete guide, Gemini 3.8 Live and Extended Thinking guide, Codex Security Cloud and Daybreak Blue guide, and GPT-6.1 Sol benchmarks and pricing guide.