Model Release

Gemini 3.8 Flash: Complete Guide, Benchmarks, and Cyber Variant

Google shipped its third Flash model in six weeks. Here is the full official benchmark table, what the introductory $0.75/$3.75 pricing actually costs from January 2027, how the Fairwind cybersecurity variant differs, and what the numbers mean for your agent stack.

Editorial illustration of a glowing Gemini model chip streaking with speed lines, surrounded by Google-color accent nodes.
Illustration: Google’s third Flash release in six weeks, built for long-running agentic loops.

Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2, 2026 — the third Flash release in just six weeks. 3.8 Flash claims the best publicly listed scores on DeepSWE v1.1 among Flash-class models at 73.7%, wins the Vals Finance Agent and Harvey Legal Agent benchmarks outright, and launches at an introductory $0.75 per million input tokens and $3.75 per million output tokens. The Cyber variant posts 47.2% pass@1 on CWE-Bench patching at significantly lower cost than a frontier model at 47.8%.

This guide follows the official Google announcement, the DeepMind methodology pages for 3.8 Flash and 3.8 Flash Cyber, and the launch threads from Google DeepMind, Google, Sundar Pichai, and Gemini API lead Logan Kilpatrick. It was checked on September 2, 2026, the day of the announcement.

What Google launched

Two models, one day. Gemini 3.8 Flash is the general-purpose release: Google’s positioning is “our most intelligent model yet” with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning — and explicitly, a model “built to scale your AI agents.” Gemini 3.8 Flash Cyber is a separate variant tailored for cybersecurity deployment: frontier-level vulnerability detection and automated patching, available to trusted defenders through a new Fairwind Program rather than the open API.

Naming note: “Gemini 3.8” in Google's messaging is shorthand for the 3.8 generation — the generally available model is gemini-3.8-flash, plus the gated 3.8 Flash Cyber variant. There is no separate non-Flash “Gemini 3.8” model ID as of September 2026.

Both share the same foundational intelligence, accelerated by long-running agentic loops — the same pattern Google introduced with agentic video understanding one day earlier. The cadence is the story too: 3.6 Flash, 3.7 Flash, and now 3.8 Flash in six weeks, each release layering on the previous one’s agentic capabilities.

Sundar Pichai’s launch post frames the cadence directly — the third Flash release in six weeks — and calls out DeepSWE v1.1: on that benchmark, 3.8 Flash “outperforms most larger frontier models in autonomously solving complex engineering problems end to end, at a fraction of the cost.”

Official launch graphic reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, with a glowing blue geometric form on a pale blue background.
Google’s official launch graphic. Source: Google launch thread

Benchmark evidence

Logan Kilpatrick’s launch post carries the master comparison card — pricing on top, then sixteen benchmark rows against Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra. Every number below is transcribed from that official card and cross-checked against the DeepMind methodology page.

Official Gemini 3.8 Flash benchmark card comparing pricing and sixteen benchmarks against Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra, with 3.8 Flash leading on DeepSWE v1.1, Vals Finance Agent, Harvey Legal Agent, Terminal-bench 2.1, CharXiv, LVBench, HLE-Verified, and LABBench2.
The official benchmark card from Logan Kilpatrick’s launch post. Source: Logan Kilpatrick (Gemini API lead)
BenchmarkGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5GPT-5.6 SolGPT-5.6 Terra
DeepSWE v1.1 (long-horizon SWE)73.7%65.3%74.0%72.7%69.8%
Terminal-bench 2.1 (agentic terminal)89.4%85.8%89.1%88.8%87.4%
Terminal-bench 4.0 (general agent)19.1%11.2%51.8%37.3%23.6%
Vals Finance Agent v261.4%59.0%58.6%53.8%54.4%
Harvey Legal Agent (pass rate)10.0%8.8%6.7%2.5%0.8%
GDPVal-AA v2 (Elo)15451482182417101528
HLE-Verified (expert reasoning)54.9%53.6%54.4%54.5%51.1%
CharXiv Reasoning (no tools)86.2%84.5%83.7%85.8%85.9%
LVBench (long video)87.8% (agentic)85.4%75.4%82.1%78.9%
OSWorld-2.0 (computer use)59.0%50.6%75.4%62.6%50.2%
GDP.PDF (doc comprehension)35.0%34.0%37.0%40.0%29.0%
BioMysteryBench (human-solvable)88.8%87.1%90.1%79.5%83.8%
BioMysteryBench (human-difficult)56.5%43.5%49.4%44.7%49.4%
LABBench2 (biology research)86.2%82.1%84.2%82.1%81.2%

Read the columns honestly. 3.8 Flash sweeps the agent-adjacent rows — DeepSWE, Terminal-bench 2.1, Vals Finance, Harvey Legal, CharXiv, LVBench, BioMysteryBench (difficult), LABBench2 — which is exactly the “scale your AI agents” positioning. But Claude Opus 5 keeps a real lead on GDPVal-AA v2 knowledge work (1824 vs 1545 Elo), OSWorld-2.0 computer use (75.4% vs 59.0%), and the new Terminal-bench 4.0 general-agent benchmark (51.8% vs 19.1%). This is a Flash-class model that wins its class; it does not uniformly beat larger frontier models on everything.

The DeepSWE efficiency plot makes the cost story visual: 3.8 Flash sits at roughly Opus 5’s accuracy (74%) at a fraction of the per-task cost, flagged “most efficient” on Google’s chart. Datacurve’s DeepSWE is an external benchmark (deepswe.datacurve.ai), not an internal Google one.

Official DeepSWE v1.1 scatter plot of benchmark score against average cost per task, showing Gemini 3.8 Flash near Claude Opus 5 accuracy at far lower cost, flagged as most efficient.
DeepSWE v1.1: score vs cost per task. Source: Google DeepMind launch thread
Official HLE-Verified bar chart showing Gemini 3.8 Flash at 54.9%, ahead of GPT-5.6 Sol at 54.5%, Claude Opus 5 at 54.4%, Gemini 3.7 Flash at 53.6%, GPT-5.6 Terra at 51.1%, and Claude Sonnet 5 at 31.0%.
HLE-Verified, multidisciplinary expert reasoning. Source: DeepMind methodology

The efficiency story

The clearest wins are where capability meets price. Harvey’s Legal Agent Benchmark is the starkest: 3.8 Flash scores 10.0% all-pass rate — double Claude Sonnet 5’s 5.0% and ten times GPT-5.6 Terra’s 1% — while costing a fifth of Opus 5 per token. On the DeepSWE cost plot, Google’s own positioning is that 3.8 Flash reaches Opus-5-class accuracy (74.0% vs 73.7% is inside noise) at a fraction of the cost per task.

LVBench is worth a separate note: 87.8% with agentic video processing vs 87.1% static — the first chart where Google publicly shows the day-old agentic video mode folded into a headline model card. 3.7 Flash managed 85.4% on the same benchmark.

Pricing and context

The card’s pricing rows are the ones to quote, with the January 2027 step-up in mind:

TierInput /1MOutput /1M
Gemini 3.8 Flash — introductory (expires Dec 31, 2026)$0.75$3.75
Gemini 3.8 Flash — regular (from Jan 1, 2027)$1.50$7.50
Gemini 3.7 Flash — introductory (same schedule)$0.75$3.75
Claude Opus 5$5.00$25.00
Claude Sonnet 5$2.00$10.00
GPT-5.6 Sol$4.00$20.00
GPT-5.6 Terra$2.00$12.00

The footnote on the official card: the introductory price for 3.7 and 3.8 Flash expires December 31, 2026; from January 1, 2027, $1.50/1M input and $7.50/1M output apply. Even at regular pricing, 3.8 Flash undercuts every competitor on the card by 4–17× on input.

3.8 Flash Cyber and the Fairwind Program

The Cyber variant is not on the public API. Google describes it as its “most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.” Deployment environments differ from the standard model, and access is gated.

Official CyberGym Pass@1 bar chart for vulnerability discovery in C/C++, showing Gemini 3.8 Flash Cyber at 86.2%, ahead of GPT-5.5-Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 77.5%.
CyberGym Pass@1, vulnerability discovery in C/C++. Source: DeepMind Cyber methodology

Three numbers define the Cyber pitch. On CyberGym Pass@1 (vulnerability discovery in C/C++), 3.8 Flash Cyber scores 86.2% vs 85.6% for GPT-5.5-Cyber and 77.5% for 3.5 Flash Cyber. On CWE-Bench (patching, run by Collinear), it lands on the Pareto frontier at 47.2% pass@1 vs 47.8% for a leading frontier model — “yet offered at a significantly lower cost.” And in deployment, Google says the Chrome Security team got 2.6× more correct patches than from “the best commercial models that are much larger,” while Wiz measured +7.5–9.7% higher recall on its internal penetration-testing benchmark at 2.3–5.2× lower cost.

What the Fairwind Program is.

Google’s announcement names Fairwind as the access path for 3.8 Flash Cyber — a program for “trusted defenders” rather than a public API surface. If you want the model for defensive security work, that program — not AI Studio — is the door. The announcement does not publish Fairwind admission criteria or pricing.

Getting started

3.8 Flash is served through the standard Gemini API surface. The model string follows the family convention (gemini-3.8-flash); reasoning effort, thinking budget, and tool configuration work as on 3.7 Flash. Developers on the 3.7 Flash agentic-video configuration from yesterday’s launch can switch the model string and keep the processing: "agentic" parameter.

from google import genai

client = genai.Client()

response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents="Find and fix the failing test in this repository."
)

print(response.text)

Context is 1M tokens with 65,536 max output tokens, carried over from the 3.7 generation. That window is what makes the long-context MRCR gains usable in practice — and it is the same window the MRCR rows below were measured against.

Agentic video, now on the flagship card

One day before this launch, Google introduced agentic video understanding — the model deciding which parts of a video to watch instead of ingesting at a fixed 1 FPS. 3.8 Flash’s card is the first to ship with that mode baked into the headline numbers: LVBench at 87.8% agentic vs 87.1% static. For teams processing long-form video, the practical takeaway is that the agentic parameter now carries over to the newest Flash without re-validation.

Caveats

All numbers are vendor-reported. The benchmark table, the CyberGym and CWE-Bench charts, and the Chrome/Wiz deployment figures come from Google’s launch materials and DeepMind’s methodology pages. Google chose the competitor configurations. Directionally consistent with independent leaderboards (DeepSWE is Datacurve’s external benchmark), but if you are migrating a pipeline, benchmark on your own tasks.

The Terminal-bench 4.0 gap is real. 19.1% vs Opus 5’s 51.8% on general agent capabilities is not a rounding artifact — for open-ended computer-use agent work today, Opus 5 remains ahead. 3.8 Flash’s wins are concentrated in software engineering, terminal coding, and domain-specific agent tasks.

Introductory pricing is temporary. The $0.75/$3.75 rate disappears on January 1, 2027. Budget your unit economics on the $1.50/$7.50 regular rate, not the launch price.

Use it if

Use Gemini 3.8 Flash if you run coding agents, terminal workflows, or domain-specific agent pipelines (finance, legal, biology) and cost per task matters — it matches or beats frontier models on exactly those rows at a tenth of the token price. Its 1M context and agentic-video support make it the default Flash for long-document and long-video pipelines too.

Look elsewhere if your workload is open-ended computer use (OSWorld-2.0), broad knowledge-work Elo (GDPVal-AA v2), or the new Terminal-bench 4.0-style general agency — Claude Opus 5 leads those rows today. And if you need the cybersecurity variant, the open API is not the path: that is what the Fairwind Program gates.

FAQ

What is Gemini 3.8 Flash?

Google’s Flash-class model released September 2, 2026 — the third Flash release in six weeks — focused on software engineering, agentic tasks, and multi-step reasoning, with a companion cybersecurity variant.

How much does Gemini 3.8 Flash cost?

Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. From January 1, 2027, regular pricing is $1.50/$7.50 per million.

What is the context window?

1M input tokens with 65,536 max output tokens.

What is Gemini 3.8 Flash Cyber?

A separate variant for vulnerability detection and automated patching, available to trusted defenders through Google’s Fairwind Program rather than the public API. It scores 86.2% on CyberGym Pass@1 and 47.2% on CWE-Bench patching.

How does it compare to Gemini 3.7 Flash?

Same price at launch, same 1M context, with gains across software engineering (DeepSWE 73.7% vs 65.3%), terminal coding (89.4% vs 85.8%), and long video (LVBench 87.8% vs 85.4% with agentic processing).

Does it support agentic video understanding?

Yes — the official card lists LVBench at 87.8% with agentic processing vs 87.1% static, one day after Google introduced the mode.

Sources

Sources checked September 2, 2026:

Related reading: Gemini agentic video understanding guide, Gemini 3.7 Flash guide, and Gemini Omni complete guide.