AI Deep Dive

Agentic Media Economics: What Omni Flash and Nano Banana 2 Lite Actually Cost

Google gave agentic media a price list on June 30, 2026: $0.10 a second for Omni Flash video, $0.034 an image for Nano Banana 2 Lite. This is a plain breakdown of what a real pipeline costs, where the pennies-per-loop thesis holds, and where every conversational edit quietly bills you again.

Editorial hero illustration for Agentic Media Economics: Omni Flash + Nano Banana

Get the latest on AI, LLMs & developer tools

New MCP servers, model updates, and guides like this one — delivered weekly.

The Price List

Agentic media has a price list now, and it is short. Google Omni Flash generates 720p video at $0.10 per second, about $1 for a ten-second clip. Nano Banana 2 Lite generates a 1K image in under four seconds for roughly $0.034. Chain them the way Google intends, with Nano Banana 2 Lite drafting stills and Omni Flash animating and editing them, and a full 60-second product reel with a few conversational edits lands near $9.34. That is the whole economic story in three numbers. Everything else is about when those numbers stay small and when they quietly do not.

60-second product reel — worked cost (official rates)

10 candidate stills, NB2 Lite (1K):  10 x $0.0336  = $0.336
6 hero clips x 10s = 60s Omni video:  60 x $0.10    = $6.000
3 conversational edits x 10s = 30s:   30 x $0.10    = $3.000
Input tokens:                                        < $0.100
------------------------------------------------------------
Total                                               ~ $9.34

cost ~= (N_images x $0.0336) + (M_video_seconds x $0.10)

Two things fall out of that math immediately. Video seconds dominate the bill; images are a rounding error. And the "few edits" line is not free: the three edits above cost $3.00 on their own. Hold that thought, because it is the crack the whole optimistic narrative runs through. For the API mechanics behind the chain, the task types, and the store-equals-true requirement that makes editing possible, see our Omni Flash developer guide.

Per-Second Cost vs Veo 3.1

Priced per second, Omni Flash sits exactly on top of Veo 3.1 Fast and undercuts standard Veo 3.1 by three quarters. The catch is the ceiling: Omni Flash is 720p-only. Veo keeps the higher resolutions, and that single fact is what keeps it employed. The Omni Flash $0.10 figure is official; the comparison grid below is VentureBeat's reporting on published Veo rates.

Per secondOmni FlashVeo 3.1 LiteVeo 3.1 FastVeo 3.1
720p$0.10$0.05$0.10$0.40
1080pn/a$0.08$0.12$0.40
4Kn/an/a$0.30$0.60

Read the table the right way. At 720p, Omni Flash matches Veo 3.1 Fast, doubles Veo 3.1 Lite, and costs a quarter of standard Veo 3.1, but it cannot produce 1080p or 4K at all. For a vertical clip that will be watched on a phone, 720p is not a compromise. For a broadcast spot or a hero video on a landing page, it is a wall. Price is only half the decision; if you care about the quality gap against Veo, Sora, Seedance, and Kling rather than the sticker, read our Omni Flash vs Veo, Sora, Seedance and Kling comparison.

The Pennies-Per-Loop Thesis

The bull case is a loop. When each unit costs cents, you can let an agent iterate without watching a meter: generate ten stills, throw away nine, animate the winner, try three directions, keep one. Shubham Saboo put the anchor version of this argument online, and it spread because the arithmetic is genuinely hard to argue with at the generation step.

Rohan Paul adds the missing half: the product is the chain, not either model alone. The Interactions API keeps session state on Google's servers, so an agent stacks sequential edits without re-uploading the clip each turn. That is what turns two cheap primitives into a pipeline an autonomous agent can actually run.

Every Edit Is a New Bill

Here is the part the loop math skips: every conversational edit is a brand-new, separately billed generation. There are no free tweaks. The clearest statement of the counterargument came from RobCDoesAI:

"720p only, no 1080p/4K. Priced same as Veo 3.1 Fast, but every conversational edit is a brand-new paid generation. Solid for social. Not the tool for anything needing real resolution."— RobCDoesAI

VentureBeat drew the line precisely: a stateful model "does not change the cost of an edit, it changes the number of wasted ones." Server-side session state is a real saving, but it is a saving on rework, not on the unit. If you re-render a ten-second clip four times to get the shot, you paid for forty seconds of video no matter how conversational the interface felt.

There is a second multiplier almost nobody prices in: refusals. Omni Flash blocks editing of uploaded video across the EEA, Switzerland, the UK, and some regions entirely, and it refuses a meaningful share of clips on safety grounds, often with a generic non-explanatory error. A failed generation you have to redo is still spend. A working cost guide suggests multiplying your clean estimate by 1.5x to 2x to cover regenerations forced by filters. Budget for the refusals, not just the successes.

Integration Economics

The per-unit price is only half the economic picture. The other half is what it costs to wire these models into a pipeline an agent can operate. This is where MCP earns its keep. Higgsfield exposes Omni Flash and Nano Banana 2 Lite through an MCP server, so a coding agent plans and runs a marketing-asset pipeline end to end instead of a human clicking through a UI.

Now the myth that has to die. Vercel AI Gateway does not discount these models. It is no-markup: you pay the same as calling Google directly. What you buy is unification, one API key, one invoice, and automatic failover across providers, not a lower per-second rate. Anyone telling you a gateway halves your cost is wrong. That distinction matters more than it sounds, because once an autonomous agent is spending real money per second, the billing and metering surface is part of the architecture, not an afterthought; see our explainer on agentic payments.

Who Is Actually Paying

Who is paying for this at scale? On launch day the named list was not hobbyists. SiliconANGLE reported day-one adopters including WPP (through WPP Open), Figma (in Weave), Artlist, Manus AI, and invideo. Michael Gerstenhaber, VP of Product at Google Cloud, called the pricing "aggressive," which is the kind of thing a company says when it is buying market share.

The most revealing line came from Artlist: generation is now "faster than ideation," which keeps creators "inside the idea." That is the actual value proposition and the actual risk in one sentence. When rendering a variant is cheaper and quicker than deciding whether you want it, the constraint moves from budget to judgment. The cost of a bad idea is no longer money; it is the seconds you spent looking at it.

Five Tools Into One Model

The cost comparison most people get wrong is per-second against per-second. The real comparison, and VentureBeat's sharpest structural point, is one model against the five-tool pipeline it replaces: an LLM for the script, a text-to-image model, an image-to-video model, a lip-sync model, and a voice model. Each of those was a separate vendor, a separate key, a separate integration to maintain, and a separate seam where the output of one stage failed to match the input of the next.

Omni Flash collapses that into a single natively multimodal model that generates video with audio in one call. So the savings story is as much about vendor and overhead consolidation as it is about the $0.10 line item. You stop paying five subscriptions, stop maintaining four brittle handoffs, and stop losing quality at every conversion boundary. For a small team, that consolidation can dwarf the raw per-second math.

When the Math Works, When It Does Not

The math works when the output is short-form and social: vertical clips, faceless educational content, product reels, ad variants, storyboards, and anything watched on a phone where 720p is invisible. It works when an agent runs the loop unattended and a few wasted generations are cheaper than a human's attention. And it works hardest when you are replacing a five-tool stack, because the consolidation savings stack on top of the low unit price.

The math breaks when you need real resolution: 1080p or 4K sends you back to Veo and a different price sheet. It breaks when your workflow lives in a blocked region or trips the safety filters, because the regeneration tax is real and unbudgeted. And it breaks when a human is doing the editing by hand, because conversational edits are not free tweaks; every "actually, make it warmer" re-bills the whole clip. Price the loop, not the click.

FAQ

How much does Gemini Omni Flash cost per second?

Omni Flash generates 720p video at $0.10 per second, which is about $1 for a ten-second clip. That is the official Google API rate and the same per-second price as Veo 3.1 Fast.

How much does Nano Banana 2 Lite cost per image?

Roughly $0.034 for a 1K image (officially $0.0336 per 1K), generated in under four seconds. Batch pricing halves it to $0.0168, but batch is not available through the Interactions API yet, so interactive pipelines pay the standard rate.

What does a full 60-second product reel actually cost?

Using Google published rates: about $0.34 for ten candidate stills, $6.00 for sixty seconds of Omni Flash video, and $3.00 for three ten-second conversational edits, plus under $0.10 of input tokens. That is roughly $9.34 end to end. Video seconds dominate; images are a rounding error.

Is Omni Flash cheaper than Veo 3.1?

At 720p, Omni Flash matches Veo 3.1 Fast, is double Veo 3.1 Lite, and costs a quarter of standard Veo 3.1. But Omni Flash is 720p only. If you need 1080p or 4K you are back on Veo, and the comparison no longer applies.

Does every conversational edit cost money?

Yes. Every edit is a brand-new, separately billed generation that re-renders the full clip duration. Three ten-second edits cost $3.00, not free tweaks. A stateful session reduces the number of wasted edits; it does not make any single edit free.

Does Vercel AI Gateway make Omni Flash cheaper?

No. Vercel AI Gateway is no-markup: you pay the same as calling Google directly. Its value is unification (one key, one invoice, automatic failover), not a discount. Claims that a gateway halves your cost are false.

Glossary

  • Omni FlashGoogle video model (gemini-omni-flash-preview). 720p only, 3 to 10 second clips, 24 FPS, priced at $0.10 per second of output.
  • Nano Banana 2 LiteGoogle fast image model (gemini-3.1-flash-lite-image). Generates a 1K image in under four seconds for roughly $0.034; the cheap first stage of a chained pipeline.
  • Interactions APIThe stateful endpoint both models run through. It keeps session history server-side so a follow-up edit does not have to re-upload the clip.
  • Conversational editA plain-language change to an existing clip. Each one is a fresh, fully billed generation that re-renders the whole duration, not a free tweak.
  • MCPModel Context Protocol. The way a coding agent such as Claude drives an external tool (Higgsfield, Omni Flash) directly instead of a human clicking through a UI.
  • AI gatewayA routing layer, for example Vercel AI Gateway, that unifies keys, billing, and failover across providers. Typically no-markup: it does not change the per-unit price.
  • SynthID + C2PAThe invisible watermark and Content Credentials attached to every Omni Flash output so the clip can later be identified as AI-generated.

Verdict

The pennies-per-loop thesis is true and incomplete. It is true that generation is now cheap enough to run in a loop, and true that consolidating five tools into one model is a bigger saving than the sticker price suggests. It is incomplete because a conversational interface makes it feel like edits are free when each one is a fresh render, because the 720p ceiling quietly disqualifies anything that needs real resolution, and because refusals add a regeneration tax nobody puts in the spreadsheet.

The honest bottom line: for short-form social media run by an agent, the economics are genuinely new and genuinely good. For everything else, do the boring thing and price the whole loop, including the wasted seconds and the refused clips, before you believe the demo.

Sources

Pricing and specs are from official Google documentation. The comparison table, enterprise adopters, and structural framing are from named reporting. Social posts are attributed inline.

Get the latest on AI, LLMs & developer tools

New MCP servers, model updates, and guides like this one — delivered weekly.