Viral AI Trends
By Viral AI Trends · Published

Gemini Omni: what it is, how to use it free, and what it really costs

Gemini Omni is the model Google built instead of a Veo 4 — an any-to-any system that starts with video and edits it by conversation. It is also the model with the friendliest free route in AI video and one of the least forgiving paid ones, because every generation has a 10-second ceiling and every request is billed. This page covers both: how to try it at no cost, what a second actually costs at each resolution, and the numbers that Google has not published.

What Gemini Omni is

Gemini Omni is the family name for Google's any-to-any generative model, announced at Google I/O on 19 May 2026. Google's own framing is that it can create anything from any input and that it is starting with video — image and audio outputs are named as later additions rather than shipped features.

The first model in the family is Gemini Omni Flash. It accepts text, images, video and voice references in a single prompt, and returns a short video with synchronised audio. Google's model card describes it as transformer-based with native multimodal support for text, vision, video and audio, trained on TPUs at Google DeepMind.

Three things separate it from the video models it gets compared with:

One thing it is not. Gemini Omni is not Veo 4. Veo is Google's dedicated video line, still at Veo 3.1, and Google has not announced a successor to it. The two run side by side inside Google Flow. If you searched for "Veo 4" and landed here, the model you are looking for is called Gemini Omni — and its pricing, its clip length and its prompting style are all different from Veo's.

The free routes — and the catch

This is the part that makes Gemini Omni unusual, and it is stated plainly in Google's own I/O announcement.

For most people reading this, the YouTube route is the whole answer: it is a genuinely free way to use a frontier video model, which did not exist a year ago. The catch is that the free surfaces are built for remixing short-form video, not for producing assets to spec. You are working inside YouTube's flow, on YouTube's aspect ratios, with YouTube's templates — so it is excellent for learning what the model can do and poor for building a deliverable.

Why the free tier is the most interesting thing here. Every other leading video model gives you a credit allowance that runs out, or a watermark, or both. Putting a frontier model inside YouTube Shorts at no cost is a distribution decision rather than a pricing one — Google is buying habit. If you want to understand the model before paying for it, do your testing there and treat the API as the production step.

What it actually costs

Google's developer announcement gives the headline number plainly: $0.10 per second of video output. The API page explains the mechanism behind it, and the mechanism is what matters, because it tells you where the money goes.

Billing is per token, not per second. Input — text, images, video and audio you send — is $1.50 per million tokens. Video output is $17.50 per million tokens, and one second of 720p video is 5,792 output tokens. That multiplies out to roughly $0.1014 per second, which is about $1.01 for a 10-second clip.

What you generateApproximate costWhy
10-second clip at 720pabout $1.01The baseline. 5,792 tokens per second × 10 seconds at $17.50 per million.
10-second draft at 360pabout $0.34The 360p fast mode costs roughly one third of 720p and runs up to 60 percent faster. This is the single biggest lever on your bill.
A 40-second sceneabout $4.06 outputFour chained requests, not one pass — four generations of 10 seconds each. Every turn also re-bills the context you carry forward as input.
1080p and 4Khigher than 720pThese tiers exist from the 1.1 release. Reported figures put a 10-second clip around $1.50 at 1080p and $3.00 at 4K, but treat those as third-party reporting rather than a published rate card — and note that 1080p and 4K are upscales of the generated frames rather than native renders.
The one habit that changes the economics. Video prompting takes retries — a realistic session is a dozen attempts before you keep one shot. If all twelve run at 720p you spend about $12 to arrive at a clip worth $1. If you draft all twelve at 360p and only render the winner at 720p, the same result costs about $5. That is a 60 percent saving for changing one setting, and it is the reason Google built the 360p mode at all. Draft cheap, finish once.

Gemini Omni Flash vs Gemini Omni 1.1 Flash

There are two model IDs in circulation and one of them no longer works. This is the difference.

Gemini Omni Flash (preview)Gemini Omni 1.1 Flash (GA)
Announced19 May 2026, at Google I/O—
API statusPublic preview from 30 June 2026 — endpoint retired 30 September 2026Generally available since 27 August 2026
Model IDgemini-omni-flash-previewgemini-omni-1.1-flash
Resolution720p native, no higher tierAdds 360p fast mode for drafting, plus 1080p and 4K output
Scene extensionNot availableIn 10-second increments to about 40 seconds cumulative
Frame controlPrompt onlyFirst and last frame can both be specified; the model generates what connects them
Reference inputImages and voiceAdds up to three seconds of video reference

The practical consequence of the retirement: anything still wired to gemini-omni-flash-preview stopped working on 30 September 2026. The fix is a model-ID change plus a check on the new resolution parameter, which did not exist when the preview shipped. If you inherited an integration from the summer, that is the first thing to look at.

It is also worth reading the 1.1 release as a cost-control release rather than a quality release. The headline additions — a cheap draft resolution, a way to pin the first and last frame so you stop burning generations on camera moves, a longer context window so extensions stop drifting — all reduce the number of paid attempts a shot needs. Google moved the spend from the search to the final render.

Where it ranks — and why we report two answers

Gemini Omni arrived at the top of the boards and then the boards were rebuilt underneath it. Both of the following are accurate, which is why we show them together instead of picking the flattering one.

BoardSnapshotResult
Arena AI — text-to-video, blind human voting22 September 2026First and second. gemini-omni-1.1-flash at an Elo of 1,516, the original gemini-omni-flash at 1,513. It had debuted first on the same arena back on 2 July 2026 with a 101-point margin.
Artificial Analysis — rebuilt text-to-video board30 September 2026Eighth, both with audio (1,113) and silent (1,070). The board was rebuilt on a new scale at the end of September, so its earlier scores are not comparable.

The reconciliation is that these measure different things on different prompt sets, and a rebuilt board with a new scale will always reshuffle ranks. What is worth taking from it is narrower and more useful: Gemini Omni's lead lasted about ten weeks, and it no longer has one. It went from sweeping the boards in July to mid-table on one of them in September, while competitors kept shipping. Judge it on what it does for your workflow — conversational editing, native audio, text rendering — rather than on a number that moved.

What nobody has published

Google is more forthcoming about Gemini Omni than ByteDance is about Seedance, but there are still gaps, and they are easy to trip over because launch coverage tends to fill them in silently.

What you would expectWhat actually exists
Published benchmark scoresNot yet, and Google says so. The model card states that evaluations for text-to-video-with-audio, image-to-video-with-audio, reference-to-video-with-audio, video editing and image generation will be shared "when we roll out to developers and enterprise customers via APIs". The third-party board scores above are the only numbers in circulation.
A clear per-second price at 1080p and 4KNot clearly published. The $0.10 per second headline is for 720p. Higher tiers exist from 1.1 and cost more, but the figures circulating come from third-party write-ups rather than a rate card you can check.
Whether 1080p and 4K are nativeThey are described as upscales of the generated frames rather than native renders — which materially changes what you are paying for. Confirm on the pricing page before building a 4K pipeline around this.
A Gemini Omni ProAnnounced as teased, no date. A higher-end tier has been referenced repeatedly, the same way Gemini 3.5 Pro was. Nothing has shipped.
Length beyond roughly 40 secondsNot demonstrated. Generation is 3–10 seconds; extension chains to about 40. Google's own framing of longer durations is "coming soon", which for planning purposes means: cut in an editor.

There is also one genuine contradiction in Google's own launch material worth knowing about, because you will hit it if you go looking. The I/O announcement says Google will "soon support output modalities like image and audio", which reads as though audio does not exist yet. The developer announcement says Gemini Omni Flash "natively generates audio with every video output". The model card resolves it: the output is video with audio. Sound arrives attached to a video; it is not available as a separate output you can request on its own, and audio input at launch was limited to voice references.

Getting a usable clip

Most of the difference between a good Gemini Omni result and a wasted afternoon is decided before you generate anything.

  1. Learn it on YouTube first.

    Shorts Remix is free and it is the same model. Spend your free time there finding out what this model is good at — conversational edits, text rendering, physics — before you spend a single paid second discovering it is not a cinematography tool.

  2. Draft at 360p, always.

    One third the cost, up to 60 percent faster. Run every experiment at draft resolution and only render the version you are keeping at 720p or above. This is the difference between a $12 session and a $5 one.

  3. Design for 10 seconds, not against it.

    A 10-second ceiling is not a limitation if you write for it. Decide up front whether your shot is one generation or a chain of extensions; each extension is a separate paid request that re-bills the context, so a careless 40-second scene costs four times what you budgeted.

  4. Pin the frames when the camera moves.

    The 1.1 release lets you specify both the first and the last frame and have the model generate what connects them. That is how you get an orbit, a zoom or a clean loop without spending twenty generations prompting for it in words.

  5. Edit by asking, not by re-rolling.

    This is the model's actual advantage. Before you regenerate a clip because one element is wrong, try telling it what to change. You keep the scene, the character and the audio, and you pay for one turn instead of a whole new generation.

  6. Label it and plan for the watermark.

    Every output carries SynthID and Content Credentials, verifiable through the Gemini app, Chrome and Google Search. For commercial work, build the assumption into your workflow rather than discovering it after delivery — and check the channel's own AI-disclosure rules before publishing.

Frequently asked questions

What is Gemini Omni?

The family name for Google's any-to-any generative model, announced at Google I/O on 19 May 2026 and described by Google as a model that can create anything from any input, starting with video. The first member is Gemini Omni Flash, which takes text, images, video and voice references in one prompt and returns a short video with synchronised audio. It is not Veo 4 — Veo 3.1 remains Google's Veo model.

Is Gemini Omni free?

Partly, and the free part is unusually good. Google states that Gemini Omni Flash is available at no cost inside YouTube Shorts Remix and the YouTube Create app to users aged 18 and over, and that it is rolling out to AI Plus, Pro and Ultra subscribers globally through the Gemini app and Google Flow. The API has no free tier — your first request is billed.

How much does Gemini Omni cost?

Google's announcement states $0.10 per second of video output, and the API bills by token: $1.50 per million input tokens and $17.50 per million video output tokens, at 5,792 tokens per second of 720p — about $0.1014 per second, or roughly $1.01 for a 10-second clip. A 360p draft costs about a third of that and is up to 60 percent faster. Higher tiers cost more, and revising a prompt also pays for the input tokens carrying the conversation forward.

How long can a Gemini Omni video be?

3 to 10 seconds per generation. The 1.1 release added scene extension, which continues a clip in 10-second steps using up to 10 seconds of preceding footage, to roughly 40 seconds cumulative. That is four chained requests rather than one generation, and each is billed separately.

What is the difference between Gemini Omni Flash and Gemini Omni 1.1 Flash?

Flash is the model launched at I/O on 19 May 2026. Its preview endpoint, gemini-omni-flash-preview, ran from 30 June and was retired on 30 September 2026. The production model is gemini-omni-1.1-flash, generally available since 27 August 2026, which added scene extension, first-and-last-frame control, video reference input, a 360p draft mode and 1080p/4K output tiers. Anything still calling the preview ID stopped working at the end of September.

Does Gemini Omni generate audio?

Yes — Google's model card lists the output as video with audio, and its developer announcement states the model natively generates audio with every video output. What does not exist yet is audio as a standalone output: Google's I/O announcement says image and audio outputs will come later, and audio input at launch was limited to voice references. Sound comes attached to a video, not on its own.

How good is Gemini Omni compared with other video models?

Two reputable boards now disagree, so we report both. On Arena AI's text-to-video snapshot of 22 September 2026, gemini-omni-1.1-flash ranked first and the original second. On Artificial Analysis's rebuilt text-to-video board of 30 September 2026, it ranked eighth with audio and eighth silent. The honest reading: it led decisively in July and no longer does. Judge it on conversational editing, native audio and text rendering rather than on a rank that moved.

What are the limitations of Gemini Omni?

Google's model card names three: consistency across edits, complex motion, and accurate text remain challenging. The 3-to-10-second cap is structural. Image and audio are not available as standalone outputs. And the independent evaluation numbers were absent at launch — the model card says they will be shared when the model reaches developers and enterprise customers through APIs.

Is Gemini Omni the same as Veo?

No. Veo is Google's dedicated video line, currently Veo 3.1, with no announced successor. Gemini Omni is a separate family that Google launched as its any-to-any model, for video first. Both live inside Google Flow and they are priced and prompted differently.

Do Gemini Omni videos carry a watermark?

Yes. Every Omni video carries the SynthID watermark, which is imperceptible, and Google has extended Content Credentials verification across its products so a file's origin can be checked in the Gemini app, Chrome and Google Search. For commercial work, assume the marking is there and confirm the channel's AI-disclosure requirements before you publish.

Also on this site

Disclosure. We are an independent guide and not affiliated with Google, Google DeepMind, Flow, YouTube or any platform listed. Dates, prices and terms on this page were read from Google's own model card, developer announcement, pricing page and I/O announcement, and from third-party leaderboards that are named where used. Prices and availability change frequently; verify before purchase. Where a claim could not be confirmed against a primary source, the page says so. Some links may be affiliate links; they do not change what we recommend or how we rank anything.