Gemini Omni: what it is, how to use it free, and what it really costs
Gemini Omni is the model Google built instead of a Veo 4 — an any-to-any system that starts with video and edits it by conversation. It is also the model with the friendliest free route in AI video and one of the least forgiving paid ones, because every generation has a 10-second ceiling and every request is billed. This page covers both: how to try it at no cost, what a second actually costs at each resolution, and the numbers that Google has not published.
What Gemini Omni is
Gemini Omni is the family name for Google's any-to-any generative model, announced at Google I/O on 19 May 2026. Google's own framing is that it can create anything from any input and that it is starting with video — image and audio outputs are named as later additions rather than shipped features.
The first model in the family is Gemini Omni Flash. It accepts text, images, video and voice references in a single prompt, and returns a short video with synchronised audio. Google's model card describes it as transformer-based with native multimodal support for text, vision, video and audio, trained on TPUs at Google DeepMind.
Three things separate it from the video models it gets compared with:
- It is conversational. Rather than re-rolling the whole clip when you want a change, you ask for the change and the model keeps the scene. Google's phrase is the useful one: change the environment, angle, style or a specific detail "without ever losing the thread of your original scene".
- It reasons about the world before it renders it. Google describes combining an understanding of physics — gravity, kinetic energy, fluid dynamics — with Gemini's general knowledge of history, science and culture. This is why it was launched as a world model rather than a video filter.
- It puts legible text into footage. Google lists rendering text and graphics directly into video, with kinetic typography synced to on-screen movement, as one of its four headline areas. Historically the weakest part of AI video, and the reason explainer clips were unusable.
The free routes — and the catch
This is the part that makes Gemini Omni unusual, and it is stated plainly in Google's own I/O announcement.
- YouTube Shorts Remix — free, to users aged 18 and over. You pick an eligible Short, prompt what you want changed, and get a new version without leaving YouTube. Google also ships a version that puts you into the clip.
- The YouTube Create app — also free, also 18+.
- Gemini app and Google Flow — included with Google AI Plus, Pro and Ultra subscriptions. Not free, but included if you already pay for one of those tiers.
- The API — paid from the first request. There is no free tier, and there is no trial allowance to burn through while you read the docs.
For most people reading this, the YouTube route is the whole answer: it is a genuinely free way to use a frontier video model, which did not exist a year ago. The catch is that the free surfaces are built for remixing short-form video, not for producing assets to spec. You are working inside YouTube's flow, on YouTube's aspect ratios, with YouTube's templates — so it is excellent for learning what the model can do and poor for building a deliverable.
What it actually costs
Google's developer announcement gives the headline number plainly: $0.10 per second of video output. The API page explains the mechanism behind it, and the mechanism is what matters, because it tells you where the money goes.
Billing is per token, not per second. Input — text, images, video and audio you send — is $1.50 per million tokens. Video output is $17.50 per million tokens, and one second of 720p video is 5,792 output tokens. That multiplies out to roughly $0.1014 per second, which is about $1.01 for a 10-second clip.
| What you generate | Approximate cost | Why |
|---|---|---|
| 10-second clip at 720p | about $1.01 | The baseline. 5,792 tokens per second × 10 seconds at $17.50 per million. |
| 10-second draft at 360p | about $0.34 | The 360p fast mode costs roughly one third of 720p and runs up to 60 percent faster. This is the single biggest lever on your bill. |
| A 40-second scene | about $4.06 output | Four chained requests, not one pass — four generations of 10 seconds each. Every turn also re-bills the context you carry forward as input. |
| 1080p and 4K | higher than 720p | These tiers exist from the 1.1 release. Reported figures put a 10-second clip around $1.50 at 1080p and $3.00 at 4K, but treat those as third-party reporting rather than a published rate card — and note that 1080p and 4K are upscales of the generated frames rather than native renders. |
Gemini Omni Flash vs Gemini Omni 1.1 Flash
There are two model IDs in circulation and one of them no longer works. This is the difference.
| Gemini Omni Flash (preview) | Gemini Omni 1.1 Flash (GA) | |
|---|---|---|
| Announced | 19 May 2026, at Google I/O | — |
| API status | Public preview from 30 June 2026 — endpoint retired 30 September 2026 | Generally available since 27 August 2026 |
| Model ID | gemini-omni-flash-preview | gemini-omni-1.1-flash |
| Resolution | 720p native, no higher tier | Adds 360p fast mode for drafting, plus 1080p and 4K output |
| Scene extension | Not available | In 10-second increments to about 40 seconds cumulative |
| Frame control | Prompt only | First and last frame can both be specified; the model generates what connects them |
| Reference input | Images and voice | Adds up to three seconds of video reference |
The practical consequence of the retirement: anything still wired to gemini-omni-flash-preview stopped working on 30 September 2026. The fix is a model-ID change plus a check on the new resolution parameter, which did not exist when the preview shipped. If you inherited an integration from the summer, that is the first thing to look at.
It is also worth reading the 1.1 release as a cost-control release rather than a quality release. The headline additions — a cheap draft resolution, a way to pin the first and last frame so you stop burning generations on camera moves, a longer context window so extensions stop drifting — all reduce the number of paid attempts a shot needs. Google moved the spend from the search to the final render.
Where it ranks — and why we report two answers
Gemini Omni arrived at the top of the boards and then the boards were rebuilt underneath it. Both of the following are accurate, which is why we show them together instead of picking the flattering one.
| Board | Snapshot | Result |
|---|---|---|
| Arena AI — text-to-video, blind human voting | 22 September 2026 | First and second. gemini-omni-1.1-flash at an Elo of 1,516, the original gemini-omni-flash at 1,513. It had debuted first on the same arena back on 2 July 2026 with a 101-point margin. |
| Artificial Analysis — rebuilt text-to-video board | 30 September 2026 | Eighth, both with audio (1,113) and silent (1,070). The board was rebuilt on a new scale at the end of September, so its earlier scores are not comparable. |
The reconciliation is that these measure different things on different prompt sets, and a rebuilt board with a new scale will always reshuffle ranks. What is worth taking from it is narrower and more useful: Gemini Omni's lead lasted about ten weeks, and it no longer has one. It went from sweeping the boards in July to mid-table on one of them in September, while competitors kept shipping. Judge it on what it does for your workflow — conversational editing, native audio, text rendering — rather than on a number that moved.
What nobody has published
Google is more forthcoming about Gemini Omni than ByteDance is about Seedance, but there are still gaps, and they are easy to trip over because launch coverage tends to fill them in silently.
| What you would expect | What actually exists |
|---|---|
| Published benchmark scores | Not yet, and Google says so. The model card states that evaluations for text-to-video-with-audio, image-to-video-with-audio, reference-to-video-with-audio, video editing and image generation will be shared "when we roll out to developers and enterprise customers via APIs". The third-party board scores above are the only numbers in circulation. |
| A clear per-second price at 1080p and 4K | Not clearly published. The $0.10 per second headline is for 720p. Higher tiers exist from 1.1 and cost more, but the figures circulating come from third-party write-ups rather than a rate card you can check. |
| Whether 1080p and 4K are native | They are described as upscales of the generated frames rather than native renders — which materially changes what you are paying for. Confirm on the pricing page before building a 4K pipeline around this. |
| A Gemini Omni Pro | Announced as teased, no date. A higher-end tier has been referenced repeatedly, the same way Gemini 3.5 Pro was. Nothing has shipped. |
| Length beyond roughly 40 seconds | Not demonstrated. Generation is 3–10 seconds; extension chains to about 40. Google's own framing of longer durations is "coming soon", which for planning purposes means: cut in an editor. |
There is also one genuine contradiction in Google's own launch material worth knowing about, because you will hit it if you go looking. The I/O announcement says Google will "soon support output modalities like image and audio", which reads as though audio does not exist yet. The developer announcement says Gemini Omni Flash "natively generates audio with every video output". The model card resolves it: the output is video with audio. Sound arrives attached to a video; it is not available as a separate output you can request on its own, and audio input at launch was limited to voice references.
Getting a usable clip
Most of the difference between a good Gemini Omni result and a wasted afternoon is decided before you generate anything.
-
Learn it on YouTube first.
Shorts Remix is free and it is the same model. Spend your free time there finding out what this model is good at — conversational edits, text rendering, physics — before you spend a single paid second discovering it is not a cinematography tool.
-
Draft at 360p, always.
One third the cost, up to 60 percent faster. Run every experiment at draft resolution and only render the version you are keeping at 720p or above. This is the difference between a $12 session and a $5 one.
-
Design for 10 seconds, not against it.
A 10-second ceiling is not a limitation if you write for it. Decide up front whether your shot is one generation or a chain of extensions; each extension is a separate paid request that re-bills the context, so a careless 40-second scene costs four times what you budgeted.
-
Pin the frames when the camera moves.
The 1.1 release lets you specify both the first and the last frame and have the model generate what connects them. That is how you get an orbit, a zoom or a clean loop without spending twenty generations prompting for it in words.
-
Edit by asking, not by re-rolling.
This is the model's actual advantage. Before you regenerate a clip because one element is wrong, try telling it what to change. You keep the scene, the character and the audio, and you pay for one turn instead of a whole new generation.
-
Label it and plan for the watermark.
Every output carries SynthID and Content Credentials, verifiable through the Gemini app, Chrome and Google Search. For commercial work, build the assumption into your workflow rather than discovering it after delivery — and check the channel's own AI-disclosure rules before publishing.
Frequently asked questions
What is Gemini Omni?
The family name for Google's any-to-any generative model, announced at Google I/O on 19 May 2026 and described by Google as a model that can create anything from any input, starting with video. The first member is Gemini Omni Flash, which takes text, images, video and voice references in one prompt and returns a short video with synchronised audio. It is not Veo 4 — Veo 3.1 remains Google's Veo model.
Is Gemini Omni free?
Partly, and the free part is unusually good. Google states that Gemini Omni Flash is available at no cost inside YouTube Shorts Remix and the YouTube Create app to users aged 18 and over, and that it is rolling out to AI Plus, Pro and Ultra subscribers globally through the Gemini app and Google Flow. The API has no free tier — your first request is billed.
How much does Gemini Omni cost?
Google's announcement states $0.10 per second of video output, and the API bills by token: $1.50 per million input tokens and $17.50 per million video output tokens, at 5,792 tokens per second of 720p — about $0.1014 per second, or roughly $1.01 for a 10-second clip. A 360p draft costs about a third of that and is up to 60 percent faster. Higher tiers cost more, and revising a prompt also pays for the input tokens carrying the conversation forward.
How long can a Gemini Omni video be?
3 to 10 seconds per generation. The 1.1 release added scene extension, which continues a clip in 10-second steps using up to 10 seconds of preceding footage, to roughly 40 seconds cumulative. That is four chained requests rather than one generation, and each is billed separately.
What is the difference between Gemini Omni Flash and Gemini Omni 1.1 Flash?
Flash is the model launched at I/O on 19 May 2026. Its preview endpoint, gemini-omni-flash-preview, ran from 30 June and was retired on 30 September 2026. The production model is gemini-omni-1.1-flash, generally available since 27 August 2026, which added scene extension, first-and-last-frame control, video reference input, a 360p draft mode and 1080p/4K output tiers. Anything still calling the preview ID stopped working at the end of September.
Does Gemini Omni generate audio?
Yes — Google's model card lists the output as video with audio, and its developer announcement states the model natively generates audio with every video output. What does not exist yet is audio as a standalone output: Google's I/O announcement says image and audio outputs will come later, and audio input at launch was limited to voice references. Sound comes attached to a video, not on its own.
How good is Gemini Omni compared with other video models?
Two reputable boards now disagree, so we report both. On Arena AI's text-to-video snapshot of 22 September 2026, gemini-omni-1.1-flash ranked first and the original second. On Artificial Analysis's rebuilt text-to-video board of 30 September 2026, it ranked eighth with audio and eighth silent. The honest reading: it led decisively in July and no longer does. Judge it on conversational editing, native audio and text rendering rather than on a rank that moved.
What are the limitations of Gemini Omni?
Google's model card names three: consistency across edits, complex motion, and accurate text remain challenging. The 3-to-10-second cap is structural. Image and audio are not available as standalone outputs. And the independent evaluation numbers were absent at launch — the model card says they will be shared when the model reaches developers and enterprise customers through APIs.
Is Gemini Omni the same as Veo?
No. Veo is Google's dedicated video line, currently Veo 3.1, with no announced successor. Gemini Omni is a separate family that Google launched as its any-to-any model, for video first. Both live inside Google Flow and they are priced and prompted differently.
Do Gemini Omni videos carry a watermark?
Yes. Every Omni video carries the SynthID watermark, which is imperceptible, and Google has extended Content Credentials verification across its products so a file's origin can be checked in the Gemini app, Chrome and Google Search. For commercial work, assume the marking is there and confirm the channel's AI-disclosure requirements before you publish.
Also on this site
- Seedance 2.5 — ByteDance's video model: 30-second single generations, 50 reference assets, and the four specifications no engineering source confirms. The other half of the current top tier.
- Nano Banana 2.1 — Google's image model, and the other half of this family's story: what changed, what it costs, and how it compares with the Pro tier.
- MiniMax H3 — the model that beats Gemini Omni on Arena's image-to-video board and trails it on text-to-video. One of the few frontier video models you can download and run yourself.
- GPT Image 2.5 — the image-side counterpart: two API models at one price, native transparent backgrounds, and a free route this model does not have.
- Kling 4.0 — the competitor promising 30-second generations. What Kuaishou has shipped versus what it has only announced.
- Higgsfield AI review — the aggregator route: what each plan grants, what a clip really costs, and which surfaces bill even on an unlimited model.