Google introduced Gemini Omni, a multimodal model that generates and edits video from almost any input, at its I/O developer conference on May 19, 2026, moving the company's generative-video effort out of the standalone Veo line and into the core Gemini system. That's not a minor product update. It's an architectural shift in how Google thinks about generative AI.
Until now, Google ran a split stack: Veo for video, Imagen for images, and separate systems for audio. Omni collapses that into one model that can reason across modalities. In practice, that translates to more coherent edits and fewer pipeline artifacts.
Google describes Omni as the point where "Gemini's ability to reason meets the ability to create." That framing is deliberate. This isn't just a video generator bolted onto an LLM. It's a unified system where understanding and generation happen inside the same weights.
What Is Gemini Omni
Gemini Omni is Google DeepMind's first natively multimodal generative media model. The first variant in the family is Gemini Omni Flash, which is now live. Omni accepts any combination of text, images, audio, and video as input and produces a video. The key here is that there's no relay happening across different systems — this is all one model.
Google's AI portfolio now includes Omni, a world model designed to simulate physical environments, predicting what happens next based on a user's actions. That "world model" framing is worth paying attention to. It signals that Google isn't just targeting content creators. It's targeting anyone who needs a system that understands causality, physics, and context well enough to generate believable outputs.
Gemini Omni Flash is the first version to debut, with a Pro model to follow later.
How Gemini Omni Works
The core mechanic is conversational, multi-turn video editing. Users can combine images, audio, video, and text in a single prompt. Rather than stitching those inputs together, the model reasons across them to produce one output and then accepts further changes through conversation.
Every edit you make builds on the one before, maintaining a consistent, coherent scene. Gemini Omni combines an intuitive understanding of physics with Gemini's knowledge of history, science, and cultural context.
Every conversation with the model layers changes and transformations according to the last request. This allows users to change specific details or broader visual elements. The model also takes into account the physics and consequences of requests, allowing users to change the environment, angle, style, and action, as well as add new characters, objects, and details.
Key Technical Highlights
- With Omni, you can combine images, audio, video, and text as input and generate high-quality videos grounded in Gemini's real-world knowledge.
- Omni Flash improves character consistency, meaning identity and voice are preserved across every scene.
- Omni has an improved intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics, allowing you to create more realistic scenes.
- Every Omni output ships with two layers of provenance. SynthID is an invisible watermark embedded directly into the pixels at generation time — imperceptible to viewers, designed to survive cropping, filters, and re-encoding. C2PA content credentials sit alongside it as a signed cryptographic manifest attached to the file.
- Flash clips are capped at 10 seconds, a deployment decision rather than a model constraint — a way to widen access while compute demand is high.
- Omni is also coming to Google Flow Music, allowing users to work conversationally with the agent to direct shareable music videos. With Omni Flash you can guide the styles, subjects, and scenes to match the narrative and pacing of your track.
Availability and Pricing
Gemini Omni Flash is rolling out to all Google AI Plus, Pro, and Ultra subscribers globally through the Gemini app and Google Flow. It's also rolling out at no cost to users on YouTube Shorts and YouTube Create App starting this week. In the coming weeks, it will also be rolling out to developers and enterprise customers via APIs.
Google AI Ultra subscriptions now start at $100/month. Subscribers get higher access to advanced Gemini models and powerful features like video generation with Gemini Omni. The Ultra tier also includes 20TB of cloud storage and YouTube Premium.
Credit allocations scale with the tier: Plus gets 200 monthly AI credits, Pro gets 1,000. Developer API access is confirmed but not yet live at launch.
Safety and Content Provenance
Google put meaningful effort into the safety architecture here. Automated red teaming was used to dynamically evaluate Gemini Omni Flash for safety and security considerations at scale, complementing human red teaming and static evaluations. Ethics and safety reviews were conducted ahead of the model's release.
Google said at I/O that SynthID has now marked more than 100 billion AI-generated images and videos, and that OpenAI, ElevenLabs, and Kakao are adopting the standard. That's a notable signal. If SynthID becomes an industry-wide provenance layer, it shifts from a Google-specific feature into actual infrastructure.
Google DeepMind product management director Nicole Brichtova told TechCrunch the avatar onboarding requires recording yourself and speaking a series of numbers aloud. The avatar is then stored for reuse, an anti-deepfake step modeled loosely on the Cameos feature from OpenAI's now-discontinued Sora app.
How It Stacks Up Against Competitors
Omni enters a crowded field. ByteDance's Seedance 2.0 has led public quality benchmarks, and Kling 3.0 remains dominant in the Chinese market. Independent testers have suggested Flash's raw generation quality may trail those competitors even if its conversational editing is stronger.
Google's answer to that gap is distribution, not raw output quality. Google's strategic edge is distribution: Omni ships inside Search, the Gemini app, Flow, and YouTube rather than as a standalone product.
Google's distribution includes 2 billion daily Gemini users, YouTube integration, Google Workspace, Android, and Vertex AI for enterprise. No other company has that full-stack leverage. That matters more than benchmark scores when you're trying to get a model in front of creators at scale.
Final Thoughts
The most technically interesting thing about Omni isn't any single capability. It's the architectural decision to unify generation and reasoning into one model rather than chaining specialized systems. Its editing-first philosophy and unified multimodal architecture represent a genuinely different bet on AI video's future — one that prioritizes workflow over raw generation quality. Whether that bet pays off depends on whether users actually want iterative, conversational editing or whether they just want the best-looking output on the first try.
The 10-second clip cap and the absence of formal public benchmarks at launch are things I'd watch closely. Several technical claims circulating alongside the launch are not confirmed by Google. A widely shared explainer describes Omni Flash output as capped at 720p and quotes generation times of roughly 60 to 90 seconds per clip, but neither figure appears in Google's official materials. Until Google publishes numbers on VBench 2.0 or the Artificial Analysis Video Arena, the quality story stays incomplete.
What's clear is that Google is done treating video as a separate product. Omni is the bet that one model, deeply integrated across Search, YouTube, and the Gemini app, beats a collection of best-in-class specialists. I'll be watching how developers use the API once it opens up. What do you think? Drop your thoughts in the comments.
Frequently Asked Questions
5 questions
1What is Gemini Omni?
Gemini Omni is Google DeepMind's first natively multimodal generative media model. It accepts text, images, audio, and video as input and generates video output, all within a single unified model rather than a chain of separate systems.
2How is Gemini Omni different from Veo?
Veo was a standalone video generation model. Omni integrates video generation directly into the core Gemini system, enabling multi-turn conversational editing where each instruction builds on the previous one, with persistent character and scene consistency.
3Who can access Gemini Omni Flash right now?
Gemini Omni Flash is available to Google AI Plus, Pro, and Ultra subscribers via the Gemini app and Google Flow. YouTube Shorts and YouTube Create App users get access for free. Developer API access is coming in the following weeks.
4How does Google handle deepfakes and content authenticity in Omni?
Every video generated with Omni carries an imperceptible SynthID watermark and C2PA content credentials. For personal avatars, users must record themselves and speak a series of numbers aloud during onboarding, an anti-deepfake verification step.
5Does Gemini Omni beat Seedance 2.0 in video quality?
Not definitively, based on current information. Independent testers suggest Omni Flash's raw generation quality may trail Seedance 2.0, though its conversational editing workflow is considered stronger. No formal public benchmark results have been published by Google as of launch.






