Contents
Gemini Omni is Google DeepMind’s unified multimodal video model, and it does something most video models cannot: it generates a video with sound from a prompt, then lets you keep editing that video by conversation, one plain-English instruction at a time. The easiest way to use it is on Picsart, which offers it as Google Omni, with no install and no API setup: open the Picsart AI Video Generator, select Google Omni, describe the shot you want, and generate.
This guide covers what Gemini Omni is, how to make your first video step by step, how to edit it by chat, how to write prompts, and what you can create. By the end, you will have a finished, sound-on clip you can refine without starting over.
What is Gemini Omni?
Gemini Omni is Google’s native multimodal video model, offered on Picsart as Google Omni. “Multimodal” is the key word: it takes text, image, video, and audio as input, and it generates video with synchronized audio in a single pass. Where most models hand back a silent clip, Gemini Omni delivers the scene and its soundtrack together, including dialogue, ambient sound, and true-to-life sound effects.
Its signature is conversational editing. After a video is generated, you refine it by chat, and the model changes only what you ask while keeping the rest of the scene stable, turn after turn. It is also a world model, with an intuitive grasp of physics like gravity and fluid dynamics, which helps motion and materials behave believably.
On Picsart, Gemini Omni arrives across four places creators already work: the AI Playground, the AI Video Generator, the AI Video Editor, and Picsart Flow. You can also run it against other video models on the same prompt. For the full rundown, see the Gemini Omni launch in Picsart.
Gemini Omni at a glance
| Feature | Detail |
|---|---|
| Developer | Google DeepMind |
| Type | Unified multimodal video model |
| Inputs | Text, image, video, audio |
| Editing | Conversational, multi-turn; edits uploaded video |
| Audio | Synchronized in one pass; lip-sync in 6 languages; Foley |
| Clip length | Up to 10 seconds |
| Aspect ratios | 16:9 and 9:16 |
| Reference images | Up to 5 |
| Strengths | Physics-aware motion, multi-shot storytelling, on-screen text |
| Best for | Edit-after-generate, storytelling, marketing content |
How to use Gemini Omni on Picsart, step by step
You can make a Gemini Omni video in a few minutes, with no account juggling or setup. Here is the full flow:
How to use Gemini Omni on Picsart, step by step
1. Open the AI Video Generator
Open the Picsart AI Video Generator or AI Playground.
2. Select Google Omni
Select Google Omni as the model.
3. Write your prompt
Type a prompt describing the scene, the action, and the sound you want.
4. Set aspect ratio and references
Set the aspect ratio, and add a photo or reference image if you want to guide the result.
5. Generate with audio
Generate the video, with audio, in one pass.
6. Refine by chat or download
Refine it by chat with follow-up instructions, or download it.
Everything runs on one credit balance, and it is free to try, so you can experiment before committing.
Generate your first Gemini Omni video: a worked example
The fastest way to learn Gemini Omni is to make something, then shape it. Say you want a cozy scene. Start with a prompt that names the subject, the setting, the action, and the sound:
“A tabby cat curled on a sunny windowsill, tail flicking slowly as birds chirp outside, warm afternoon light, gentle ambient sound.”
Select Google Omni, set the aspect ratio, and generate. In a few minutes you get a short clip with the motion and ambient sound already mixed in. Now the part that sets Gemini Omni apart: instead of re-writing the prompt, you talk to the video. Type “add a soft blanket under the cat,” then “make the light golden hour,” then “add quiet piano in the background.” Each instruction changes only what you named and leaves the rest intact, so the clip evolves without ever starting over.
How to edit videos with Gemini Omni
Conversational editing is the reason to reach for Gemini Omni. Once a clip exists, you refine it by describing changes in plain English, and the model rewrites only the affected frames while keeping the rest of the scene stable. You can:
- Change the look: swap the environment, adjust the camera angle, shift the style, or change the lighting.
- Change the content: add new characters or objects, or change what is happening in the shot.
- Apply effects: ask for physics-based touches, like a surface rippling like liquid.
Because it is stateful, each edit builds on the one before, so characters stay consistent, the physics hold up, and the scene remembers what came before as you work across several turns. A useful habit: keep edit prompts simple, and add “keep everything else the same” when you want to change one thing and leave the rest untouched. Gemini Omni can also edit videos you upload yourself, not only ones it generated.
How to write Gemini Omni prompts
Gemini Omni follows a clear, structured prompt closely. A few habits get better results:
- Describe the whole shot: the subject, the action, the setting, the camera, and the mood.
- Direct the audio: name the dialogue, ambient sound, or music you want, since Gemini Omni generates sound too.
- Control the shots: by default it builds a short multi-shot sequence, so add “in a single continuous shot” if you want one unbroken take.
- Time events: you can call out timing in plain language, such as “after 3 seconds, the camera pulls back.”
- Put on-screen text in quotes: it renders readable text, so quote the exact words you want to appear.
For editing, do the opposite and keep it simple: short instructions work best, and overly detailed edit prompts can cause unintended changes.
Combine images, video, and audio as references
One of Gemini Omni’s most powerful moves is building a video from things you already have. It reads references across formats, image, video, and audio, and blends them into one cohesive clip:
- Start from photos. Bring up to 5 images of a character, product, or scene into a generation to guide the result.
- Borrow motion. Point Gemini Omni at a video and ask it to apply that movement or camera style to a new subject.
- Guide with a sketch. Use a drawing as a movement guide, then ask for realistic footage that follows it without showing the drawing itself.
- Match a style. Hand it a reference image and have the model carry that look across the whole clip.
Describe how each reference should be used, and Gemini Omni combines them into a single, consistent video, the heart of its create-anything-from-any-input design.
What you can create with Gemini Omni
Gemini Omni is built for finished, sound-on video you can keep shaping:
- Multi-shot stories and explainers, built as a short sequence rather than a single take.
- Talking videos, with dialogue and lip-sync across six languages.
- Marketing and product clips, refined by chat until the details are right.
- Image to video, turning up to five photos into a moving scene.
- Edited and restyled clips, including reworking videos you upload.
Tips for the best results
A few habits get more out of every generation:
For a change, use a follow-up chat instruction instead of a new prompt. Simple instructions land cleaner than long, detailed ones. when you want to change one element and hold the rest. Name the dialogue, ambience, or music, since audio is part of what it makes. Add up to 5 photos to guide a subject, character, or product. Gemini Omni is Google DeepMind’s unified multimodal video model, offered on Picsart as Google Omni. It generates video with synchronized audio in one pass and lets you edit the result by conversation, keeping the rest of the scene stable.Tips for the best results
Edit, do not restart
Keep edit prompts short
Add "keep everything else the same"
Direct the sound
Use references
Get answers to common questions
Start creating with Gemini Omni on Picsart
Using Gemini Omni is as simple as describing a shot, then talking to the result until it lands. Open the Picsart AI Video Generator, select Google Omni, and make your first sound-on clip in a few minutes.