Contents
A Gemini Omni prompt has to carry more than an ordinary AI video prompt. Omni makes picture and sound in the same pass, so silence is something you ask for. Quote a line of text and it renders legibly too. Unless you say otherwise, it also decides on its own how many shots your clip has. Instead of approximating motion, it reasons about weight, gravity and fluid behaviour. And once the clip exists, you keep editing it by chat.
Google calls it Nano Banana for video, and the comparison holds. So anyone already writing Nano Banana prompts will find the instinct transfers straight across. Even so, the extra work sits in three places: the soundtrack, the shot count, and the timing. Get those right and even a flat clip starts to look filmed.
Below are 30 prompts, grouped by the capability each one exercises. So copy any of them, swap the subject for yours, and generate. Readers who want the mechanics first should instead start with how to use Gemini Omni, then come back to the library afterwards. Everyone else can dive straight in.
What makes a Gemini Omni prompt different
Six things an ordinary video prompt leaves out, and an Omni prompt should carry:
- The sound. Ambience, dialogue, Foley, or music. Name it, or the model chooses for you.
- Any on-screen words. Quote them exactly and say where they sit in frame.
- The shot count. Omni builds a short multi-shot sequence by default.
- The timing. You can say when things happen, in plain language or in timecode.
- What to leave out. There’s no negative-prompt field, so negatives go in the prompt itself.
- The physical behaviour. Name the force or the material, not the look.
Two limits shape everything you write. First, clips run 3 to 10 seconds. Second, the frame is either 16:9 or 9:16, with no square option. Also keep spoken lines in English, since that is the language Google has fully evaluated.
Prompts that direct the audio
This is the signature. Picture and synchronized sound come out together, so the prompt is where the sound gets decided. Name the source inside the scene rather than a genre in the abstract, because “tyres hissing through water” lands better than “atmospheric”. Music especially has to be asked for, otherwise Omni picks its own bed.
A rain-soaked bus shelter at night, one commuter scrolling a phone, headlights sliding past behind them. Sound: heavy rain drumming on the plastic roof, tyres hissing through water. No dialogue. Close-up of a barista sliding a cup across the counter, steam rising off it. She says, “Careful, it’s hot.” Sound: espresso machine hiss, low café chatter under the line. A watchmaker’s bench, tweezers lowering a tiny gear into a movement, macro lens. Sound: the click of the gear seating, tweezers tapping the tray, a clock ticking somewhere behind. No music. A skate park at golden hour, one rider carving the bowl, camera tracking low and level. Include calm lo-fi music under the whole clip, and keep the board sounds audible over the top of it.Ambience and room tone
A spoken line
Object Foley
Music under the scene
Prompts that render on-screen text
Omni renders quoted words legibly. That is what makes price cards, signage and title cards workable. In practice, keep the string short, quote it exactly, and say where it sits in frame. Also define any text that would appear naturally, such as a shopfront or a number plate, because anything left unspecified gets invented for you.
Overhead shot of a pair of running shoes on wet concrete, camera pushing in slowly. A clean card slides into the lower third reading “£89” and holds there. Sound: light wind, one soft synth note as the card lands. The camera drifts past a small bakery window at dawn. The hand-painted sign on the glass reads “Open since 6”. Inside, a baker sets a tray down. Sound: door bell, distant traffic. A chef looks up from a chopping board and nods at camera. A lower third appears reading “Maya, head chef” and stays for two seconds. Sound: knife on board, extraction fan hum. Waves flattening onto black sand, handheld, slightly underexposed. The words “Cold Water Club” appear centre-frame in thin white type, then the shot carries on. Sound: surf, wind buffeting the mic.A price card
A sign inside the scene
A lower-third name
A title card
Single-shot and multi-shot prompts
Left alone, Omni cuts your clip into a short sequence and builds a small narrative from it. Therefore a single continuous take has to be requested. Three phrasings do it: “in a single unbroken scene”, “in a single continuous shot”, and “no scene cuts”. Each pair below is one idea written both ways, so the phrasing that flips it is easy to spot.
In a single continuous shot, a cyclist rides through a tunnel toward daylight, camera trailing just behind the rear wheel. Sound: tyre hum rising in pitch, echo dropping away as they exit. A three-shot sequence of the same tunnel ride: wide on the entrance, close on the pedals turning, then the rider bursting into daylight. Keep the same bike and jacket in all three. Sound: tyre hum, echo, then open air. One unbroken scene: a florist wraps a bouquet from start to finish while the camera slowly orbits the table. No scene cuts. Sound: paper, scissors, a tap running in the background. A short sequence: hands cutting stems, then paper folding around the bouquet, then the finished wrap handed across the counter. Same hands and same apron in every shot. Sound: scissors, paper, a quiet “thank you”.One take: tunnel ride
Sequence: tunnel ride
One take: florist
Sequence: florist
Prompts with timed events
Call out when something happens and Omni handles it. No special syntax is needed, since plain language works and so does a timecode block. Timecode especially helps when the whole clip is really a mini shot list. Still, one timed event per prompt reads more reliably than three stacked together.
A dancer stands still in an empty studio. After 2 seconds the music kicks in and she starts moving; at 5 seconds the camera pulls back to a wide. Sound: one beat drop, floor squeak. A dust sheet covers something large in a garage. At 3 seconds a hand pulls the sheet away to reveal a restored motorcycle. Sound: fabric sliding off, then a low sustained note underneath. A living room at dusk. At 3 seconds the lamps come on one after another around the room. Sound: three soft clicks spaced apart, evening birds outside the window. [0-3s] A kettle on a stove starts to steam. [3-6s] A hand lifts it and pours into a mug. [6-10s] The camera settles on the mug as steam curls upward. Sound: rising kettle whistle, water pouring, then quiet.A camera move on the beat
A reveal
A light change
A full clip in timecode
Prompts that lean on physics and world knowledge
Omni was built with an intuitive grasp of forces like gravity, kinetic energy and fluid dynamics. On top of that sits Gemini’s knowledge of how things actually work. As a result, it reasons about what should happen next, rather than approximating a motion that looks roughly right. That is exactly where most AI video falls apart. That is why you name the force or the material rather than the look, because “heavy, viscous, slow to settle” gets further than “realistic”.
A stack of wooden blocks on a table is nudged at the base and topples, blocks bouncing and settling with real weight. Locked-off macro shot. Sound: hollow wooden knocks, then stillness. No music. Thick honey pouring off a spoon into a glass bowl, close-up. The stream coils as it lands and folds slowly into itself: heavy, viscous, slow to settle. Sound: a low wet fold. No music. A paper-cutout explainer of how a lock and key work: the key slides in, the pins lift to the shear line, the cylinder turns. Flat overhead view, one continuous shot, mechanically accurate. Sound: paper sliding, a click as it turns.Gravity and collision
A fluid pour
An explainer that has to be correct
Prompts that rule things out
Omni has no negative-prompt field. Instead, anything you want excluded goes into the prompt as an instruction. Three phrasings carry most of the weight: “no dialogue”, “no extra sound effects”, and “no embellishments”. Use them, for example, to stop the model adding a voice, a music bed, or a flourish nobody asked for.
A woman reading on a balcony in late afternoon light, a page turning, the city below her. Sound: distant traffic and the page turn only. No dialogue, no music. Macro of a record needle dropping into a groove and beginning to track. Sound: surface crackle and the first bar of a quiet piano. No extra sound effects. A plain white ceramic mug on a windowsill, camera pushing in slowly as steam rises. Keep it simple and still: no embellishments, no added props, no dialogue.Keeping it silent
Keeping the mix clean
Keeping the frame plain
Chat-edit follow-up prompts
These are second-turn prompts, the ones you type after the clip exists. Each instruction builds on the last. As a result, characters stay consistent, the physics hold, and the scene remembers what came before. That memory therefore turns a single generation into a working session, where the first result is a starting point rather than a verdict.
One rule matters more than the rest, and it inverts everything above: keep edit prompts simple. A long, descriptive edit prompt causes changes nobody asked for. So name one change, leave the rest unstated, and also add “keep everything else the same” to hold the frame steady. Voices and speech are the exception, though, since Omni cannot edit those yet.
Move this scene to a snowy street at night. Keep everything else the same. Change the camera angle to look up at her from waist height. Keep everything else the same. Make this video look like hand-drawn animation. Keep everything else the same. Add a dog asleep under the table. Keep everything else the same. Change the text on the sign to say “Closed Mondays”. Keep everything else the same.Swap the environment
Change the angle
Restyle it
Add something
Rewrite the signage
Prompts that use a reference image
In Picsart, Omni takes one reference image, so these prompts are about describing how a single uploaded frame should move. Don’t re-describe what’s already visible in the picture. Instead, spend the words on the camera move, the subject’s motion, and what’s happening in the environment. “Make it move”, for example, consistently underdelivers.
Using the uploaded product photo, hold the bottle still and move the camera in a slow arc around it while the light shifts from cool to warm. Sound: low room tone, one soft chime as the warm light lands. From the uploaded portrait, have him turn his head toward the window and let his coat settle as he moves. Handheld, shallow depth of field. Sound: fabric shift, quiet street noise. Using the uploaded street photo, keep the buildings fixed and bring the scene to life with pedestrians crossing, a bus passing behind, flags moving in the wind. Sound: traffic, footsteps, a bicycle bell.Animate a product still
Put a portrait in motion
Wake up a still street
Write your own: the Gemini Omni prompt formula
One formula covers most of it:
subject + action + setting + camera + sound + any quoted text + timing + what to leave out
Not every slot earns a place in every prompt. However, the two people skip most often are sound and shot count, and those are the two that make a result feel unfinished. Stay silent about sound, for instance, and it gets chosen for you. Leave shot count alone and you get a sequence instead. Filling both in costs a clause each, so here is a thin prompt next to the same idea written properly:
A man walking a dog in a park. In a single continuous shot, a man walks a wiry-haired terrier along a wet park path at dusk, camera tracking beside them at knee height. Low sun breaks through the trees behind him and the puddles hold the light. At 4 seconds the dog stops to sniff a bin and he waits, the lead going slack. A small park sign in the background reads “No cycling”. Sound: paws and shoes on wet gravel, a dog tag ticking, distant traffic. Keep the expressions detailed and the background busy with real park life. No dialogue, no music.Thin
Filled out
Every slot is doing something there. Specifically, the camera instruction stops it becoming a sequence, the timed beat gives the clip a shape, and the quoted sign keeps background text under control. The closing exclusions then keep a voice and a music bed out. Strip any one of them and the result drifts further from what you pictured.
Omni also responds to instructions about the quality of the output, not just its content. For example, ask for micro-detail in expressions or richer background elements, and the whole scene shifts toward feeling filmed. Moreover, these lines cost nothing and stack with the rest of the formula. Keep them on the end of prompts that matter, then.
What Gemini Omni won’t do
- Clips run 3 to 10 seconds, and the frame is 16:9 or 9:16 only
- You describe the sound you want; you can’t upload your own audio as a reference
- Voices and speech can’t be edited after the fact
- No video extension, and no generating the middle between a first and last frame
- It can’t reference or reason across more than one video at a time
- Every clip carries an invisible SynthID watermark identifying it as AI-generated
Where to use these prompts in Picsart
Google Omni is available in Picsart across three surfaces, so these prompts travel wherever the work already happens. The Picsart AI Playground, the AI Video Generator and Picsart Flow all generate video with synchronized audio, and each also lets you refine it right where you work. Still, pick the surface that matches how finished the idea already is.
First, the AI Playground is for testing a prompt properly. It puts Google Omni behind the same prompt bar as the rest of the video lineup, so one prompt can run through several models. Second, the AI Video Generator is the paste-and-go route: pick Google Omni, drop the prompt in, and choose 16:9 or 9:16. Also add a reference image where the prompt calls for one. Finally, Flow is where a clip stops being a one-off, since it lands the result beside the images, edits and exports it belongs with. For chat edits on any surface, keep the clip in front of you and send the next instruction as a fresh, short prompt.
The Google Omni model page covers what it takes as input, what it sends back, and where it sits alongside Picsart’s other video models. So it is worth a look before committing to a long prompt. The roster, after all, moves faster than most documentation does.
Get answers to common questions
Yes. It produces speech, music and sound effects in the same pass as the video, based on what your prompt describes. Leave sound out of the prompt entirely and it will still generate an audio track, just not the one you had in mind. Therefore naming the sound is always cheaper than regenerating.
Start generating with Google Omni in Picsart
Pick a prompt, change the subject to yours, and run it in the Picsart AI Playground. The first result is rarely the final one, and that is the point. Of course, keep editing by chat until the clip matches what you had in mind. Thirty starting points are certainly plenty to find the two or three that suit the way you work.