Gemini Omni 1.2 Flash is available in Picsart, and its headline number is 4K: this release raises both generation and editing to that ceiling, where the previous one stopped at 1080p. Picsart had early access, so the model arrived here as soon as Google shipped it.
Resolution is the news, but it is not the reason Omni is unusual. Most AI video models produce a silent clip that somebody else has to score, caption, and cut. Omni was built to hand back something closer to finished, and the sections below cover the three behaviors that make it feel that way.
Sound arrives with the picture, not after it
Audio is generated alongside the frames rather than layered on once the render completes. There is no separate voice engine to line up and no effects library to borrow from.
That shared origin is what makes the output cohere. Footsteps land on the frame where the foot lands. Spoken lines match the mouth shapes producing them. Room tone belongs to the space on screen and stays there as the shot moves. Speech, music, and atmosphere all trace back to the same instruction that described the picture, and it is why the model takes no audio file at all: whatever should be heard belongs in the written brief.
Revisions happen in plain language
A finished clip can be changed by describing the change. No timeline, no masking, no starting over.
The instructions really are that ordinary. Swap a red car for a black one. Take a logo off a shirt. Recolor a jacket. Make a delivered line sound more apologetic. Omni rewrites the frames that the request touches and leaves the rest pixel-stable, so whatever was already correct stays correct. This is the capability with no equivalent elsewhere in Picsart, and 4K is what promotes it from a drafting trick to something usable on a final master, since the revised frames come back at full size.
Text inside the frame stays readable
Type is where AI video usually gives itself away, with letters that warp between frames or spell a word two ways in four seconds. Omni holds typography steady instead.
Captions keep their shape. A call to action on an ad stays legible at the end of the clip as well as the start. An equation on a blackboard, a label in a product demo, a lower third on an explainer: all of it survives the motion. Video that has to carry words stops requiring a separate design pass to fix them.
Who the release is for
- Anyone delivering above 1080p. Paid placements, homepage banners, event screens, and broadcast inserts routinely specify more, and this release finally satisfies those specs.
- Editors working in one aspect ratio and shipping in two. A 4K landscape render has enough detail to yield a sharp vertical crop from inside the same frame.
- Teams grading or stabilizing footage. Those passes eat detail, and a 4K source has enough to give away.
- Anyone who needs sound on the first pass. Dialogue, atmosphere, and effects come out of the same generation as the picture.
- People revising to a brief rather than a whim. The plain-language edit loop suits work that gets notes and needs several rounds.
Alongside 4K, a single generation accepts up to five reference images and three reference videos, takes independent start and end frames, and can extend footage you already shot. Clips run at any whole second between three and ten, in landscape or portrait.
Hear the difference on one clip
The fastest way to judge a one-pass model is to close your eyes for the second half. Generate this, then listen to whether the sound belongs to the picture or merely accompanies it.
One-pass audio test, Gemini Omni 1.2 Flash
A blacksmith brings a hammer down on glowing metal in a dim workshop, sparks scattering across the anvil, bellows breathing in the background. Each hammer strike rings out and decays into the stone room, the fire hisses when the metal is quenched, and a low murmur of wind passes the open door between blows
Where to find it
Gemini Omni 1.2 Flash sits in Picsart AI Playground, listed under its full name in the model menu.
The Playground puts 176 AI models from 32 providers behind a single prompt bar, so weighing Omni against a competing model costs a menu change rather than a second subscription, and one credit balance funds all of them with per-model pricing shown up front. Background on the wider family sits on the Google Omni model page, and the video models directory maps the rest of the roster.
Get answers to common questions
Google’s unified multimodal model, which emits video together with matching audio in one pass instead of combining two separate systems. This release raises generation and editing to 4K.
Put it to work on something real
Pick a clip that needs to arrive with its own sound and land above 1080p. Open Picsart AI Playground and start there.