Flux 3 prompts have to carry the soundtrack. The model makes picture and sound in one pass. So every line of description also tells the clip what to sound like. Say nothing about the audio and Flux 3 chooses it for you.

That single fact changes the whole job. Most AI video prompt lists target models that hand back silent footage. So those prompts transfer badly. Instead, a Flux 3 prompt names the room tone, the impact, the line of dialogue. It also gets twenty seconds to work with instead of the five or ten most models allow. That is long enough for a character to walk in, do something, and say something.

The other thing that changes is where the prompt starts. Flux 3 Video takes three kinds of input, and the prompt looks different for each one. Start from words alone, start from reference images, or start from a clip you already have and keep it going. So the 25 prompts below follow that split. The mode decides what the prompt still needs to say, and what the model can already see for itself.

Copy any of them, swap in the subject you care about, and generate.

What makes a Flux 3 prompt different

A Flux 3 prompt has four jobs that a prompt for a silent model never has to do:

  • Decide the audio. Room tone, impacts, music, or a spoken line. Name what makes the noise, or the model picks something for you.
  • The length of the beat. Twenty seconds holds a whole small action. Say what fills it.
  • The number of shots. Flux 3 can cut between angles inside one generation, so state whether you want one continuous take or several.
  • Any words on screen. Reproduce the string exactly and place it in the frame.

Two product limits shape everything below. Clip length comes from a short list rather than a typed number: auto, 5, 10, 15, or 20 seconds. Frame shape, however, is wider open. It runs from 21:9 down to 9:16, so a wide film crop and a vertical social crop are both available. Resolution is hd or fhd.

Text-to-video prompts

Words alone. This is where prompt craft matters most, because the model has nothing except what you type. Name the sound source rather than a mood. “Gravel crunching under boots” gives Flux 3 something to render. “Atmospheric” does not.

Prompts led by sound

The audio is the point of these. In short, each one names what makes noise, and says what stays quiet.

Room tone and one small action

A locksmith’s workshop late in the evening, one bare bulb over the bench, brass filings on the wood. A hand turns a half-cut key against a grinding wheel, sparks arcing off in short bursts. Sound: the rising whine of the wheel each time the key touches it, filings ticking onto the bench, a radio murmuring in another room too quietly to make out words. No music.


An impact that has to land on frame

A slow push-in on a cast iron skillet over a gas flame, butter already browning. A steak comes down into the pan and holds. Sound: the sear blooming loud exactly as the meat touches the surface, then settling into a steady crackle, the gas ring hissing underneath. No music, no voices.


Weather doing the work

A tin-roofed bus shelter on an empty coast road, one passenger with a hood up, the sea flat and grey behind. Rain arrives partway through, first a few taps and then a downpour. Sound: rain building on the tin from scattered hits to a solid roar, wind pulling at the hood fabric, distant surf. No music.

Prompts with dialogue

Speech and frames come out of Flux 3 together, so the lip-sync comes built in rather than patched on afterwards. Quote the line exactly. Then put it in the mouth of a described person, so the model knows who speaks.

One line, one speaker

A night fishing boat tied up at a small harbour, a woman in oilskins coiling rope under a sodium lamp. She looks up at the camera and says, “You are three hours late.” Sound: her line clear over rope creaking, water slapping the hull, a halyard tapping a mast. No music.


A two-hander across a cut

Two mechanics either side of an open bonnet in a daylit garage. The first, wiping his hands on a rag, says, “It is not the alternator.” Cut to the second, crouched at the wheel arch, who answers, “I know what it is.” Both shots handheld, same warm daylight, same garage. Sound: the two lines in sequence, a compressor cycling in the background, tools shifting on a tray.

Prompts that use more than one shot

Flux 3 can cut inside a single generation and hold the character and the look across the cut. So say how many shots and what each one shows.

Three shots, one continuous sound bed

Three shots of the same baker, early morning, same flour-dusted apron throughout. Shot one, wide: she pulls a tray of loaves from a deck oven, steam rolling out. Then shot two, close: her hands turning a loaf over to knock the base. Finally shot three, medium: she sets it in a basket by the window as light comes through. Sound: one continuous bed across all three, oven door clunking, the hollow knock on the crust, a street waking up outside. No music, no dialogue.


A reveal built across a cut

Shot one: a tight close-up of a woman’s face lit from below in flickering orange, eyes moving, unreadable expression, no context. Shot two: cut wide to reveal she is standing in a field at night in front of a burning haystack, alone, hands in her coat pockets. Hold the wide. Sound: fire roaring up in the cut to the wide shot, crickets underneath, no music.


One continuous take instead

ONE continuous unbroken shot, no cuts. The camera follows a courier on foot through a covered market at speed, weaving past crates and hanging fabric, the courier always centre frame with the market blurring past on both sides. Ends when they push through a doorway into hard daylight and stop. Sound: footsteps, crates knocking, overlapping market chatter rising and falling as they pass each stall, then sudden open-air quiet.

Prompts with words on screen

Quoted strings render as part of the scene. Also keep the string short and state its position in the frame.

A title card inside the shot

A record shop window at dusk seen from the pavement, rain on the glass, sleeves propped up inside. The camera pushes in slowly. Hand-painted white lettering on the glass reads “CLOSED SUNDAYS” and stays legible and steady through the whole move. Sound: rain on the awning, a bus passing behind the camera, faint music from inside the shop.


Signage that has to stay stable

A slow lateral track past a row of market stalls at night, string lights overhead. A chalkboard propped on the middle stall reads “TWO FOR FIVE” in clear handwriting, and it holds legible as the camera passes it. Sound: sizzling from a griddle, coins in a tin, low crowd noise. No music.

Prompts in styles other than cinematic

Flux 3 covers animation, motion design, and stop motion as well as photographic looks. Name the medium first, and describe the motion in that medium’s own language.

Hand-drawn cel animation

1990s hand-drawn cel animation, painted backgrounds, bold ink lines, visible film grain. A cartoon fox in a striped scarf drags a sled twice its size up a snowy hill, slipping back a step for every two it gains, the effort exaggerated with squash and stretch and smear frames on the fast slips. It reaches the crest and immediately loses the sled over the far side. Sound: a comedic orchestral score matching each slip, snow crunching, a long descending whistle as the sled goes over. No dialogue, no on-screen text.


Stop motion

Stop-motion clay animation, visible fingerprints in the clay, slightly uneven frame timing, shot on a small tabletop set. A stout clay figure in a knitted hat tries to post an oversized letter into a postbox slot, folding it smaller and smaller until it fits, then looking pleased with itself. Sound: the dry scrape of clay against the postbox, paper creasing, a jaunty ukulele under the whole thing.


Motion design

Clean 3D motion design on a deep navy background, soft studio lighting, no camera shake. Six matte-white spheres drop into frame from above, bounce with real weight, and settle into a perfect hexagon. As the last one lands, a thin gold ring draws itself around the whole group. Sound: a soft low thud per sphere as it lands, then a single bright chime as the ring completes. No voices.

Image-to-video prompts

Flux 3 Video takes up to ten reference images in a Reference slot. The references carry the look, the subject, and the character. So the prompt no longer has to say what things look like. It says what happens instead, which is a different writing job.

Instead, keep the prompt on motion, camera, and sound. Describing what the model can already see in the reference wastes the prompt, and sometimes fights it.

Animating a single reference

Give a still photograph motion and sound

Bring the reference to life with minimal invention. Keep the subject, wardrobe, framing, and colour exactly as they are. Add only what the moment needs: a slow breath, a small shift of weight, hair and fabric moving in a light breeze, one blink. The camera drifts in almost imperceptibly across the clip. Sound: the ambience this location would actually have, nothing dramatic, no music.


Let the camera do the work

Hold the subject in the reference exactly as framed and move the camera around it in one slow arc, roughly ninety degrees, keeping the subject centred and the same size in frame throughout. Lighting shifts naturally as the angle changes. Nothing else in the scene moves. Sound: quiet location ambience and the faint low rumble of the move. No music.

Using several references for a look

References as a style guide

Use the references for palette, lighting, texture, and general atmosphere, not as frames to reproduce. Generate a new shot in that same visual world: a cyclist riding a narrow lane between tall hedges at the same time of day, camera tracking alongside at wheel height. Match the grade and the grain of the references. Sound: tyres on loose gravel, chain and freewheel ticking, birds in the hedges. No music.


References as a wardrobe and set bible

Take the clothing, props, and set dressing from the references and stage a new action with them. A market trader in that same apron and cap stacks crates of apples into a pyramid, checks it, and pulls one crate back off the top when it looks unstable. Same location, same light as the references. Sound: wooden crates knocking together, apples shifting, a low market hum.


One reference for the subject, the rest for the world

Keep the person from the references exactly as they appear, and place them into a wider version of the same environment. They walk from the back of frame toward the camera and stop just short of it, looking off to one side rather than at the lens. Camera static at chest height. Sound: footsteps on the surface underfoot, ambience appropriate to the location, no music.

Keeping a character consistent

One person, three shots, one generation

Keep the person in the references identical across every shot: same face, same hair, same jacket. Three shots in sequence. Wide, they stand at a bus stop checking a paper timetable. Medium, they fold it and put it in a pocket. Close, they look up as headlights wash across their face. Same overcast light throughout. Sound: one continuous street bed, paper folding, a bus approaching under the last shot.


Consistency through a change of location

Keep the character from the references unchanged in face, build, and costume. Shot one: they push open a heavy door from inside a dim stairwell. Shot two: the same person outside on a bright street, squinting, adjusting to the light, walking out of frame left. The light changes hard between the two shots but the person does not. Sound: the door mechanism and stairwell echo, then open street noise arriving abruptly on the cut.

Video continuation prompts

Feed a clip into the Video slot and Flux 3 carries on from where it ends. Momentum, framing, and scene logic all come across too. So the prompt only has to say what happens next. In fact, describing the clip again is the most common way to waste one of these.

Carrying the action forward

Straight continuation

Continue from the final frames with no cut and no change of camera. Keep the same motion, framing, lens, grade, and pace. Let the action already in progress simply finish, then hold the shot for a beat on the result. Sound: continue the existing ambience without a seam.


Continue and resolve

Pick up exactly where the clip ends, same shot, and bring the movement to a stop. Whatever is travelling slows and settles. The camera settles with it and holds steady. Nothing new enters frame. Sound: the existing bed thinning out as motion stops, ending on quiet room tone.

Turning the scene

Something arrives

Continue from the last frames in the same unbroken shot, then let one new element enter: a second person walks in from frame right, stops beside the subject, and looks at the same thing they are looking at. Neither speaks. The camera stays where it is. Sound: the existing ambience continuing, footsteps arriving and stopping.


Widen out

Continue the clip and pull the camera back steadily to reveal the wider location around the subject, ending on a wide that shows where this has been happening all along. Keep the subject in frame throughout the move. Sound: the close ambience opening out into a broader, more distant version of itself.

Carrying the sound forward

Continue the audio, change the picture

Continue from the final frames, keeping the soundtrack unbroken and consistent with what came before. Cut once, to a different angle on the same scene, then hold. The sound must not break or restart across the cut. Sound: one continuous bed running through the cut, with the perspective of the noise shifting slightly to match the new angle.

The Flux 3 prompt formula

Every prompt above uses the same six parts. Drop any of them, and Flux 3 then fills the gap with a decision of its own.

  1. Shot and camera. Wide, close, tracking, static, one take or several.
  2. Subject and action. Who or what, doing one specific thing with a beginning and an end.
  3. Setting and light. Where, what time, what the light is doing.
  4. Style or medium. Photographic, animated, stop motion, motion design. Name it first.
  5. Sound. Ambience, impacts, dialogue in quotes, and whether music belongs.
  6. What to exclude. Say “no music”, “no dialogue”, or “no on-screen text” out loud.

The sixth part matters more than it looks. There is no separate negative field, so put anything you want ruled out in the prompt itself.

Where Flux 3 Video runs in Picsart

Flux 3 Video runs in three places. The choice between them is about how much work happens around the prompt.

  • The AI Playground is the place to test a prompt against other models before committing to it, with 150+ available from a single prompt. Pick Flux 3 Video from the model list, set resolution, frame shape, and duration, and attach references or a clip in the slots beside the prompt box.
  • The AI video generator is the faster route for paste-and-go work when the model choice is already made.
  • Picsart Flow is where a prompt becomes one step in something larger, feeding its output into the rest of a multi-step project.

Nothing needs setup in any of the three. Flux 3 also reaches past the web app and the desktop app. Also available: a CLI, an MCP server, a REST API, and an SDK. Anyone can build Flux 3 straight into a project of their own.

Finally, the Flux 3 model page has the full details. It also covers what the model handles beyond the prompts here.

Get answers to common questions

A Flux 3 prompt is a text description that produces video and its soundtrack in one generation. Unlike a plain video prompt, it also carries the sound. Flux 3 makes picture and audio together in a single pass, so the prompt decides both.

Start writing prompts that carry sound

One test shows the difference fast. Take any prompt from the text-to-video section and generate it as written. Then generate it again with the sound line deleted. Everything else stays the same, yet the two clips will not feel like the same clip. That gap is the part of Flux 3 worth learning. Open the AI Playground, pick Flux 3 Video, and paste the first one in.