A MiniMax prompt has to answer one question before anything else. Which of the two models is going to read it? MiniMax H3 understands text, images, video, and audio in one context, and returns video with native stereo sound. MiniMax H3 Max reads words and frames only, and reads them fast.

That split decides what your prompt needs to say. Write a soundtrack into an H3 prompt and the clip arrives with audio already in sync. Write the same line for H3 Max and you have spent words on something its controls do not expose. So the ten prompts below are sorted by the model that runs them best.

Both models live in Picsart AI Playground, which means you can paste a prompt, generate, and switch models without changing anything else. Copy any prompt here, swap in your own subject, and run it.

What every MiniMax prompt has to carry

Both models want the same four things named plainly, in this order:

  • The subject and the setting. One clear thing in one clear place. “A ceramicist at a wheel in a dusty studio” beats “an artist working”.
  • The action, with a beginning and an end. A clip runs 5 to 15 seconds, so pick an action that fits inside one. One completed move reads better than three rushed ones.
  • The camera. State whether it holds still, pushes in, or tracks alongside. Say nothing and the model chooses for you.
  • The light. Time of day, direction, and hardness. This is the fastest way to change how a shot feels.

After those four, the models diverge. MiniMax H3 generates native stereo audio in the same pass as the picture, so an H3 prompt gets a fifth job: name what makes the noise. MiniMax H3 Max was tuned to follow a prompt more closely and to look better doing it, so its prompts reward precise camera and framing language instead.

There is one more habit worth building for H3. When you attach a reference, describe the relationship between that input and the clip you want, rather than just describing the clip. “Move the camera the way the reference video does, but around the man in the reference image” is the kind of instruction it was built to follow.

Which model should run your prompt

Reach for MiniMax H3 when the clip has to arrive finished. It generates 2K by default at 24 fps, produces matching stereo audio in the same pass, models multiple shots natively inside one generation, and takes reference images, video, and audio together. It is also strong at rendering legible text and brand marks, and at transferring motion from one clip to another. Clips run 5, 10, or 15 seconds.

Reach for MiniMax H3 Max when you are still deciding. It turns a prompt around in seconds at 480p or 768p, accepts any whole second count in the 5 to 15 range, and pins the first and last frame. There are no reference slots, so the prompt carries everything.

The practical order is to draft on H3 Max and finish on H3. Full specs for each sit on the MiniMax H3 and MiniMax H3 Max model pages.

MiniMax H3 prompts: sound, references, and 2K

These six lean on what only H3 does. Each one either names its own audio or puts a reference to work. Notice how the sound line describes a source rather than a mood, because “a wooden rib scraping the clay” is something a model can render and “atmospheric” is not.

One continuous take with its own soundtrack

A ceramicist’s studio at first light, wet clay turning on the wheel, grey light coming through one dusty window. Her hands close around the rising wall of the pot and steady it, both thumbs pressing a groove into the rim. The camera holds at hand height and does not move. Sound: the low hum of the wheel, water sliding under her palms, a wooden rib scraping once against the clay, birds outside the glass. No music.


Several shots inside one generation

A night market in the rain, three shots in one clip. First a wide of the lane, canopies dripping, string lights doubled in the puddles. Then a close on a wok as the noodles hit the oil and flare up. Finally a medium of the cook handing a paper box across the counter to a customer in a yellow raincoat. Sound: rain drumming on canvas, the burst of the wok, low crowd chatter underneath.


Motion borrowed from a reference video

Take the camera movement from the reference video and apply it to a new subject: a lone red tractor parked in a harvested field at dusk. Match the reference for speed, direction, and the moment the move settles, but change nothing about my subject or setting to suit it. Low sun behind the tractor, long shadows across the stubble. Sound: wind across open ground, metal ticking as the engine cools.


A character locked by a reference image

Keep the woman in the reference image exactly as she is: her face, her hair, and the green corduroy jacket stay identical from the first frame to the last. Put her in a second-hand bookshop, walking the length of the aisle, pulling a paperback from a high shelf and reading the back cover as she keeps walking. Warm tungsten light, tall stacks either side, shallow depth of field. Sound: floorboards creaking under her boots, a page turning, a radio playing quietly at the front of the shop.


Text and a logo that stay legible

The camera creeps toward a cafe’s front window, early morning, the street still empty behind it. The words “OPEN FROM SEVEN” are painted on the glass in cream serif capitals, arched, and they stay sharp and correctly spelled for the whole clip. Reproduce that string exactly. Below it, a small circular logo in the same cream, centered under the arch. Soft overcast daylight, faint reflections of the street in the glass. Sound: a distant bus, a shutter rolling up somewhere off screen.


A spoken line matched to a reference voice

Have the mechanic in the reference image speak the line “It was never the alternator” in the voice from the reference audio, matching its pace and delivery rather than reading it flat. She is in a lit garage bay, oil on her forearms, wiping a wrench on a rag, and she looks up at someone off camera before saying it. Then she turns back to the engine. Medium shot at chest height, hard overhead work light, deep shadow behind her. Sound: her line clear over the ring of a dropped socket and a compressor cycling behind her.

MiniMax H3 Max prompts: fast, framed, and repeatable

These four are built for the drafting pass. They stay short on story and long on framing, because framing is what H3 Max has to work with. Run them at 5 seconds and 480p first, then rerun the one that works at 768p and full length.

Five seconds to test one idea

A cyclist crests a coastal road at golden hour, the sea on her left, dry grass bending in the wind. She stands on the pedals for the last of the climb, then sits back down as the road levels out. The camera tracks alongside her at the same speed, low and close to the wheels.


A product reveal pinned between two frames

Start frame: the closed box on a concrete surface. End frame: the bottle standing upright beside the open box. Between them, a pair of hands lifts the lid, sets it aside, and draws the bottle out in one continuous move. Keep the surface, the background, and the light identical from the first frame to the last. Camera locked off, no cuts.


Photo animation from a start frame

Start frame: the uploaded photo. Hold the framing, the clothing, and the light exactly as they are, then let the scene run forward. Steam lifts off the cup, the sitter turns her head toward the window and settles back, traffic crosses the street beyond the glass. The camera drifts in slightly and stops. Nothing else in the frame changes.


A camera move stated exactly

A slow dolly in on one empty chair in a school gymnasium, mid-afternoon, dust hanging in the light from the high windows. Start wide enough to see the painted lines on the floor and end tight on the chair back. Constant speed, no easing, no handheld shake. Muted colors, long shadows running away from the windows.

Prompt expansion, and how much to write

MiniMax H3 Max adds a control the prompts above assume you will touch. Prompt expansion has three modes: disabled, balanced, and quality. Disabled runs your words as written, which is what you want once a prompt is doing exactly what you asked.

Balanced and quality let the model elaborate, which helps a short prompt and can overwrite a long one. So the rule is simple. The more detail you have written, the further down that scale you should sit.

There is also a seed field. Fix it and a rerun of an identical prompt lands in the same place, which is how you change one word at a time and see what that word actually did.

Where to run these MiniMax prompts

Both models sit in AI Playground, where you can run one prompt through several models and compare the results side by side. That is the fastest way to see the difference between an H3 clip with sound and an H3 Max clip without it.

MiniMax H3 is also in the AI video generator if you want to paste a prompt and go. Either way, the rest of the AI models catalog is one click away when a shot calls for something else.

Get answers to common questions

A MiniMax prompt is the text you give MiniMax H3 or H3 Max to generate a video clip. A good one names the subject, the action, the camera, and the light.

Start generating with MiniMax

Pick the prompt closest to the shot you want, change the subject, and generate. Draft it on H3 Max, then run the winner through MiniMax H3 for 2K and stereo sound.

Try MiniMax in AI Playground