A face that shifts halfway through a shot ruins the shot. Thirty seconds is a long time to ask a video model to hold onto a person, and the failure always looks the same. The rider who takes off is not quite the rider who lands. A jaw softens, a hairline creeps, a beard fills in somewhere over the second act. Nobody watching can name what went wrong, and everybody feels it.

There is a Picsart Flow template that solves this before a single frame of video exists. It spends its first half building something that never appears in the finished clip, then generates a continuous 30-second flight over a fantasy landscape on the back of a dragon. The thing it builds first is a character reference sheet, and that sheet is the entire trick.

Why one photo cannot carry a 30-second shot

A portrait is one angle, one lens, one lighting setup. It shows a person from the front, at eye level, holding still. That is roughly two percent of what a flight sequence needs.

A camera that circles a rider mid-air asks for the back of a head, a profile against the sky, a three-quarter view from below, a shoulder in the corner of a first-person frame. None of that exists in the source photo, so the model invents it. Invention is exactly where drift comes from, and the longer the clip runs, the further the invention travels from the person it started with.

The fix is not a longer prompt. Piling adjectives onto a description of a face gives the model more words to interpret, not more information to copy. The fix is to give it more angles before you ask for motion.

What a character reference sheet actually is

A character reference sheet is a single image holding the same character in several aligned views, made to be looked at by whoever renders that character next. Animation and game studios have used them for decades under the name turnaround, and the job has never changed: remove guesswork from everyone downstream.

The only difference here is who reads it. Instead of an animator matching a model sheet by hand, a video model receives the sheet as reference and matches it frame by frame. Every angle sitting on that sheet is an angle the model no longer has to guess.

The five views the sheet has to cover

The template writes the sheet from a text prompt that is specific about what belongs on it. Five views, all on one canvas:

  • Full-body front view
  • Full-body side view, facing right
  • Full-body back view
  • Front close-up of the head and shoulders
  • Top-down view of the whole figure

That last one earns its place the moment the camera climbs above a flying dragon and looks down.

The prompt is equally specific about how the figure stands: a neutral relaxed pose, arms naturally at the sides, a calm neutral expression, every view aligned in a clear grid like a studio character design sheet. The neutrality is doing real work. A reference sheet is source material, not a finished picture, and any drama baked into the pose becomes drama the video model faithfully copies into all 30 seconds.

Building the flying dragon video step by step

The canvas runs left to right in four movements: reference images feed a character sheet, the sheet feeds a video, and the video feeds an upscale. Here is what happens at each node.

Step 1. Open the template as a live canvas

Open the flying dragon video template. It arrives fully wired, with every node connected and every prompt already written, so nothing has to be built from an empty canvas. What you are looking at is a working pipeline you can inspect, run and rewrite.

Step 2. Supply the character references

Two image inputs feed the character sheet. The first is the portrait that becomes the rider, and that is the one to swap for your own, so use a clear, well-lit, front-facing shot where the face is unobstructed. The second is an outfit reference showing a costume from the front, the side and the back, and it is what dresses the character on the sheet.

Splitting these apart is a deliberate design choice. The portrait carries identity and the outfit reference carries wardrobe, which means the sheet arrives already dressed and the video prompt never has to describe clothing in words. Wardrobe described in adjectives is wardrobe that drifts.

Step 3. Generate the character reference sheet

A text node holds the sheet prompt, roughly 240 words of it, and that prompt is what turns two loose references into a structured turnaround. It asks for a single person wearing the reference outfit, rendered in the five views listed above, aligned in a clear grid.

This node runs on GPT Image 2. What comes back is the multi-view sheet plus a strip of close crops along the bottom picking out fabric and detail, which gives the video model texture information as well as silhouette.

Step 4. Check the sheet before spending anything on video

Read the sheet properly before moving on. The front, side, back, close-up and top-down views should all look like the same person wearing the same clothes, and the top-down view deserves the closest look because it is both the hardest for the model to produce and the one an aerial sequence leans on hardest.

Regenerate the sheet if any view disagrees with the others. A bad sheet fixed here costs one image generation. The same bad sheet discovered after the video renders costs the entire clip.

Step 5. Configure the Seedance 2.5 video node

The video node is where the sheet gets turned into motion, and it runs on Seedance 2.5. Its settings, as the template ships them:

Setting Value What it controls
Model Seedance 2.5 Generates the flight sequence
Duration 30s The full length of the clip, generated as one take
Audio On The clip returns with sound rather than silent
Resolution 720p The setting the template ships with. 1080p is also selectable

The run bar carries two further toggles, displayed as Frame and No EP, both left at their template defaults.

Reference capacity is the other reason this chain holds together. Seedance 2.5 accepts up to 50 references, so a sheet carrying five aligned views plus a strip of fabric and detail crops sits comfortably inside what the model can consider at once.

Resolution is a choice rather than a ceiling here. The template ships at 720p and 1080p is available on the same node, but the upscale waiting at the end of the chain is what determines the resolution of the finished piece either way. That makes the lower setting the sensible place to work while the shot is still being judged.

The duration is the setting worth pausing on. Thirty seconds as a single unbroken take is a Seedance 2.5 capability specifically, and it is what makes the template’s central request possible at all. Ask most video models for half a minute and you get a sequence of cuts stitched together, each one a fresh opportunity for the face to change. One take means one continuous piece of footage and one continuous character.

Step 6. Run the video and read the result

Run the node and watch the whole clip rather than scrubbing it. Drift shows up in transitions, so the moments to check are where the camera changes relationship to the rider: the switch from first-person to an external view, the pass behind the dragon’s wing, the climb where the rider is briefly seen from above.

If the rider changes, go back to step 3. That is the loop the workflow is built around, and it is why the video renders cheaply.

Step 7. Upscale the take you are keeping

The finished clip goes through the upscale node, which lifts it from the working resolution to a 4K output. This is the step that makes a low-resolution generation worth running in the first place.

Step 8. Export

Export the finished video from the canvas. Exporting usually means signing in first.

Inside the prompt that keeps the character locked

The video prompt is not a paragraph. It is written as labelled blocks, each responsible for one decision, which is why it survives being 30 seconds long.

The first block sets the shot: “One continuous 30-second flight sequence, no hard cuts – the camera flows between locked first-person POV and detached free-flying camera moves as if a single unbroken drone-rig follows the ride.” One sentence, and it has already ruled out the thing video models love to do unprompted, which is cut. It also defines the camera as a single physical rig, which is why the perspective changes read as one continuous move rather than as edits.

The block after it handles the rider, and it does something worth stealing. Rather than describing a face in adjectives and hoping, it names a specific input image with a reference token and instructs the model to preserve that image exactly. The character is defined by pointing, not by describing. Every other block in the prompt can then talk about wind, altitude and light without ever touching the question of who is in the frame.

That is the difference between asking for consistency and building it. Adjectives are a request. A reference token is an instruction.

Render small, then upscale once

The video node runs at 720p even though 1080p is available on it and the finished piece is 4K, and that ordering is deliberate rather than a limitation. Iterating at a higher resolution would multiply the cost of every take that gets discarded, and most takes get discarded.

So the workflow generates cheap, judges the result, regenerates until the flight reads properly, and only then spends the upscale on the single version worth keeping. Resolution is the last thing to buy, not the first.

Tips

Keep the reference pose neutral

A dramatic stance on the sheet becomes a dramatic stance in every frame of the video. Neutral input, flexible output.

Let an image carry the outfit

Wardrobe described in words drifts between shots. Wardrobe supplied as a reference stays put.

Point at an image instead of describing it

A reference token beats a paragraph of facial description every time, and it keeps the prompt short enough to stay readable.

Fix the input, not the prompt

When a face comes back wrong, the sheet is usually the problem. Regenerate the sheet rather than adding more instructions to the video prompt.

Judge at low resolution and upscale once

Resolution does not tell you whether a shot works. Save it for the take you are keeping.

Watch the transitions

Character drift shows up where the camera changes its relationship to the subject, not in the calm middle of a move.


Get answers to common questions

A character reference sheet is one image showing the same character from several aligned angles, built as source material for whoever renders that character next. Studios call it a turnaround. In an AI workflow, the video model reads it instead of an artist.

Start building

The sheet is the part worth taking with you. Any workflow that has to keep one person recognizable across a long shot gets better the moment you stop describing that person and start showing them from every angle first.

Open the Picsart Flow template, swap in a portrait, and see what comes back. From there the workflow editor lets you rewire any part of it.