AI motion control copies the movement from one video onto a still image of a completely different character. You upload two things: a portrait, and a clip of the action you want. The model then rebuilds your character performing that same action, matching the hand positions and the expression on the face. There is no rigging, no keyframes and no motion capture suit involved. Two things decide whether it works. The first is the reference clip you pick. The second is how well that clip matches your image. This guide covers both, along with the settings that most people get wrong on the first try.
What AI motion control actually does
AI motion control is motion transfer. The model watches the body movement and facial performance in a reference video, then maps it onto a still image of a different character. Everything about the movement comes from the video: the pose, the timing, the weight shifts, the small changes in expression. Everything about the look comes from the image: the face, the outfit, the art style. You end up with a clip of your character doing something they were never filmed doing.
That split is what makes motion control different from ordinary animation. A standard image to video model invents the movement from a text prompt. The result looks plausible, but it is never exact, and it changes every time you run it. Motion control takes the guesswork away, because the movement is already recorded. The model only has to re-perform it. You feel the difference most on choreography, on a specific gesture, on a particular walk, or on a hand sign that has to be right.
The feature works on realistic humans, stylized characters, anime figures and 3D renders. It also handles humanoid animals, plus animals built on roughly human proportions. Accuracy does drop as the body plan moves further from a human skeleton. For the cleanest results, film a real person for the reference clip.
Motion control sits in the AI Video Generator and the AI Playground. It is listed as its own model, not as a toggle buried inside another one. Kling powers both versions, one built on Kling V2.6 and one on Kling 3.0. Both run on Picsart AI Credits, the same balance every other model uses.
How to use AI motion control in Picsart
1. Open the model chooser
Start in the AI Video Generator or the AI Playground. Motion control appears there as a category of its own.
2. Pick your motion control model
Two options show up. Kling Motion Control 2.6 transfers body movement from a reference video onto a portrait photo. Kling Motion Control V3 does the same job at V3 quality. The comparison further down covers which one suits which shot.
3. Add the character image
One image goes into the Reference slot. This is the character you want to see moving. Use a clean shot with nothing blocking the body or the head.
4. Add the motion video
One clip goes into the Video slot. This is the performance you want copied. Nothing else you do here matters as much, so the next section is entirely about getting this right.
5. Prompt the background, not the movement
The prompt cannot change the action, because the action comes from the video. Use it for everything around the character instead: the setting, the wardrobe, the lighting, anything else in the scene. Something like "Standing on a wet neon-lit street at night, in an oversized denim jacket, steam rising from a vent behind her" is what belongs here.
6. Set your rendering mode and generate
Kling V2.6 gives you two rendering modes. Standard produces fast HD drafts and Pro produces Full HD final output. Test your reference clips in Standard, then re-run the keeper in Pro.
Notice what is missing from that list. There is no duration picker and no aspect ratio picker. Both are inherited from the reference video, so your clip decides how long the render runs and what shape it comes out in.
Most failed renders are reference video problems, not model problems. The clip has to be one unbroken take, with the character visible the whole way through. Cuts, shot changes and camera moves all break it, and the usual symptom is a video that comes back truncated. A phone clip of someone doing the action in one go beats a polished edit every time.
Your reference video decides almost everything
Match the framing to the image
A half-body portrait needs a half-body reference. A full-body portrait needs a full-body reference. Driving a chest-up character with a clip of someone dancing head to toe is the most common mistake here, and the distortion shows immediately.
Keep the motion moderate
Wide movement works well. Speed does not. Fast, frantic action gives the model less to read, and a character who stays roughly in place beats one who sprints out of shot.
Leave room to move
Big gestures need space around the character in the reference image. A tightly cropped portrait leaves nowhere for an arm to swing.
Use one character
A clip with two or more people in it will not fail. Only one performance gets used, and it belongs to whoever fills the most of the frame. Crop to a single subject and you remove the guesswork.
Mind the technical floor and ceiling
Clips run from 3 to 30 seconds. On resolution, 340px is the minimum on the short side and 3850px is the maximum on the long side.
Test difficult movement on a short clip first
Your video matches the length of the clip you upload, but only across the parts where the motion is usable. Anything complex or fast gets trimmed back to the stretch the model can actually read. A 20-second upload can come back as 12 seconds, and those credits are not refunded. The model needs at least three seconds of clean continuous movement to produce anything at all.
Character orientation is the setting that changes the shot
Character orientation decides where the camera lives. It has two settings, and they produce genuinely different videos from identical inputs.
Character orientation matches video is the default. Everything follows the reference clip: the movement, the expressions, the way the character faces, and the camera work. Film the reference with a slow push-in and you get a slow push-in. Use this setting to copy a performance whole.
Character orientation matches image keeps the character facing the way they face in your still. The movement and expressions still come from the video. You then direct the camera through the prompt, choosing zoom in, zoom out, camera up, camera down or a fixed position. Use this setting to reuse one good performance across several shots without refilming anything.
The rule is simple. Pick the video setting for a faithful copy. Pick the image setting when you want control over the composition. Both give you the same performance, and the only real difference is who holds the camera.
Kling Motion Control 2.6 vs V3
Both models take the same two inputs. Both produce a clip that matches the reference length. What separates them is where each one puts its accuracy.
| Motion Control 2.6 | Motion Control V3 | |
|---|---|---|
| Best for | Full-body action, dance, hand-heavy performance | Close-ups, head turns, emotional range |
| Facial handling | Solid on steady, front-facing motion | Holds identity through angle changes, partial occlusion and sudden jumps in shot size |
| Standout strength | Full-body sync and hand precision over long takes | Stable features and smooth expressions in complex, multi-angle motion |
| Inputs | One character image, one reference video | One character image, one reference video |
| Output length | Matches the reference clip | Matches the reference clip |
| Choose it when | The action is the point | The face is the point |
The short version is that 2.6 is the workhorse for movement and V3 is the upgrade for performance. A dance clip where the character stays front-on suits 2.6 well. A monologue that swings from happy to sad while the head turns is a V3 job. So is anything where the character passes behind an object or jumps between close-up and wide mid-shot. Kling 3.0 built its motion control around holding a face together in exactly those situations.
Get answers to common questions
AI motion control transfers the movement and facial performance from a reference video onto a still image of a different character. The image supplies the identity and the style. The video supplies the action, and the model combines the two into a new clip. It does the job that manual animation and motion capture used to do.
Start animating your characters
One portrait and one clean reference clip. That is the whole input list, and the model handles the rest. Pick your motion control model in the AI Video Generator and bring a character to life.