Contents
Kling 2.6 generates video and sound together in one pass. A single prompt returns a clip with dialogue, sound effects and room tone already in it. There is a catch worth knowing first. The default settings produce a silent video, because sound arrives only at 1080p and the model starts at 720p with audio off.
The audio is the part everyone talks about. Control is the more interesting part. Kling 2.6 lets you pick which voice comes out of which character, and drive a character with movement taken from a real video. Most of this guide is about using those two well.
What Kling 2.6 actually generates
The model handles a wider range of sound than most descriptions suggest, and it layers them, so one scene can carry a voice, a room and an effect at once.
- Voices. Narration, conversation between characters, singing and rap with real lyrics.
- Ambience. Wind, traffic, waves and other background beds.
- Effects. Specific actions like glass breaking or footsteps on gravel.
The basics are quick to state. Clips run 5 or 10 seconds, at 720p or 1080p, in 16:9, 9:16 or 1:1. Prompts stretch to 2,500 characters, far more room than most people use. Pick 10 seconds for singing or a back-and-forth, since a conversation rarely resolves in five.
There are two routes in. Text to video builds the whole scene from a description and offers the square format. Image to video animates a still you supply, takes its shape from that image, and accepts a specific voice.
Getting the audio settings right
Sound and resolution are linked, and this trips up nearly everyone. Native audio works at 1080p and nowhere else. Asking for audio at 720p is not a lower-quality version of the feature. It is an invalid combination, and since generation starts at 720p with audio off, both settings need changing before anything makes a sound.
| What you want | Settings | Worth knowing |
|---|---|---|
| Video with sound | Audio on, 1080p | 720p is silent, and both defaults have to be changed |
| A specific voice | Image to video, audio on, 1080p | Two voices maximum per clip. Audio cannot be off |
| A start and end frame | Image to video, 1080p | Produces a silent clip. Sound is not available here |
| Movement from a real clip | Motion control | Keeps the source clip’s real sound. Reaches 30 seconds |
| Dialogue or singing | 10 seconds | Five seconds rarely completes an exchange |
What you can make with Kling 2.6
- A person talking to camera. Product demos, lifestyle clips, presenting and reporting. Lip movement tracks the speech, and the delivery note carries the genre: steady and measured reads as reporting, warm and quick reads as a vlog. Give the speaker one physical action so the shot has something to do besides talk.
- Narration over a scene. An unseen voice explains while the camera stays on the subject. The easiest category to get right, since no lip sync has to hold up. Pace is the whole game.
- Two people in conversation. Interviews, scripted scenes, everyday exchanges. Needs the dialogue rules below and benefits most from 10 seconds. Comic timing lives here, and a beat of silence before a reply is worth writing in.
- Music and performance. Singing with real lyrics, rapping over a described beat, group vocals, instrumental playing. Name the genre, the technique and the accompaniment separately. Ten seconds is close to essential.
- Atmosphere and effects. Observational scenes, close-up texture, ASMR, effects-led ads. The cheapest way to learn how the model handles sound, because there is no dialogue to go wrong.
The prompt formula that works
Sound has to be written, not assumed. Five things go in order, and leaving the audio out is the most common reason a clip comes back flat. The model fills that gap with whatever the scene implies, rather than with what you had in mind.
- Scene. Location and time of day, named in one clause. Example: a narrow ramen counter late at night, steam on the window.
- Element. Who the shot is about. Give it a label you can reuse. Example: a chef in a navy apron.
- Movement. One physical action, not three. Example: lifts a basket of noodles and taps it twice against the rim.
- Audio. Quote spoken lines, name the source of every sound. Example: the hiss of broth, the double knock of the basket, low chatter behind.
- Other. Style, mood, camera. Lens and light do more than adjectives. Example: warm tungsten light, shallow focus, handheld.
Drop one and it shows. Skip the movement and the frame goes static. Skip the style note and the model picks a look for you.
A vocabulary for directing sound
Swapping a vague word for a precise one changes the output more than adding another sentence does. Most prompts describe speech and ignore everything else, and that is why so many clips sound like a voice recorded in a vacuum.
- Speech. Whispering, softly speaking and clearly speaking set the level. Excitedly, complaining and sighing set the feeling. Hoarse and deep describe the instrument. Fast and slow talking set the rhythm. Reciting, monologue and voiceover set the mode.
- Interaction. Answering, arguing, shouting, discussing, crying, screaming, laughing, chuckling.
- Music. A capella, humming, loud singing, opera, pop vocals, vibrato, falsetto, harmony, rapping, fast rap, heavy beat.
- Objects. Knocking, footsteps, chewing, glass shattering, metal clanging, friction, thunder, fire crackling, bubbling, sirens, braking, gears whirring.
- Space. Traffic noise, crowd murmur, subway noise, ocean waves, bird chirping, wind, rainforest, library silence, cafe background, air conditioner hum, fireplace burning.
Space is the group most often skipped, and it sells a shot. A library reads as a library because of the silence and one dropped book. Naming the reverb helps too, since a hall and a small room carry a voice very differently.
Writing dialogue between characters
Conversations are where prompts fall apart, and the failure mode is always the same. Both lines come out in one voice, or the wrong character speaks, because the model could not tell who was who. Four habits remove that ambiguity.
| Habit | Do this | Not this |
|---|---|---|
| Naming | The same unique label every time | Switching to he, she, or a synonym |
| Anchoring | Action first, then the line | A line with no action attached to it |
| Voice | A distinct tone and emotion per character | A man says, a woman replies |
| Order | Linking words, plus a note that the speaker changes | Two quoted lines stacked with nothing between |
Keep the cast small. Two speakers is where the model is reliable, and quality drops once a third voice enters. A crowded exchange works better split across two clips than forced into one.
Choosing a voice for a character
Voices come from three places, and only one needs a recording.
- A ready-made library. The fastest route, and the answer for anyone without clean audio to hand.
- A clip you upload. Audio or video, 5 to 30 seconds, one speaker, no background noise.
- A video you already made. Any Kling 2.6 clip generated with sound on. Generate a character once, keep its voice, and every later clip can reuse it.
That third option is the interesting one. A recurring host, a brand mascot or a series character stays recognisable across a whole run of videos. A saved voice also performs in English and Chinese without being set up twice. That takes most of the work out of a localised version.
Binding is where mistakes cluster. Attach the voice to a character, never to a voice name standing alone, and never to an object or a sound effect. Put the tag beside the character rather than inside the quoted line, since a tag left trailing at the end attaches to the wrong person. Give each character its own voice, and never bind one to someone who does not speak.
Two limits are worth planning around. Two voices per clip is a hard cap, and 200 saved voices is the ceiling. A voice that contradicts the character will also fight the model, so a tall adult paired with a small child’s voice is really an instruction to resolve a conflict. Singing holds a voice less reliably than speech.
Kling motion control, explained
Motion control gets the least attention and solves the hardest problem. Supply a picture of a character and a video of someone moving, and the model transfers that movement onto the character. Choreography, gesture and timing come from real footage instead of a written description, and that is the part text prompting struggles with most.
The source video decides the result, so it needs to meet a few conditions.
- One unbroken take. No cuts, no camera moves, with the person visible throughout including limbs and head.
- Steady movement. Smooth beats fast. Very quick action can come back shorter than the source, since only the usable stretch gets extracted.
- Three seconds minimum of continuous motion.
- Matching framing. Full-body motion driven onto a half-body portrait produces exactly the mess it sounds like, and the character needs to fill a reasonable share of the frame.
Realistic and stylised characters both work, including humanoid animals and figures with roughly human proportions. One setting controls the ceiling: orienting the character to the reference video allows a source up to 30 seconds, while orienting to the still image caps it at 10. That 30 second route is the longest output Kling 2.6 produces. Motion control keeps the source video’s original sound rather than generating new audio, so treat it as a separate tool.
Generating sound on its own
Sound does not always need a video attached. Kling can produce audio by itself, either from a written description or by generating effects to match a video you upload. That earns its place when the video already exists. Phone footage, an older clip with unusable audio, or an animation needing a foley pass can all take generated sound without regenerating the footage, and it sidesteps the 1080p rule entirely.
Limits worth knowing before you start
- No 4K. 1080p is the ceiling, and it doubles as the requirement for any clip with sound.
- One shot per clip. Multi-shot generation is absent, so a cut sequence has to be assembled in an editor.
- No extension. A clip cannot be stretched past its original length after the fact.
- Missing controls. Camera control, motion brush and end-frame-only generation are all unavailable. A start frame is required whenever you work from an image.
- Two spoken languages. English and Chinese. Prompts in other languages are translated to English for the audio.
- 30 day retention. Generated files are cleared after 30 days, so download anything worth keeping.
Results also improve when a prompt does one thing well. Pick a single core idea, use a reference image that matches the description instead of contradicting it, and choose settings deliberately. Stacking three ambient layers and a two-hander into five seconds is the most reliable way to get a muddle back.
How to use Kling 2.6 in Picsart
Kling 2.6 is available in AI Playground, where it sits alongside the rest of the model lineup and runs against the same prompt as anything else. Comparing outputs side by side is the quickest way to learn which model suits a given shot, and it costs less than forming an opinion one generation at a time. The settings that matter most are the two that control sound, so set those before anything else.
- Open AI Playground and select Kling 2.6. It sits with the other video models, so the same prompt can be run against any of them for comparison.
- Set the resolution to 1080p. Sound is unavailable at 720p, so this comes before the audio setting rather than after it.
- Turn native audio on. The default is off, and this is the step most people miss on a first attempt.
- Choose the clip length. Five seconds suits a single action or an atmosphere shot. Pick ten for dialogue, singing or anything that needs a reply.
- Write the prompt in five parts. Scene, subject, movement, audio, then style. Quote any spoken lines and name every sound you want to hear.
- Generate, then listen before you look. Sound is the part most likely to need another pass, and a flat result usually traces back to a missing audio element in the prompt.
Get answers to common questions
Native audio works at 1080p only, and generation defaults to 720p with audio switched off. Both settings need changing. Selecting audio at 720p is an invalid combination rather than a lower-quality one.
Start with sound turned on
Kling 2.6 rewards a little setup. Move to 1080p, switch audio on, and write the sound into the prompt rather than hoping for it. The difference shows up on the first generation. Open AI Playground and give it a scene with something worth hearing.