HeyGen Video is a text-to-video model in Picsart AI Playground. It makes short clips with the sound already built in. It works from three kinds of input: a written prompt, an optional first-frame image, and up to 12 references in any mix of images, video, and audio. This guide covers each input and setting. It also compares HeyGen Video with an avatar tool and walks through a first HeyGen text to video clip.

What is HeyGen Video?

HeyGen Video generates picture and sound in the same run. Describe a scene, and the clip comes back with its audio included. Nobody has to add a voiceover or music track afterward.

Clips last 5 to 15 seconds. They come out at 480p or 768p, in seven aspect ratios plus an adaptive option. The model sits in Picsart AI Playground, so a finished clip can go straight into editing.

What sets this HeyGen AI video generator apart is how much it accepts besides text. A prompt alone works. Files can show what words struggle to describe, like a product’s exact shape or the mood of a soundtrack.

How HeyGen Video works, from input to clip

Every HeyGen Video clip starts from the same three ingredients. Only the prompt is required. The other two add control.

Input Required? What it controls
Text prompt Yes The action, the scene, and what should be heard
First-frame image No How the clip opens
References (up to 12) No How the subjects, settings, and sound look and feel

All of them go into one generation, which returns one clip. Video and audio arrive in a single pass. So a prompt that mentions rain, footsteps, or a café hum comes back with those sounds attached.

Think of it as a small production. The prompt is the director’s note. The first frame is the opening shot. The references are the cast, the set, and the score.

References explained

A reference is simply a file you upload to show the model what something should look or sound like. HeyGen Video takes up to 12 of them per clip. They can be images, video clips, audio files, or any mix of the three.

How references get their names

Each reference gets a name based on the order it was uploaded: Picture 1, Video 1, and so on. The prompt uses those names to point at a file.

That changes how a prompt reads. Skip “a matte sage-green ceramic mug with a thin handle and a speckled glaze.” The prompt can simply say “the mug from Picture 1.” The look lives in the file, so the sentence is free to focus on what happens.

What each type of reference is for

  • Images suit subjects and styling: a product, a character, an outfit, a poster.
  • Video clips suit settings: a street at night, a shopfront, a harbor.
  • Audio files suit the sound: a melody, a mood, a signature brand sound.

Habits that keep reference prompts readable

  • Introduce each file once. A short clause like “Picture 1 is a ceramic mug” shows anyone reading the prompt which file does what.
  • Give every file one job. Subject, setting, or sound. A file doing two jobs is harder to swap out later.
  • Write the sound as its own sentence. Background sound, small effects, and music are easier to change when they are not buried mid-paragraph.
  • Upload in the order you write. Names come from upload order. Upload the subject first and the setting second, and the prompt stays in step with the files.

HeyGen Video prompt examples

Three prompts that put references to work. Swap in your own files and keep the structure.

Product scene from two pictures

Picture 1 is a ceramic coffee mug in matte sage green. Picture 2 is a linen tablecloth in warm sand tones. Show the mug from Picture 1 on the tablecloth from Picture 2, beside a sunny kitchen window. Steam rises slowly as morning light drifts across the table. Slow push-in toward the mug. Sound: soft kitchen ambience, a spoon tapping the rim twice, birdsong outside the window.

A character placed in a video location

Use Video 1 as the setting: a rainy city street at night with neon reflections on wet pavement. Place the character from Picture 1, wearing a yellow raincoat, walking toward the camera under a clear umbrella. Keep the mood quiet and cinematic. Sound: steady rain drumming on the umbrella, distant traffic, footsteps splashing through shallow puddles.

A poster brought to life with an audio reference

Picture 1 is a hand-drawn travel poster of a coastal town. Bring the scene from Picture 1 to life: fishing boats rocking gently in the harbor, gulls circling overhead, laundry swaying in the breeze between pastel houses. Let the sound follow the mood of the audio reference, a light accordion melody. Add soft waves and gull calls underneath.

HeyGen Video settings explained

Three settings decide where a clip can go. They are worth choosing before generating rather than fixing afterward.

Aspect ratio. A vertical 9:16 suits Reels, Shorts, and TikTok. 1:1 and 3:4 fit feeds. 16:9 and 4:3 cover landscape screens, and 21:9 gives a wide cinematic frame. An adaptive option is there too.

Length. Anything from 5 to 15 seconds. Short loops and teasers sit at the low end. A scene with a small arc of action needs room toward 15.

Resolution. Two tiers, 480p and 768p.

Left untouched, the settings produce a 5-second clip at 768p in 16:9.

Setting Options in HeyGen Video
Resolution 480p, 768p (default)
Length 5 to 15 seconds, 5 by default
Aspect ratio 9:16, 1:1, 3:4, 4:3, 16:9 (default), 21:9, adaptive
Starting point Text prompt, with an optional first-frame image
References Up to 12, mixing images, video, and audio, named by upload order
Sound Generated together with the picture
Prompt help Prompt Enhancement, turbo or quality mode
Repeatability Fixed seed

Prompt Enhancement and seed explained

The first try is rarely the last. Two controls help with the next ones.

Prompt Enhancement sharpens a description before the model starts. It comes in a turbo mode and a quality mode.

A fixed seed brings back a take worth keeping. It also gives a steady base for improving that take, one prompt change at a time.

HeyGen Video vs the AI avatar video generator

Plenty of searches for “heygen video” come from people who want a digital presenter reading a script to camera. HeyGen Video in Picsart does a different job. It builds scenes with sound from a prompt, a first frame, and references. For a presenter speaking straight to the viewer, Picsart has a separate AI avatar video generator.

Feature HeyGen Video AI avatar video generator
Starts from A prompt, plus an optional first frame and up to 12 references A script, or a concept turned into a script
On screen Any scene the prompt and references describe A digital presenter speaking the script
Sound Generated with the picture The presenter delivering the script
Made for Social clips, product promos, reference-guided scenes UGC-style ads, social posts, product explainers
Where it lives AI Playground Its own tool, with the result added to the Picsart canvas

The two pair well. An avatar can introduce a product, and a HeyGen Video scene can show it in use.

Where to use HeyGen Video

The home base is Picsart AI Playground. There, HeyGen Video sits among more than 100 other models from the AI model catalog. It works in the browser and in the desktop app.

Teams that make lots of clips can reach the same model from the Picsart CLI, MCP, the REST API, and the SDK. A prompt tested in the Playground can then run inside an automated workflow.

Generated clips stay inside Picsart for finishing. Cut them down in the AI video editor, and reshape them for another placement with the video resizer.

What to make with HeyGen Video

A vertical clip for Reels, Shorts, and TikTok

Set 9:16, write the scene, and write what the scene sounds like. The clip arrives ready to post with its audio in place. Try a bakery opening its shutters at dawn, with trays clattering and a radio playing low. Or a skateboarder carving through an empty parking lot.

A promo built around a product photo

Upload the product shot as Picture 1 and let the prompt stage it. Think of a fragrance bottle on a marble ledge with water dripping nearby. Or a skincare jar opening by a bathroom window at golden hour. One generation becomes several placements after a pass through the resizer.

A scene assembled from a full kit of references

This is where the 12-file ceiling pays off. A small brand can load product shots, a clip of its shop, and an audio file with its signature sound. Then each new prompt can combine the same kit in a new way.

How to make a HeyGen text to video clip

Open HeyGen Video in AI Playground to follow along. Keep the AI video editor and video resizer open for the last part.

1. Write the scene and the sound

Describe what happens on screen and what the viewer should hear. Choose a length from 5 to 15 seconds and a resolution of 480p or 768p. Then pick the aspect ratio that fits where the clip will play.

2. Upload a first frame or references

An optional first-frame image sets how the clip opens. Up to 12 images, video clips, and audio files can also be added. The prompt refers to them by upload order as Picture 1, Video 1, and so on.

3. Generate the clip

Run HeyGen Video in AI Playground and the clip comes back with picture and audio together. Turn on Prompt Enhancement to tighten the description, or fix the seed to recreate a take and build on it.

4. Finish it in Picsart

Trim the clip, switch its aspect ratio for another channel, and add it to the rest of the project.


Get answers to common questions

It turns a text prompt into a 5 to 15 second video with sound. An optional first-frame image and up to 12 image, video, and audio references add control over the result.

Start your first HeyGen Video clip

Gather a few files and give each one a job. Write a line about the sound, and pick where the clip is headed. Open HeyGen Video in AI Playground and make the first take.