LTX 2.3 AI Model | Audio-Driven Video Generation | Picsart
LTX 2.3: the audio-to-video AI model
Most video models start from a prompt or a still. LTX 2.3 starts from sound. Feed it speech, music or sound design and the audio becomes the conditioning signal - it sets the motion, the pacing, the scene structure and how hard the visuals push, frame by frame.
LTX 2.3 uses audio to guide what happens on screen. Voice, music, and sound effects can shape the structure, pacing, and motion of the video, making the visuals respond to the sound rather than simply playing alongside it
Go beyond basic lip sync
Instead of focusing only on mouth movements, LTX 2.3 uses audio to influence the wider scene. Speech, music, and sound effects can guide movement, timing, and visual structure, creating a video that feels connected to the audio from start to finish.
Get sharper details and better prompt results
LTX 2.3 introduces a rebuilt latent space and updated VAE designed to preserve fine details such as textures, hair, text, and edges. It also features a larger text connector for better handling of complex prompts, including multiple subjects, spatial relationships, and stylistic instructions.
Create video made for vertical screens
LTX 2.3 supports native portrait video, trained on vertical footage rather than relying on cropped landscape output. That means better framing and composition for mobile-first content such as Reels, Shorts, and TikTok videos. Bring in audio as WAV, MP3, M4A, or OGG and use it to guide your video. You can also add an optional starting image to give the model a visual reference for the scene.
Try LTX 2.3 in Picsart
Bring your audio into the AI Playground and turn it into a video with LTX 2.3. Add a starting image to guide the look of your video and use a prompt to describe the scene you want to create. Your audio helps shape the timing, movement, and overall flow of the result.
What LTX 2.3 Audio to Video is good for
Turn music into visuals that follow its rhythm, tempo, and intensity. It works well for music videos, lyric videos, performance clips, and experimental visualizations.
Compare LTX 2.3 audio to video model with other video models for cinematic motion, storytelling, and social clips.
LTX 2.3 Audio to Video AI model FAQ
LTX 2.3 is an AI video generation model from LTX. It supports text-to-video, image-to-video, and audio-to-video generation. With audio-to-video, you provide an audio track and LTX 2.3 generates visuals synchronized to it
LTX 2.3 can use an audio track to generate the video itself, rather than only matching mouth movements to existing footage. Dialogue, music, or ambient sound can be used as the input, with the generated visuals synchronized to the audio.
LTX supports WAV, MP3, M4A, and OGG audio files. You can also provide an image to use as the first frame and a text prompt to describe the video or guide how that image should be animated.
LTX 2.3 comes in Fast and Pro variants. Fast is designed for faster generation and supports text-to-video and image-to-video. Pro supports those modes as well as audio-to-video and additional video editing capabilities. Audio-to-video with LTX 2.3 is available through the Pro model.
LTX 2.3 supports portrait and landscape video, with resolutions up to 4K depending on the generation mode. It also supports native 9:16 vertical video generation.
Seedance 2.5 supports 20 languages with matched lip-sync, and can localize the setting, characters, and dialogue so one production works for audiences worldwide.
Seedance 2.5 doubles clip length to 30 seconds, replaces 2.0's over-saturated glow with softer, physics-accurate lighting and more natural eye and facial detail, adds realistic impact physics and 3D-reference texturing, and raises reference support from 15 to 50.
Yes. Videos created through Picsart's tools powered by Seedance 2.5 can be used for marketing, social media, brand content, and other commercial applications, subject to Picsart's terms of use.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Up to 1080p with strong physical realism and optional reference image.Reference input1080pPro qualityCinematicSee model
Naturalistic 720p video with lifelike motion and character detail.CinematicVideo generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Up to 1080p with strong physical realism and optional reference image.Reference input1080pPro qualityCinematicSee model
Naturalistic 720p video with lifelike motion and character detail.CinematicVideo generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Up to 1080p with strong physical realism and optional reference image.Reference input1080pPro qualityCinematicSee model
Naturalistic 720p video with lifelike motion and character detail.CinematicVideo generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Up to 1080p with strong physical realism and optional reference image.Reference input1080pPro qualityCinematicSee model
Naturalistic 720p video with lifelike motion and character detail.CinematicVideo generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Up to 1080p with strong physical realism and optional reference image.Reference input1080pPro qualityCinematicSee model
Naturalistic 720p video with lifelike motion and character detail.CinematicVideo generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.