How to generate AI video from the terminal with the Picsart CLI

SKILLS7 minIntermediate

Create video with Kling, Sora, Veo, Runway, and Luma. Animate images, extend clips, add soundtracks, and chain sequences.

How to generate AI video from the terminal with the Picsart CLI

What you'll learn

  • Generate video from text prompts or animate existing images
  • Choose the right video model for speed vs. cinematic quality
  • Extend 5-7 second clips into longer sequences
  • Attach AI-generated soundtracks to your video clips

What is CLI video generation?

CLI video generation creates AI video clips directly from your terminal. The Picsart gen-ai CLI supports text-to-video (T2V), image-to-video (I2V), and video-to-video (V2V) workflows across models like Kling, Sora, Veo, Runway, and Luma. You can animate a static image, extend a 5-second clip into a 30-second sequence, or attach a soundtrack. Think of it as Final Cut Pro's timeline, but the AI handles the animation and editing.

Common use cases

  • Social reels: Animate product photos into 10-second TikTok or Instagram clips
  • Hero videos: Turn a static hero image into a looping background for landing pages
  • Video ads: Generate multiple variations of a 15-second ad to A/B test
  • Storyboarding: Mock up camera movements and scene transitions before production
  • Explainer content: Animate diagrams or charts for educational videos
  • Extended sequences: Chain 7-second extensions to create 30-60 second clips

Generate AI video step by step

STEP 1: Open the CLI

  • On web: Go to picsart.com/cli → Install if you haven't already
  • On mobile: Not applicable — video generation runs in terminal only
Get the CLI

STEP 2: Choose video generation mode

Pick your input type and model:

  • Text-to-video (T2V): Generate from a prompt — use Sora for cinematic quality, Kling for faster turnaround
  • Image-to-video (I2V): Animate a static image — point to your file with --image hero.webp
  • Video-to-video (V2V): Transform an existing clip — use --video input.mp4 to restyle or extend
  • Extension mode: Add --extend to chain 7-second clips into longer sequences

STEP 3: Run the generation

Type gen-ai generate with your prompt, model, and optional image input. Video models take 30 seconds to a few minutes depending on length and resolution. The CLI shows progress updates. Your file saves to the current directory when done.

STEP 4: Review and refine

Check the video quality and motion: Not perfect? Generate again with a more specific prompt, switch models, or use --extend to chain multiple clips for longer sequences.

  • Verify smooth motion and no glitchy frames
  • Check that camera movement matches your prompt direction
  • Look for character consistency if animating people or objects
Start generating video

Tips for best results

💡 Use Sora for cinematic quality

Sora produces the highest visual fidelity with natural motion and lighting. It's slower (2-3 minutes per clip) but worth it for hero videos or polished social content. Use Kling if you need faster turnaround and can accept slightly less polish.

💡 Describe camera movement in your prompt

Add phrases like "camera slowly pans left," "zoom into subject," or "drone shot rising upward" to control motion. Without camera direction, the AI picks a default movement that might not match your intent.

💡 Chain extensions for longer clips

Most models generate 5-7 seconds per run. Use gen-ai extend to stitch multiple generations into a 20-30 second sequence. The CLI maintains character and scene consistency across extensions. Add --soundtrack to attach an AI-generated music track afterward.

💡 Animate images for product demos

If you have a hero product photo, use --image to animate it instead of generating from scratch. The AI adds realistic motion while keeping your product centered. Perfect for turning static catalog shots into social reels.

Frequently asked questions

Use the --image flag with gen-ai generate and point to your file. Example: gen-ai generate --image hero.webp --prompt "camera slowly zooms in" -m kling. The AI animates your image based on the prompt description. This works with any JPG, PNG, or WebP file. The result is a 5-7 second video clip saved to your current directory.

Sora produces cinematic-quality video with natural motion, accurate lighting, and high visual fidelity. It takes 2-3 minutes per clip and costs more credits. Kling generates faster (30-60 seconds) with good quality but slightly less polish. Use Sora for hero content or final deliverables, Kling for quick mockups or social media drafts.

Yes, but most models generate 5-7 seconds per run. Use gen-ai extend to chain multiple clips into a longer sequence. The CLI stitches them together while maintaining character and scene consistency. Example: gen-ai extend --input clip1.mp4 --count 4 creates a ~28-second video from four 7-second extensions.

After generating your video, run gen-ai generate --type audio --prompt "upbeat electronic music" to create a soundtrack, then use a tool like ffmpeg to merge them. Or use the --soundtrack flag during video generation (if your model supports it). The CLI auto-generates music that matches your video length.

Ready to create video?

Install the CLI and start animating images or generating cinematic clips from text prompts.

Get started