Contents
MiniMax Music v3 turns a written description of a song into a finished recording, sung or purely instrumental, in one pass. You describe the genre, the mood, and the feel you are chasing, and the model composes the melody, arranges the instrumentation, performs the vocal, and produces the mix. The track arrives whole.
That last part is what separates this from a clip generator. A four-bar loop leaves you with stitching work, a tempo to match and an ending to invent. A complete arrangement leaves you with a song you can drop straight into an edit. That is why the model fits a reel, a brand film, an ad, and a song you simply wanted to exist.
This walkthrough follows the path from one text box to one audio file. What belongs in a style prompt, how section tags set the shape of the song, what changes when you hand over your own lyrics, and which export settings to lock before you generate.
What goes into a style prompt
The input is plain language, not notation. A style prompt describes the music the way a producer would describe it to a session player. The more of that description you supply, the less the model has to guess. Seven elements do most of the work.
| Part of the prompt | What it sets | Example phrasing |
|---|---|---|
| Genre and subgenre | The instrumentation and the production style | Progressive house, baroque pop, nu-disco |
| Tempo | The pace, written in BPM | 126 BPM |
| Key | The tonal centre the song sits in | B-flat major |
| Emotional arc | How the feeling changes from section to section | Reflective verses rising into a euphoric chorus |
| Voice | Range, texture, and delivery of the lead vocal | Breathy male tenor with restrained phrasing |
| Instrumentation | Which instruments carry the track | Side-chained synths, club bass, crisp drums |
| Mix | The space and finish around everything else | Wide hall reverb, polished festival mix |
Not every prompt needs all seven. Genre, mood, and tempo alone will get you a track. The rest are the levers for when a first generation lands close but not right. Naming an emotional arc is the most useful addition. It tells the model the song has to travel rather than hold one feeling throughout.
Specialist vocabulary is optional. A Prompt Enhancement System sits behind the box. It expands a short description into a fuller, musically coherent instruction, drawing on established arrangement language, so a plain sentence still produces a properly structured track. Writing “sad piano song” gets you further than it should.
Uplifting progressive house, 124 BPM, A-flat major, reflective verses that open into a euphoric chorus, warm male tenor with a slight rasp, layered synths over a punchy club bass, crisp four-on-the-floor drums, wide and polished festival mix Late-night lo-fi hip-hop, slow and unhurried, dusty vinyl crackle under a soft electric piano, muted bass, brushed drums, the feeling of a quiet apartment after a long day, no vocalsA full-formula prompt
A scene prompt, no jargon
Lyrics do more than supply words. Section tags written into the lyrics define the song’s macrostructure, so the arrangement knows where the verses end and the chorus lands. The tags available are `[intro]`, `[verse]`, `[pre-chorus]`, `[chorus]`, `[bridge]`, `[instrumental]`, `[solo]`, and `[outro]`.
This is the control most people miss. A block of unbroken lyrics asks the model to infer a structure. The same lyrics with five tags in them hand over a blueprint. The difference shows up in the places songs usually fall apart: the lift into a chorus, the drop out of a bridge, and the ending.
Lyrics with structure tags
[Intro] [Verse] Morning on the fire escape, coffee going cold / the city sounds like static through the glass [Pre-Chorus] And I have been counting all the ways to say it [Chorus] Say it out loud, say it out loud / before the light gets old [Verse] [Chorus] [Bridge] Quiet now, quiet now [Final Chorus] [Outro]
Vocals, instrumentals, and your own lyrics
The vocal question is a decision you make in the prompt, and it changes what comes back. Ask for vocals and the model writes a lead line and performs it. Ask for an instrumental and it strips the track to a backing bed, which is the version you want under a voiceover, a product demo, or a piece to camera. Both paths produce a complete arrangement. The instrumental is not the vocal version with the singing muted, it is a track built to stand on its own.
Leave the lyrics field empty and MiniMax Music v3 writes its own words to fit the style you asked for. Fill it in and the model performs your lyrics word for word. Take that second path when the song has a job to do. Naming a product, landing a campaign line, and carrying a message that cannot be paraphrased all call for exact words.
The style prompt still runs the show when you supply lyrics. Your words set what gets sung. The prompt sets how it gets sung, down to timbre, phrasing, and techniques like breathiness and falsetto, plus the harmony arrangement stacked around the lead. Short and rhythmic beats long and literary, because lines that scan and phrases that repeat give a performance something to hold onto.
Pop, hip-hop, electronic, cinematic, lo-fi, and ambient are all in range, so the same switch serves a chorus-driven single and a slow orchestral cue equally.
An instrumental bed
Cinematic ambient instrumental, no vocals, 70 BPM, slow swelling strings over a low synth pad, sparse piano notes, patient and spacious, built to sit under a voiceover without competing with it
Why a five-minute song is hard
Length is the easy part of long-form music. Coherence is the hard part. A song that runs five minutes has to build and resolve, bring instruments in and take them out at the right moments, and keep every section pointed in the same direction. Generating more audio is trivial next to keeping a musical identity intact while it plays.
MiniMax Music v3 handles complete songs as long as five minutes, which is the difference between a backing bed and an actual record. The arrangement carries a beginning, changes, and an ending, so the parts of a track that normally eat an afternoon of editing are already decided when the file lands.
None of this replaces working with a musician. A composer writing to picture, a session player, or a licensed track from a catalog are all still the right answer for plenty of projects. AI music sits alongside them. It is the option for the days when you need something original in minutes and nobody is going to score a 20-second reel for you.
How the model turns a description into audio
Three stages sit between your sentence and the file, and knowing them roughly makes prompt decisions less mysterious.
The first is description. Rather than reducing a whole track to one label, the model works from structured captions that describe the music moment by moment. These cover genre, tempo, time signature, key, and production character. They also follow the emotional contour, note where every instrument arrives and drops out, watch the groove and the low end develop, and mark how the singing shifts from one section to the next. Everything you write in the prompt feeds this layer, which is why specific phrasing survives into the arrangement and vague phrasing does not.
The second is modeling, and it runs at two scales at once. One model tracks the song as a whole and keeps the structure stable across five minutes. A smaller one works inside each frame and fills in acoustic detail. Musical information itself is stored in layers. The first layer holds the core structure and semantics, and the layers above it add progressively finer sound. That is what keeps a long track from drifting as it plays.
The third is rendering. Instead of handing off a simplified set of tokens, the system carries the continuous output of both models straight into audio generation, which preserves detail that would otherwise be rounded away. Performance techniques such as glissando and legato survive the trip, and instruments stay separable in dense arrangements. The low end lands without swallowing the mix. Vocals keep pronunciation and breathing that read as performed rather than assembled.
The export settings you control
Encoding is yours to configure. File format, sample rate, and bitrate all get decided before the track exists, so nothing has to be converted after the fact. Each of the three does a different job, and knowing which is which saves a re-render.
| Setting | What it controls | Why it matters |
|---|---|---|
| Format | The file type the audio is written as | Decides what your editor, player, or upload target will accept |
| Sample rate | How many times per second the audio is measured | Governs how much high-frequency detail the file can carry |
| Bitrate | How much data each second of audio is given | Trades file size against how close the export stays to what you heard |
How to generate a track in the Picsart AI Playground
MiniMax Music v3 sits in the Picsart AI Playground, reachable from a browser or the Picsart desktop app, next to more than 170 other AI models. The full capability breakdown sits on the MiniMax Music v3 model page. Once a track exists, it goes straight into a cut in the AI video editor, or under footage you build with the AI video generator.
1.
Open the AI Playground Nothing to install and nothing to configure, in a browser or the desktop app.
2.
Pick MiniMax Music v3 from the model list The same prompt box runs every model, so switching between them is a single click.
3.
Write the style prompt Name the genre, the tempo, and the emotional arc at minimum, then add the voice, the instrumentation, and the mix character if the first pass lands close but not right.
4.
Add your lyrics, or leave the field empty Supplied lyrics get performed word for word, and section tags inside them set where the verses, the chorus, and the bridge fall.
5.
Set the audio encoding Lock the format, sample rate, and bitrate you need for wherever the track is going.
6.
Generate and listen all the way through Endings and transitions tell you more about an arrangement than the first eight bars do. "Reflective verses into a euphoric chorus" gives the arrangement somewhere to travel, where "happy" gives it one note to hold. Five tags turn a block of words into a blueprint, and choruses stop arriving in unexpected places. Say vocals or say instrumental, because leaving it open means accepting whichever one the description implies. Range, texture, and delivery shape a lead vocal as much as the production sitting behind it. An `[instrumental]` or `[solo]` tag builds breathing space in on purpose rather than by accident. Decide where the track is going first, then set the format, sample rate, and bitrate to match. MiniMax Music v3 is the newest music model from MiniMax, and it runs inside the Picsart AI Playground. One written prompt is enough to get back a fully produced track, sung or instrumental.How to generate a track in the Picsart AI Playground
Tips for a better first track
Name an emotional arc, not just a mood
Tag the sections in your lyrics
Answer the vocal question explicitly
Describe the voice, not only the genre
Write a rest into the track
Choose the encoding for the destination
Get answers to common questions
Start with a genre and an arc
One sentence is enough to hear what this model does with an arrangement. Name a genre, name where the feeling travels, decide whether anyone is singing, and let the rest arrive finished. Write your first style prompt in the Picsart AI Playground.