Contents
Muse Image 1.0 is live in Picsart. Meta’s new image model is available now in AI Playground, and it is the first model in Picsart that works through your prompt before it draws anything: it plans the composition, looks up references it does not already know, builds any structured elements in code, and checks its own work before handing you a result.
That changes what you can reasonably ask for. Most image models take your words and render them in a single pass, which is why they guess at things they have never seen and why text inside an image so often arrives as letter-shaped smudges. Muse Image 1.0 treats a prompt as a task to work through rather than a description to match, so instructions with several moving parts survive the trip.
Below: what the model is, how its reasoning actually works, what it does well, how to prompt it, and where to find it now that it has landed.
What is Muse Image 1.0?
Muse Image 1.0 is Meta’s agentic image model, built by Meta Superintelligence Labs. One model covers the whole job. It generates images from a text prompt, edits images you already have, and composes new ones from several reference images at once.
What separates it from a standard generator is that it uses tools while it works. It can search the web for visual references and current facts, and it can write and run code to lay out charts, plots and QR codes accurately before placing them in the picture. It also reviews its own output as it goes, making a small correction when a detail is off, or starting again when something larger is wrong.
The practical effect is accuracy on the things image models usually fumble: real places, real products, legible text, and any instruction with more than one part to it.
How Muse Image 1.0 actually works
Ask a typical model for a poster with a headline, a subhead and a logo in the corner and you get an image that looks like a poster with gibberish printed on it. The model matched the mood, not the instruction.
Muse Image 1.0 breaks the request down first. It works out what goes where, resolves anything it needs to look up, renders, then checks the result against what you actually asked for. Where a conventional generator makes one pass from words to pixels, this one runs a loop, and the loop is where the quality comes from. Meta found that giving the model more time to think produces steadily better images, and that thinking harder beats simply generating more options and picking a favourite.
Three things happen inside that loop.
It looks things up
Point the model at a real landmark, product, logo or style and it can pull visual references rather than approximate from memory. Ask for something that depends on current information and it can go and find it instead of inventing a plausible answer.
This is the difference between an image that resembles a place and one that depicts it. A model working from memory alone produces a building that feels roughly Parisian. A model that can look first produces the building you named.
It builds structure in code
Some elements have to be correct rather than merely decorative. A chart’s proportions carry meaning. A diagram’s labels have to line up with what they label. A QR code either scans or it is a decorative square.
Because Muse Image 1.0 can write and run code, it constructs those elements properly and then places them in the image, instead of drawing an impression of what a chart looks like.
It corrects itself
The model reviews its own draft as it goes. When a small detail is wrong it makes a local edit. When something larger is wrong it starts that part again. When it is missing information it goes and finds it.
The interesting part is that Meta did not design this behaviour. It emerged during training, simply because a model that caught its own mistakes produced better images and was rewarded for it.
Editing with precision
Editing is where the reasoning shows up most plainly. Ask Muse Image 1.0 to clear fog from a landscape, remove someone from the background, restore a damaged family photo, or rewrite the text on a sign, and it changes what you named while leaving the rest of the frame alone.
That restraint is harder than it sounds. The common failure in AI editing is collateral damage: you ask for one change and the model quietly re-renders faces, shifts colours, or rearranges the background. Naming both halves of the instruction, the thing to change and the thing to protect, gives it a boundary it can hold.
Composing from several references
You can hand the model more than one reference image in a single prompt, and interleave your instructions between them, so each picture is captioned with the job it is doing. Use the pose from this one. The colour palette from this one. The room from this one.
That inline pairing is what makes complex composites tractable. Rather than attaching a folder of images and hoping the model infers your intent, you tell it what each reference is for.
The related trick is consistency across a set. Anchor a run of generations on a small group of references and the look holds from one image to the next. This is what makes campaign sets, product catalogues and multi-image social sequences workable, where the perennial difficulty has been keeping image five looking like image one.
Refining an idea across turns
The clearest demonstration of what the model is holding onto is a chain where each step depends on the last. Meta’s own walkthrough runs like this:
Start with two reference photos, a cat and a dog, and ask for them as best friends having a picnic on a sunny day, in a vintage 35mm style. Then ask to see that exact picnic photo as a framed print hanging on the wall of a cosy cafe, with a table and two empty chairs in front of it. Then ask for the front of the cafe, with its name, matching the vibe of the interior, with the framed photo visible through the window. Then design that cafe’s paper menu using its exact name, adding a “Picnic Special” with a small illustration of the same cat and dog. Finally, place that menu on the table from the empty-table shot made three steps earlier.
Nothing in that sequence is a fresh prompt. Each turn inherits the cafe, the animals, the style and the name invented along the way. That is the difference between a generator and something you can art-direct.
Prompting Muse Image 1.0
Because the model reads a prompt as an instruction rather than a mood, how you write changes the result more than it does elsewhere.
Caption each reference as you go. Label what every image is contributing rather than attaching several and hoping. Each reference gets a job.
Say what should stay the same. When editing, name the thing you want changed and the thing you want protected. “Add a red wool hat, keep the snowy forest background” gives the model a boundary.
Be explicit about numbers and details. If it matters that there are exactly five of something, say exactly five, and say it should stay five. Vague quantities are where any image model drifts.
Name the medium. Vintage 35mm, Korean manhwa, claymation, botanical engraving, isometric low-poly. A named style lands harder than an adjective.
Ask for the text you want. If words belong in the image, write them out exactly as they should appear, including the punctuation.
Let it think when it counts. The model has a reasoning setting. Give it room when the output has to be right, and dial it back when you are exploring quickly.
How to use Muse Image 1.0 in Picsart
Muse Image 1.0 is available now in the model picker in AI Playground, alongside the rest of Picsart’s AI models.
How to use Muse Image 1.0 in Picsart
1. Open AI Playground
Switch the mode toggle to Image.
2. Select Muse Image 1.0
It sits in the model picker under the prompt box, marked New.
3. Describe what you want
Upload an image first if you are editing rather than generating from scratch.
4. Set your shape and how many
Seven aspect ratios cover square, story, landscape and portrait, and you can return up to ten variations from a single prompt.
5. Generate, then keep going
Tell it what to change rather than rewriting the prompt, and it builds on what it already made.
The Advanced panel is where the model’s own behaviour is exposed: reasoning strength, whether it searches for images and facts, whether it uses its layout and chart tools, and your export format. The defaults leave everything on, which is what you want for most work. The case for switching the search tools off is speed on purely imaginative prompts, where there is nothing real to look up.
Five things to try first
- A launch kit for something imaginary. Invent a product, then build the packaging, a poster, a spec card and a social set, all anchored on the same references so they look related.
- A recipe or how-to card. A single image carrying legible steps, where the typography is part of the design rather than an afterthought.
- An event poster with a working QR code. The code is generated properly rather than pasted in, so it actually scans off the screen.
- A room, restyled. Photograph a space, ask for it in a different style, and keep the pieces you
liked from earlier versions as you iterate.
- An explainer diagram. Something with labelled parts and a sequence, where being readable matters more than being pretty.
Get answers to common questions
Muse Image 1.0 is an image generation and editing model from Meta Superintelligence Labs. It generates images from text, edits existing images, and composes images from multiple references, and it can search the web and run code while it works to get details right.
Start creating
Muse Image 1.0 is live in AI Playground. Pick it from the model picker, describe what you want, and let it work the problem before it draws.