Seven AI image models received the same prompt at the same time. Seedream 4.0 won outright, and it won the hard way, by following every instruction while also producing the most photographic image in the set.
Three things the grid settled straight away:
- One model answered everything. Seedream 4.0 was alone in delivering the angle, the lens, the lighting and all five named details without trading one off against another.
- Two models dropped the subject. Ideogram V3 and Runway Gen 4 both returned a robot portrait with no cowboy hat, so the brief failed before any technical scoring began.
- Realism and accuracy are different scores. The most striking frame in the grid, from Flux Kontext, is also one of the least accurate.
Run one demanding prompt across seven models at once and the differences stop being a matter of taste. Here is the test, the scorecard and what each model actually did.
How the test worked
One Flow canvas. One text prompt wired into seven image nodes. Every node received identical instructions, and none of them got a second attempt, a reroll or a tuned variant. The whole test runs from the Image Model Comparison template, so you can reproduce it with your own brief.
The seven models, in the order they sit on the canvas: Gemini 2.5 Nano Banana, Seedream 4.0, GPT Image 1, Imagen 4, Ideogram V3, Flux Kontext and Runway Gen 4.
Running them side by side matters more than it sounds. Comparing models from separate sessions introduces different prompts, different days and different settings. One canvas removes all of that, so the only variable left is the model.
The prompt every model received
Written to be demanding on purpose. It names a subject, then adds explicit technical instructions that a model can visibly follow or visibly ignore.
Try this prompt
Hyper-realistic close-up portrait of a robot cowgirl with intricate metallic textures and lifelike features. Angle: Head-on frontal view at eye level to emphasize symmetry. Lens: 85mm with a wide aperture of f/1.4 for soft background bokeh. Lighting: Golden hour natural light combined with a soft diffused fill light, casting warm highlights on the metallic surface and subtly illuminating facial features. Details: Chrome and bronze finishes with visible bolts and joints, reflective skin, and digital readout eyes. Background: Slightly out-of-focus prairie sunset for a cinematic effect.
What each result was scored against
Judging image models on whether a picture looks nice produces an argument, not a comparison. Every verdict below traces to something the prompt asked for by name:
- Subject. A robot cowgirl, which makes the hat part of the brief and not a garnish.
- Angle. Head-on frontal view at eye level, symmetric.
- Lens. An 85mm look at f/1.4, meaning a soft, shallow background.
- Lighting. Golden hour plus diffused fill, with warm highlights landing on metal.
- Details. Chrome and bronze finishes, visible bolts and joints, reflective skin, and digital readout eyes.
- Background. A slightly out-of-focus prairie sunset.
The digital readout eyes turned out to be the sharpest test in the set. It is specific, it is checkable in a second, and it split the seven models cleanly.
The scorecard
| Model | Hat | Head-on | Golden hour | Readout eyes | Bokeh |
|---|---|---|---|---|---|
| Seedream 4.0 | Yes | Yes | Yes | Yes | Yes |
| GPT Image 1 | Yes | Yes | Yes | Yes | Yes |
| Gemini 2.5 Nano Banana | Yes | Yes | Partial | Yes | Strongest |
| Imagen 4 | Yes | Yes | Partial | Yes | Smooth |
| Flux Kontext | Yes | Yes | Yes | Glow only | Yes |
| Runway Gen 4 | No | Yes | No | Partial | Yes |
| Ideogram V3 | No | No | No | On the cheek | Yes |
How each model handled the prompt
Seedream 4.0
The result that wins the test, and the only one that answers every instruction without trading one off against another.
The face is genuinely skin. Pores, freckles and fine texture across the cheekbone, framed by chrome jaw and temple plating that reads as hardware bolted to a person rather than a mask laid over one. Inside each iris sits a teal dot-matrix readout, the literal instruction rendered exactly as written. Below the jaw it opens into the most detailed mechanical build in the set: exposed pistons, braided hoses, brass fittings and hex bolts running down the throat into a chrome chest plate.
The sun sits on the horizon at frame left, throwing a warm rim along every polished edge while the fill keeps the face readable. It also returned at 4K, four times the resolution of the rest of the grid.
Best for: anything where the prompt has to be followed exactly and still look photographed. See what Seedream 4.0 does.
GPT Image 1
The purest golden hour shot of the seven, and the closest challenger on lighting.
The sun is visible in frame at the right horizon. Warm light runs down the bronze faceplate, catches every seam in the neck assembly and rims the shoulder armor, while enough fill remains to keep the front of the face from going to silhouette. Orange LED segments sit where the eyes should be, so the readout instruction lands, and symmetry is close to exact.
The honest miss is the finish, which is bronze throughout rather than the chrome and bronze mix the prompt specified. It is also the least human of the results that kept the hat.
Best for: briefs where the lighting is the brief. See what GPT Image 1 does.
Gemini 2.5 Nano Banana
The most literal reading of the details instruction, and the best bokeh in the grid.
Every named detail is there and then some. Green pixel-matrix eyes with individual characters visible. Rivets, screws and panel lines across the entire faceplate. Copper tubing at the throat. A studded leather hat and a rope-textured bandana nobody asked for but which suit the brief. The background dissolves into distinct circular highlights, which is the f/1.4 instruction rendered more visibly than any other model managed.
What it gives up is the other half of the sentence. The prompt asked for metallic textures and lifelike features, and this face is entirely plate. It also reads as a 3D render rather than a photograph.
Best for: when the subject should read unmistakably as a machine.
Imagen 4
The most symmetric result, which is exactly what the angle instruction asked for. Head-on, eye level, near mirror-perfect down the centerline.
Copper and bronze plating covers the whole head, cyan character readouts sit in the eyes, and rivets run across the chest. The hat is present and correctly shaped.
Two things hold it back. The light reads as cool dusk rather than golden hour, so the warm highlights never really arrive. And the surface quality is smooth in a way that says rendered asset rather than photographed object.
Best for: frontal, symmetric, product-style shots. See what Imagen 4 does.
Flux Kontext
Visually the most striking frame in the grid, and the one that misses the most specific instruction in the prompt.
A gold faceplate with human proportions, blonde hair moving in the wind, a dark cowboy hat and glossy black shoulder armor, all lit by a strong warm sunset with the background thrown far out of focus. As art direction it is the one people stop on.
The eyes are a flat cyan glow with no characters in them, so the readout instruction is answered in spirit and not in fact. The face is polished metal rather than the reflective skin the prompt asked for, so the lifelike half of the subject goes missing.
Best for: when art direction outranks accuracy. See what Flux Kontext does.
Runway Gen 4
The one result that changes the color story, and one of two that dropped the hat.
There is no cowboy hat here at all. What arrives instead is a bare chrome cranium with exposed panel work, cyan eyes and headphone-style ear units, set against a wheat field with the sun at frame right.
The grade is the real departure. Where every other model ran warm, this one pushes the subject cool blue against an orange background. It is a deliberate-looking choice and a good-looking image, but it works directly against an instruction that asked for warm highlights on metal.
Best for: cold and synthetic rather than warm and cinematic. See what Runway Gen 4 does.
Ideogram V3
The largest gap between what was asked for and what came back.
The skin is excellent, arguably second only to Seedream, with real texture and a convincing blend of copper panels into the cheek and cranium. Taken alone it is a strong portrait.
Measured against the prompt it misses four instructions. There is no cowboy hat, so the subject is not a cowgirl. The head is turned into a three-quarter view when the prompt asked for head-on frontal at eye level. The light is flat daylight under a pale sky. And the readout moved off the eyes onto the cheek, where it reads as a small gold ticker.
Best for: portraits rather than costumed characters.
Which AI image model is best for your job
No single model wins every brief, and a comparison that produced one would be hiding something. Pick by what your brief cannot afford to lose:
- The prompt followed exactly. Seedream 4.0, then GPT Image 1.
- Photographic realism. Seedream 4.0 by a clear margin.
- Every mechanical detail visible. Gemini 2.5 Nano Banana.
- The light you described. GPT Image 1.
- Symmetry for a frontal shot. Imagen 4.
- Art direction over accuracy. Flux Kontext.
- A cool, synthetic mood. Runway Gen 4.
- Portrait work. Ideogram V3.
How to run this comparison on your own prompt
1. Open the Image Model Comparison template
The canvas arrives with seven image nodes already wired to a single prompt input, so nothing needs connecting.
2. Replace the prompt with your own brief
Write instructions a model can visibly fail. Name the angle, the lens, the lighting and at least one detail that is easy to check.
3. Select the frame and click Run
All seven nodes fire at once from the same input, so every result answers identical instructions.
4. Score each result against your prompt
Work through your own list of instructions rather than picking the image you like most. The gap between them is the finding.
Tips for running your own AI image model comparison
- Write instructions a model can fail. Vague prompts produce seven nice pictures and no information. Name the angle, the lens, the light and the specific details.
- Include one detail that is easy to check. The digital readout eyes sorted this entire grid in about four seconds of looking.
- Check whether the subject survived. Two of seven models dropped the hat here, the kind of miss that stays invisible until you compare against the prompt rather than against the other images.
- Change nothing between models. One prompt, one run, no per-model tuning. The moment you tune, you are comparing your patience rather than the models.
- Watch the output resolution. One model returned 4K while the rest returned 1K, which matters enormously for print or a hero slot.
- Run it on your real brief. A robot cowgirl is a stress test. Your product, your brand palette and your actual subject will sort the models differently.
Get answers to common questions
It is a test that sends the same prompt to several image models at once and judges the results against each other. Running them simultaneously removes the variables that make separate sessions impossible to compare.
Start your own image model comparison
Open the template, replace the prompt with your own brief, and run all seven models at once. Try the Image Model Comparison template.