Skip to content

How to write AI image prompts that behave the same way twice

A prompt structure that survives contact with a real model, what a negative prompt is for, and why the same words behave differently in Midjourney and Flux.

By Wajahat Rasoul ·

A structure that holds up

Most unreliable prompts are unreliable for the same reason: they are a pile of adjectives with no clear subject. A prompt that reproduces reliably tends to have the same parts, in roughly this order.

  1. Subject. What the picture is of, stated plainly. One subject, not three.
  2. Action or state. What it is doing, or how it is arranged.
  3. Setting. Where, and when.
  4. Composition. Shot type and framing — close-up, wide, overhead, eye level.
  5. Light. Often the single highest-leverage phrase in the whole prompt.
  6. Style and medium. Photograph, illustration, 3D render, and the character of it.
  7. Technical parameters. Aspect ratio, model flags, seed.

Being specific about the right things

Length is not the goal. Specificity in the places that determine the image is.

VagueSpecificWhy it helps
beautiful lightinglow side light, long shadowsNames a physical setup the model can render.
high qualitysharp focus on the subject, soft backgroundDescribes an outcome rather than asking for virtue.
modern styleflat vector illustration, two-colourIdentifies a medium with real visual consequences.
nice coloursmuted sage and terracotta paletteConstrains hue instead of leaving it to chance.

Words like “masterpiece”, “8K” and “award-winning” are mostly inherited folklore. They rarely hurt, and they rarely do what people think — the specific clauses are doing the work.

Negative prompts

A negative prompt lists what should be absent. It is a real control on models that support it — Stable Diffusion derivatives and Flux interfaces, for instance — and is either unsupported or expressed differently elsewhere. Midjourney has its own parameter rather than a separate field, and several hosted models have no equivalent at all.

  • Use it for recurring faults, not as a wish list: extra fingers, text artefacts, watermarks, distorted faces.
  • Keep it short. A hundred negative terms constrains the model into blandness.
  • Do not negate the subject. Excluding “cup” while asking for a coffee cup produces exactly the confusion it sounds like.

The prompt generator writes a matching negative prompt alongside the main one and omits it for models that do not take one.

Why the same words behave differently

Each model was trained on different data with different captions, so the same sentence lands differently. The practical differences worth knowing:

  • Syntax. Midjourney takes trailing flags such as an aspect ratio parameter; most others take plain sentences and separate fields.
  • Sentence handling. Some models respond better to natural sentences, others to comma-separated tags. It is worth testing a prompt both ways once.
  • Text rendering. Some models place legible text well; others produce convincing gibberish. If the image needs words, this decides which model you use.
  • Default aesthetic. Several models apply a house look unless you steer them away from it explicitly.

That is why the prompt generator asks which model you will actually run before it writes anything, rather than producing one generic prompt and hoping.

Iterating without starting over

The most common mistake is rewriting the whole prompt after a disappointing result, which throws away the parts that worked.

  1. Change one clause at a time, and see what moved.
  2. Fix the seed if the interface allows it, so you are comparing prompts rather than comparing randomness.
  3. Keep the versions that nearly worked. The prompt generator keeps recent prompts in your browser for this reason.
  4. When something is close, stop prompting and finish it properly — crop and resize beat another twenty generations trying to get the framing by luck.

Working backwards from an image

When you can see the look you want but cannot name it, describing an existing image is faster than guessing. The image to prompt tool reads a reference and writes a prompt for it, with three modes: the whole image, the style only, or the subject only.

The style-only mode is the useful one for building a consistent set — take the lighting and treatment from a reference, then swap in your own subject. Two caveats: the result is a description, not a recipe, so it will not reproduce the reference exactly; and it deliberately does not name artists or brands, which is both a licensing matter and a habit worth keeping.

Two things to be careful about

Rights

What you may do with a generated image depends on the model's terms, not on the prompt. If the destination is a stock platform, check its rules on AI-generated content and on disclosure before you invest a day in a set — several agencies have specific requirements, and a few decline such work entirely.

Consistency across a set

A series is far harder than a single image. Keep one prompt skeleton, vary only the subject clause, hold the style and lighting clauses identical, and reuse the same model and settings throughout. If the series is destined for stock, the metadata for it is its own job — titles and keywords that hold up.

Tools for this

Everything here runs in your browser — no upload, no account.

Related guides