AI image prompt builder
Choose the subject, style, shot, lighting and lens. You get a complete prompt, formatted the way your platform actually wants it — including the negatives.
Building…
The anatomy of an image prompt
A good image prompt is not a sentence. It is an ordered list of decisions, and the skill is knowing which decisions exist.
Most disappointing generations come from leaving a decision unmade rather than making it badly. If you do not specify lighting, the model picks one — usually the flat, evenly-lit look that dominates stock photography, because that is what it saw most of. Six choices cover almost everything:
Two optional extras earn their place on photographic work: a lens, which controls compression and depth of field, and a detail treatment, which decides whether the image reads as crisp, intricate, minimal or soft. Neither helps much on illustration.
Why the order matters
Image generators weight the beginning of a prompt more heavily than the end. This is not a quirk to work around — it is how attention distributes over a sequence — but it has a direct practical consequence: whatever you put first is what you get a picture of.
"Golden hour lighting, a woman reading in a cafe" reliably produces a sunset with a figure somewhere in it. "A woman reading in a cafe, golden hour lighting" produces a portrait that happens to be lit warmly. Same words, different picture.
The rule: subject first, then what kind of image, then how it is lit and coloured. Modifiers describe something the model has already committed to. The builder above always assembles in that order.
A second consequence is that long prompts dilute. Past roughly forty or fifty meaningful words, each additional term competes with the rest for influence, and the terms at the end get very little. If a prompt is not working, removing half of it is more often the fix than adding to it.
One prompt, built up a step at a time
Each row adds exactly one decision. Paste any of them into the builder above and you will get the same string back — nothing here is illustrative.
| Step | Prompt | What changes |
|---|---|---|
| Subject only | a woman reading a paperback |
A generic stock-photo person against nothing in particular |
| + setting | …in a rain-streaked cafe window |
Now there is a place, weather and a reason for the light |
| + medium | …, cinematic film still |
Stops being a photograph of a person and becomes a frame from something |
| + lighting | …, soft diffused daylight |
The biggest single jump. Flat stock lighting becomes deliberate |
| + lens | …, shot on 85mm portrait lens, shallow depth of field |
Background falls away; the subject is isolated and the framing tightens |
| + negatives | --no watermark, text, extra fingers… |
Removes the artefacts that mark an image as machine-made |
Five words became about thirty, and every addition answered a question the model would otherwise have answered for you, differently on each generation. Notice what is not in there: no "masterpiece", no "8k", no "trending on artstation", no "beautiful". Those are the terms that fill prompt guides and move nothing.
Style keywords that do something, and ones that do not
A keyword earns its place if removing it changes the picture. By that test, most of the vocabulary circulating in prompt guides fails.
| Does nothing | Why | Use instead |
|---|---|---|
masterpiece, best quality |
Appears in captions of every quality level. Carries no direction | A named medium: oil painting |
8k, ultra HD, 4k |
Describes a file, not an image. Often pushes toward plastic renders | sharp focus or a real focal length |
beautiful, stunning, amazing |
Subjective adjectives with no visual referent | Say what makes it so — the light, the colour, the composition |
highly detailed everywhere |
Fights any minimal or soft treatment you also asked for | Pick one detail treatment and commit |
trending on artstation |
Was a real signal on older models; largely inert now | Name the actual style: concept art, matte painting |
The reliable ones are all concrete: a medium (oil, watercolour, 3D render, photograph), a named movement or period (art deco, bauhaus, ukiyo-e, brutalist), a real technique (long exposure, double exposure, cross-processed), and physical materials (brushed steel, weathered oak, frosted glass). Each of those has a visual definition the model learned from images that genuinely looked like it.
The platforms are not interchangeable
The same prompt behaves differently on each generator, and one difference is actively dangerous to get wrong.
| Platform | Negatives | Prompt style it prefers |
|---|---|---|
| Midjourney | --no parameter |
Comma-separated keyword lists. Parameters trail the prompt. |
| DALL-E / ChatGPT | None — do not try | Natural sentences. Describe the scene as if to a person. |
| Stable Diffusion | Separate negative field | Keyword lists. Responds strongly to detailed negatives. |
| Leonardo | Separate negative field | Keyword lists, similar to Stable Diffusion. |
| Ideogram | Separate negative field | Handles text in images better than most. Say what the text says. |
| Flux | Separate negative field | Prefers natural sentences over keyword soup. |
The DALL-E negative trap. DALL-E and ChatGPT have no negative field, and adding "no text, no watermark" to the prompt frequently produces text and watermarks — the words are in the prompt, and the model is not reliably modelling the negation. Describe the positive instead: "a plain unmarked surface" rather than "no text". The builder above omits negatives entirely for those two platforms for exactly this reason.
Negative prompts, and the ones worth having by default
A negative prompt lists what must not appear. On platforms that support them properly it is the highest-leverage line in the whole prompt, because the recurring failures of image generation are consistent and therefore consistently excludable.
The defaults the builder adds cover three families:
- Anatomy failures — extra fingers, extra limbs, fused fingers, deformed, mutated, malformed. Hands remain the single most common giveaway.
- Artefacts — blurry, low quality, jpeg artifacts, oversaturated, duplicate. These are what makes an image read as machine-made at a glance.
- Unwanted overlays — watermark, signature, text, logo, cropped, out of frame. Training data was full of stock imagery, and it shows.
Add your own on top for anything specific to your scene. Excluding "modern cars" from a period street scene, or "hats" from a portrait series, saves more regenerations than any positive term you could add.
Do not overload it. A negative prompt of eighty terms starts constraining things you wanted. Twenty is plenty, and the ones that matter are the ones specific to your image.
Camera and lens language
Photographic vocabulary works because the training data is full of photographs captioned by photographers. A focal length is not decoration — it changes the geometry of the result.
| Term | What it does | Use for |
|---|---|---|
24mm wide-angle | Wide field, exaggerated depth, distortion at edges | Interiors, landscapes, drama |
35mm | Slight wide, natural. The documentary default | Street, reportage, environmental portraits |
50mm | Roughly how the eye sees. Neutral | General purpose, honest framing |
85mm, shallow depth of field | Compresses features flatteringly, blurs background | Portraits |
135mm telephoto | Strong compression, background pulled forward | Isolating a subject, sport |
macro | Extreme close focus, very shallow field | Texture, objects, insects |
tilt-shift | Selective focus band; makes scenes look miniature | Cityscapes, dioramas |
Two cautions. Lens language only helps on photographic styles — asking for an 85mm watercolour confuses more than it helps. And "bokeh" is worth knowing but easily overdone: "shallow depth of field" gives you the same effect with less of the glowing-orb look that marks an image as AI-made.
Lighting is the biggest single upgrade
If you change one thing about how you write image prompts, name the light. It separates a competent generation from a good one more reliably than any other term, and most people leave it out entirely.
Aspect ratio changes composition, not just size
Cropping a square to a widescreen and generating at widescreen produce different images. The model composes for the frame it is given: a wide frame invites environment and context, a tall frame invites a single subject filling it.
| Ratio | Where it goes | What it does to composition |
|---|---|---|
1:1 | Profile pictures, album art | Centres the subject. Neutral, contained |
16:9 | Banners, thumbnails, desktop | Invites background and setting |
9:16 | Stories, reels, phone wallpaper | One subject, full height, little context |
4:5 | Instagram feed | Portrait-friendly without going extreme |
3:2 | Prints, classic photography | The 35mm frame. Reads as photographic |
21:9 | Cinematic panoramas | Forces a landscape read. Strong but limiting |
Real people and trademarked characters
Naming an identifiable public figure or a trademarked character usually produces one of three outcomes: an outright refusal, a heavily distorted near-miss, or an account warning. Most major generators restrict it, and the restriction is getting tighter rather than looser.
The productive route is to describe the attributes rather than the person. Almost always, what you actually want is not that individual — it is the era, the styling, the wardrobe, the posture, the mood. Those are all describable, and describing them gives you something you can use commercially and iterate on.
The same applies to living artists. "In the style of [living illustrator]" is contested ground both legally and ethically, and it produces a weaker result than naming the actual visual properties you admire — the linework, the palette, the level of abstraction. Describe the technique, not the person who owns it.
Historical figures, long-dead artists and generic archetypes ("a Victorian naturalist", "a 1970s session guitarist") are unrestricted and usually give you exactly what you were reaching for anyway.
Seven mistakes that waste generations
- Adjective stacking. "Beautiful, stunning, amazing, masterpiece, best quality" moves almost nothing. Those words appear in every caption, so they carry no direction. Concrete nouns and verbs do the work.
- Burying the subject. Three clauses of style before you say what the picture is of. Lead with the subject.
- Contradicting yourself. "Minimalist, intricate, highly detailed" forces the model to pick, and you cannot predict which it drops.
- Negatives on DALL-E. Covered above, and worth repeating because it fails silently.
- Asking for text. Most generators still render lettering badly. Ideogram is the exception. Otherwise, add text afterwards in an editor.
- Prompting for a composition of many things. "A dog, a cat, a parrot and a goldfish each holding a different instrument" reliably fails. Generators are weak at binding several attributes to several subjects.
- Never changing the seed. If a prompt is nearly right, re-rolling the seed is often faster than rewriting. Conversely, fixing the seed is how you change one element while keeping everything else.
Common questions
Does prompt length help?
Up to a point. Under about fifteen meaningful words you are leaving decisions to the model; past forty or fifty each new term dilutes the rest. Most good prompts sit in between.
Should I use weights like (term:1.4)?
Only on Stable Diffusion and its descendants, which support the syntax. Elsewhere it is literal text in your prompt. Use it sparingly — heavy weighting distorts.
Why do my faces look plastic?
Usually over-specified quality terms. "8k, ultra detailed, hyperrealistic, perfect skin" pushes toward an airbrushed render. Add film grain, a real focal length, and natural lighting instead.
What is the difference between style and medium?
Medium is what it is made of — oil, watercolour, 3D render, photograph. Style is how it is treated within that medium — art deco, ukiyo-e, brutalist. Specifying medium matters more; it is the first fork in the road.
Can I reuse one prompt across platforms?
The descriptive part, yes. The formatting, no — parameters, negative handling and preferred phrasing all differ. Switch platform in the builder above and the same choices come out formatted correctly.
Glossary
- Negative prompt — a list of what must not appear.
- Seed — the number that determines the random starting point. Same prompt plus same seed gives the same image.
- Aspect ratio — the shape of the frame, which changes composition and not only dimensions.
- Depth of field — how much is in focus. Shallow isolates a subject; deep keeps everything sharp.
- Bokeh — the quality of the out-of-focus area.
- Chiaroscuro — strong contrast between light and dark.
- Rim light — light from behind that outlines the subject.
- Upscaling — enlarging a generated image and adding detail.
- Style reference — an input image supplying visual style rather than content. Better than any style keyword when the platform supports it.
- Parameter — a platform flag such as
--aror--no, distinct from the descriptive prompt.
Or describe it once and let it be written
The builder assembles what you choose. Frompting’s image generator takes a plain description and works out the style, lighting and negatives for you — then formats it for the platform you are using.
Try it freeMore free tools
All free, all instant, none of them need an account.