How to write AI image prompts that actually get you what you want
Most bad AI images come from a vague prompt, not a bad model. Here is the anatomy of a prompt that works: subject, setting, style, light, framing, in that order, with worked before and after examples.

The core idea in one paragraph
An AI image prompt is not a wish and it is not a search query. It is a set of instructions that a model reads left to right, weighting the early words far more than the late ones. When your prompt fails, the model is almost never broken. You have not specified enough about the things that matter, in the wrong order, or buried the important instruction under decoration. The fix is a small structure you reuse every time: subject first, then setting, then style, then light, then framing. Get those five in the right order and your results jump immediately, on any model, with no extra skill required.
Who this is for
If you have ever typed a sentence into an image generator and gotten back something that was almost right but not quite, this is your starting point. You do not need a new tool. You need a repeatable structure for the words you already type.
Why most prompts fail
Here is a prompt I see constantly, in roughly this form:
a beautiful woman in a city, cinematic, 8k, highly detailed, masterpiece, trending on artstation
It produces a perfectly competent image. It also produces roughly the same image for everyone who types it, because none of those words are decisions. Beautiful is not a decision. Cinematic is a vibe word. 8k, highly detailed, masterpiece are quality escalators that older diffusion models needed and current ones largely ignore. You have told the model how good the picture should be without telling it what the picture actually is.
The result is that you, the person with the idea, have handed the creative decisions to the model. The model is happy to make them for you, using the average of everything it has seen. That average is exactly the flat, generic look everyone complains about. The model is not being lazy. You gave it no constraints.
The five part structure that fixes it
A prompt that works is built from five slots, in this fixed order. You fill them deliberately. If you cannot fill one, you leave it empty rather than fill it with filler.
- Subject who or what the image is of. The most important words in the whole prompt go here, first.
- Setting where the subject exists. Background, location, time of day, environment.
- Style the visual treatment. Photographic, oil painting, anime, 3D render, a named artist or movement.
- Light the single most underrated lever. Direction, quality, color of the light falling on your subject.
- Framing how the camera sees it. Shot type, lens, angle, aspect ratio.
Notice the order mirrors how a photographer actually works. You decide what you are shooting, where, in what treatment, under what light, and from where the camera sits. People who shoot real photographs already know this sequence. Prompting is the same discipline with different inputs.
Before and after: the same idea, two prompts
Take one intention: a tired nurse at the end of a long shift. Here is the typical way it gets prompted:
tired nurse, hospital, realistic, detailed, 8k
That will give you a nurse in a hospital, probably young, probably pretty, probably standing, lit by flat overhead light, framed at eye level. It is correct and it is forgettable. Now the same intention through the five slots:
a tired nurse in her forties, slumped against a wall in an empty hospital corridor at 4am, documentary photography, mixed fluorescent and a single warm emergency exit light, shot on a 35mm lens at eye level, shallow depth of field
Every word in the second version is a decision. The age and posture are decisions. Four in the morning is a decision that dictates the emptiness and the light. Documentary photography pulls the treatment away from glamour. The mixed fluorescent and warm exit light is the difference between a clean stock image and something that feels lived in. The 35mm lens and shallow depth of field tell the model how space should collapse around her.
Same model. Same intention. The only thing that changed was that someone made decisions and put them in the right order.
Why word order matters more than people think
This is the part most guides gloss over. Image models are built on transformers and attention mechanisms. Attention is not equal across a prompt. A token near the front of the prompt is attended to by a larger set of later tokens than a token near the end. The practical effect: the model treats your first phrase as the spine of the image and treats later phrases as modifiers. Put your subject first and it owns the image. Bury it after three adjectives and the adjectives own the image instead.
I tested this mental model on a deliberately bad ordering. Same words, scrambled:
cinematic, shallow depth of field, 35mm, documentary, a tired nurse in her forties slumped against a wall, hospital corridor at 4am, mixed fluorescent light
The subject still lands because it is a strong phrase, but the image becomes more about the cinematic documentary treatment than about her. The face gets more idealized, the framing tighter, the mood more filmic. The model weighted the front loaded style words as the primary instruction and treated the nurse as the subject of that film. Order changed meaning.
How long should a prompt be
There is a fashion for very long prompts, and a counter fashion for very short ones. Both miss the point. A prompt should be exactly as long as the number of decisions you have made, and no longer. The model has a finite attention budget. Every decorative word you add dilutes the words that matter, because attention is roughly a fixed pie divided across more slices.
Here is a rough way to think about the math. Suppose a model attends to your prompt with a fixed pool of attention that we can call 100 units. In a ten word prompt that contains five real decisions and five filler words, each decision gets roughly 10 units of attention and each filler word steals 10 units it did not earn. In a thirty word prompt with the same five decisions, those decisions now compete with twenty five other tokens and each one gets roughly 3 units. You have not given the model more information. You have given it more noise and less weight per decision.
The lesson is not that short prompts are better. It is that every word should be load bearing. If you cannot say why a word is in your prompt, cut it. The quality escalators (masterpiece, trending, highly detailed, 8k) almost never earn their place on a modern model. They were useful crutches in 2022. They are filler now.
A test you can run in two minutes
Take your last prompt and read it out loud. Every time you hit a word, ask: did I choose this, or did I copy it from a list? If you copied it, delete it and generate again. You will usually find the image improves, because you removed attention that was competing with your real decisions.
The vocabulary you should actually build
The five slots only help if you have words to fill them. Most people are weak in two places: light and framing. They can describe a subject and a setting all day, but when they get to light they type good lighting and when they get to framing they type nothing. That is exactly backwards, because light and framing are where most of the visible improvement lives.
Build your vocabulary deliberately, one slot at a time:
- Light learn ten lighting setups by name and what they do. Golden hour, blue hour, Rembrandt, split, butterfly, softbox, rim, volumetric, low key, high key. Each one is a complete mood in a single phrase.
- Framing learn the shot types (extreme wide, wide, medium, medium close, close, extreme close) and a few lens behaviors (wide angle distortion, telephoto compression, shallow depth of field, deep focus).
- Style collect the names of treatments and the artists or films that embody them. Documentary photography is more useful than realistic because it implies a whole tradition of how real subjects are shot.
Notice that all three of these are nouns with meaning, not adjectives with vibes. The model can act on Rembrandt lighting because Rembrandt lighting is a defined thing. It cannot act on nice lighting because that is an opinion.
The checklist before you hit generate
If you only take one thing from this, take the habit of running your prompt through five questions before you generate. It takes ten seconds and it will save you more time than any tool:
- Is my subject in the first phrase, named specifically?
- Have I said where this happens, including time of day?
- Have I named a style or treatment, not just an adjective?
- Have I described the light, in a word that means a real lighting setup?
- Have I said how the camera frames it?
If you can answer yes to all five, your prompt is in the top tier of what most people type, and you have not used a single trick. You have just made decisions and put them in order.
What changes when you switch models
Different models read prompts differently, and a small amount of model awareness prevents a lot of frustration. The five slot structure holds across all of them, because it reflects how the underlying attention mechanisms work, but the dials move.
Older and more fine tuned diffusion models tend to reward keyword style prompts with comma separated phrases, because they were trained heavily on caption data that looked that way. Newer transformer based image models, especially the ones that also do language, reward more natural sentences because they parse grammar and weighting from structure, not just tokens. The practical rule: if your model is also a chat model, write to it like a brief. If it is a dedicated image model with a long lineage of fine tuning, the keyword style still works well.
None of this changes the five slots. It only changes whether you join them with commas or with prose.
A starting template you can steal
If building from scratch feels heavy, start from this skeleton and fill in the brackets. It encodes everything above. Over time you will stop needing it.
[specific subject, with age or defining detail] in [specific setting, with time of day], [named style or treatment], [named lighting], [shot type] on [lens], [depth of field or other camera behavior]
Example, filled in:
a retired fisherman in his seventies, scarred hands, mending nets on a wooden dock at dawn, documentary photography, soft side light through morning fog, medium shot on an 85mm lens, shallow depth of field
That is six decisions in twenty five words. Every one of them is load bearing. There is not a single quality escalator in it. And it will produce an image that belongs to you, not to the average of everyone who ever typed fisherman, detailed, 8k.
The mindset shift
The deepest change here is not a technique. It is who is making the decisions. A weak prompt asks the model to invent the image. A strong prompt tells the model the image you already see in your head, in enough detail that it has very little room to invent. The model is a brilliant, fast, obedient renderer of decisions. It is a poor substitute for a director. The job is yours.
Once you internalize the five slots, prompting stops feeling like guessing. It starts feeling like briefing a very talented collaborator who simply needs you to be specific. That is the whole game. Make decisions, put them in order, and cut anything that is not a decision. Everything else in this blog builds on that foundation.


