Skip to content
Sign inStart free
FeaturesPricingAPIAboutBlogChangelogStart free. 5 images on us
Sign inContactChangelog
Back to blog
ConsistencyNovember 7, 202410 min read

Character consistency in AI images: keeping one face across a set

The hardest and most requested skill in AI images is keeping the same character across many frames. A protocol built on prompt discipline, seeds, character references, and a decision tree for when each tool applies.

By The AIUnmark Team

The core idea in one paragraph

Character consistency is the single hardest and most requested skill in AI image generation, and almost everyone approaches it the wrong way. They try to describe the same face with more and more words, and the model gives them a different person every time, because words are a terrible way to specify a face. The fix is to understand that identity lives in four separate layers, and that the more of those layers you lock, the more consistent your character becomes. Prompt discipline handles the first layer. Seeds handle the second. Character references handle the third. Training handles the fourth. A real protocol uses them in order, escalating only when the previous layer is not enough.

Why this is worth getting right

Agencies, brands, illustrators, and anyone making a series all hit the same wall: the same character has to survive across dozens of images. The people who solve it do so with a repeatable system, not luck. This is that system.

Why words alone cannot hold a face

Start with why the naive approach fails, because the reason points to the solution. A face is a high dimensional pattern of proportions, bone structure, skin, and asymmetry. The English language has almost no vocabulary for those specifics. You can say thin nose, strong jaw, blue eyes, freckles, and that narrows the space somewhat, but it still leaves the model free to interpret those words across thousands of plausible faces. Run the same prompt a second time and you get a different plausible face that also matches those words.

This is not a model weakness. It is a fundamental limit of using text to specify something text was never designed to specify. The way out is to stop relying on text alone and start using the tools that carry identity directly: image references, seeds, and trained adapters. Each of those encodes the face in a far denser representation than language ever could.

An abstract illustration of a single identity being preserved across many different variations.
Consistency is not one technique. It is four layers stacked. The more layers you lock, the more the same face survives across different scenes.

The four layers of identity

Think of consistency as four independent locks. Each one you engage tightens the result. Engage them in order of cost and power.

  1. Prompt discipline. A fixed, detailed, reusable description of the character, used verbatim in every prompt. The cheapest layer, and the foundation everything else builds on.
  2. Seed locking. Reusing the same random seed so the model starts from the same noise and lands in the same region of face space. Cheap and powerful within one model and prompt.
  3. Character reference. Feeding an image of the character back to the model so it copies the actual face, not your description of it. The biggest single jump in consistency on models that support it.
  4. Training. Fine tuning a small adapter, often called a LoRA, on a set of images of your character so the model knows them by name. The most powerful and the most work, reserved for characters you use at scale.

Layer one: prompt discipline, done properly

Before any tool, your character needs a written identity block that you never change. This is not a loose description. It is a fixed phrase, stored somewhere you can copy it, and pasted into every prompt unchanged. The discipline is in the never change part. If you paraphrase the description between generations, you have broken consistency at the cheapest layer before reaching any of the expensive ones.

A strong identity block is specific along the dimensions the model can actually use. Age, ethnicity, build, hair color and style, eye color, a defining facial feature, a recurring clothing choice, and any marks like scars or tattoos. Vague blocks like young woman, pretty give the model maximum freedom and therefore maximum drift. Specific blocks like a woman in her early thirties, Korean, collarbone length black hair, single eyelid, small scar through left eyebrow, wearing a faded olive field jacket constrain the model into a narrow, repeatable target.

The through line principle

The identity block is a through line that appears in every prompt unchanged. Only the action, setting, and framing change around it. Treat the block as a literal string you copy and paste, not a sentence you reword.

Layer two: seed locking and what it actually does

A seed is the starting random noise an image model builds from. The same prompt plus the same seed plus the same model produces the same image, deterministically. This is why seed locking is the cheapest powerful consistency tool: it costs nothing and it removes a huge source of variation.

What a seed does for character consistency is subtler than people assume. The seed does not encode your character. It encodes a starting point in the model's latent space. When you hold the seed and prompt constant, you stay in the same neighborhood of that space, which tends to reproduce the same face because the model is solving the same denoising problem from the same starting noise. Change the prompt even slightly, though, and the seed pulls you to a related but not identical result. The seed is a strong stabilizer within a near constant prompt, and a weak one across very different prompts.

The practical use is to find a generation you like, note its seed, and reuse that seed for variations. Many models expose the seed in their output or let you set it directly. When you want the same character in a new pose or scene, lock the seed and change only the surrounding prompt. You will get drift, but far less than seeding fresh each time.

Layer three: character references

This is where consistency takes its biggest single jump. A character reference, sometimes called a character reference image or by model specific names, lets you hand the model an actual picture of your character and say use this face. The model copies the identity directly from pixels instead of inferring it from words, and pixels carry thousands of times more information about a face than any prompt.

The reason this works so well is that the reference image is a complete specification. Every proportion, every asymmetry, every skin detail is present. The model does not have to interpret strong jaw across a range of possibilities. It has the exact jaw. Models that support character references typically let you control how strongly the reference influences the result, usually as a weight from zero to one. A high weight copies the face faithfully but can make every image feel like the same photograph with a new background painted behind it. A lower weight keeps the identity while letting pose, light, and expression vary more naturally.

The workflow is to generate one strong, clean, front facing portrait of your character first, using layers one and two. That portrait becomes your reference image for every subsequent generation. Keep the reference image stable. If you swap reference images between generations, you will reintroduce drift, because each reference pulls toward a slightly different face. One canonical reference, reused everywhere, is the discipline that makes this layer hold.

The canonical portrait deserves real care, because every later image inherits its qualities. Aim for a neutral, evenly lit, front facing headshot with the identity block fully visible: hair, eyes, any defining marks. Avoid heavy expression, dramatic light, or costume, because those bake themselves into the reference and then fight you when you want different ones later. Think of the canonical portrait as a passport photo for your character, not a hero shot. Boring and complete beats dramatic and incomplete here, because the reference's job is to be a clean source of identity, not a finished image in its own right.

Words
Cheapest. Constrains the model. Cannot hold a face alone.
Seed
Cheap. Stabilizes within a near constant prompt. Weak across very different prompts.
Ref
The biggest jump. Copies identity from pixels. One canonical reference, reused.
Train
Most powerful and most work. Reserved for characters used at scale.

Layer four: training a character adapter

For a character you will use across hundreds of images, a brand mascot, a comic protagonist, a recurring model, training a small adapter becomes worth the effort. The adapter, often a low rank adaptation, is trained on a set of images of your character and teaches the model to reproduce that specific identity when triggered by a keyword. Once trained, you can summon the character with a single word and get strong consistency without a reference image in every prompt.

The cost is a training set of perhaps fifteen to thirty images of the character from different angles, in different light, with different expressions, all labeled consistently. The quality of the training set determines the quality of the adapter. Garbage in, garbage out applies forcefully here: a training set of similar angles and flat light produces an adapter that only knows the character from that angle and that light. Diversity in the training set is what gives the adapter flexibility at inference time.

This layer is overkill for a one off project and essential for ongoing work. The decision is simply volume. If you will generate dozens or hundreds of images of one character, the training cost amortizes down to almost nothing per image. If you need the character once, stop at layer three.

A decision tree for which layer to use

Because the layers have different costs and power, you should escalate deliberately rather than reaching for the heaviest tool first. Use this decision sequence.

  1. One image only? Stop at layer one. A single generation does not need consistency with anything.
  2. A few variations of one scene? Lock the prompt verbatim and reuse the seed. Layers one and two are usually enough for small variations.
  3. The same character across different scenes and poses? Generate one clean canonical portrait, then use it as a character reference for every subsequent image. This is where most professional work lives.
  4. The same character at very large scale, or with extreme angle and expression range? Train an adapter. Layers one through three start to strain when the reference image cannot cover the angle or expression you need, and training closes that gap.

What still breaks, and how to handle it

Even with all four layers, a few things remain hard and are worth knowing in advance. Hands and small props drift more than faces, because the model has less consistent training data for them. Complex interactions between two characters tend to blend their features, a known failure where face A bleeds into face B. Extreme angles the reference did not cover will produce a plausible but not identical face, because the model is extrapolating beyond what it was given.

The professional response to these is acceptance plus selective rework. Generate more candidates than you need, keep the ones where the face held, and rework the ones where it slipped. Consistency at scale is a numbers game softened by good technique, not a guarantee. The aim is to raise your hit rate from one in ten to eight in ten, not to reach ten in ten. That improvement is the difference between an unusable workflow and a production ready one.

Building a character sheet

One practice ties all four layers together and is worth adopting from day one: maintain a character sheet. This is a single document per character that holds the verbatim identity block, the canonical seed, the path to the canonical reference image, and the trigger keyword if you trained an adapter. Anyone on your team can open the sheet and reproduce the character without guessing, because every layer is specified. The character sheet is what turns an individual's skill into a team capability.

Character consistency is genuinely hard, and the people who make it look easy are not using better models. They are using a better system. Build the identity block, lock the seed, choose a canonical reference, escalate to training only when volume demands it, and keep a character sheet that survives across projects. Do that, and the same face will follow you across a hundred images, which is the foundation every series, brand, and story actually needs.

Start free. 5 images on us

More from the blog

Use cases

AI prompts for backgrounds and environments with depth and atmosphere

Composition and camera

Aspect ratio and framing: how shape changes the image

Styles

Cinematic AI images: film stock, grading, and lens character