We shot the same person twelve times. Six of them were not the same person.
Two descriptions, six scenes each. What holds a face together between shots is the words you repeat, and one of the five things we asked for never turned up at all.
Experimentquality·FreeImgGen Team·Updated ·Tested on Z-Image Turbo
The short answer
A text-to-image model remembers nothing between shots, so an AI photoshoot holds together only because you described the same person twice. We ran twelve shots on 2026-08-12: six repeating five identity markers word for word, six with a loose description. The detailed six read as one person across a cafe, a street, a studio and a park. The loose six are six different women with similar hair.
MethodSix scenes, each generated twice: once from a description carrying five identity markers, once from a loose description of the same woman. Twelve images, one generation each, Z-Image Turbo, 2026-08-12, nothing re-rolled. Median 11.5 seconds, range 7.6 to 42.0. These six are the marker set.
cafe, morning
Jacket, glasses and hair arrive exactly as written, which is the whole mechanism: the description is the only thing carrying identity.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, seated at a cafe table by a window
A different angle and distance, and the face still reads as the same person. Nothing was carried over from the previous shot except the sentence.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, walking along a city street, three quarter view
A small mole appears high on the forehead. We asked for one above the left eyebrow, so the mark is honoured and the position is not.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, standing against a plain studio backdrop
The hardest lighting change in the set and the least faithful result: the face is thinner here than in the other five.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, at a desk in an office, side light from a window
Flat light, and the mole has moved to the other side of the face. Still recognisably the same woman.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, outdoors in a park on an overcast day
The tightest crop in the set, which is where the missing marker shows: the skin is clear, with none of the freckles we asked for.
Photograph of a woman in her early thirties with short dark curly hair, round tortoiseshell glasses and a mustard yellow jacket, close-up portrait in evening light
We asked for five specific things in every one of the six detailed shots: short dark curly hair, round tortoiseshell glasses, a mustard yellow jacket, a mole above the left eyebrow, and freckles across the nose.
Hair, glasses and jacket came through in all six. The mole appeared in the three shots framed closely enough to see one, but never twice in the same place, wandering from the forehead to the opposite brow. The freckles appeared zero times out of six, including in the close-up where they would have been unmissable.
That pattern is worth more than the pass rate. Markers that are objects survive. Markers that are small skin detail get averaged away, because the model is drawing a plausible face rather than assembling one from a checklist.
The loose set is six different women
The other six shots used the same scenes and a description with nothing distinguishing in it. Every one returned a woman with long dark curly hair, and beyond that they share nothing: different face shapes, different ages, different clothing in each frame.
MethodThe same six scenes with the identity markers removed, one generation each, Z-Image Turbo, 2026-08-12. The description given was a woman in her early thirties with dark curly hair, and nothing else.
cafe, morning
The scene is right and the person is an invention of this frame alone.
Photograph of a woman in her early thirties with dark curly hair, seated at a cafe table by a window, morning light
Some products keep a character stable by conditioning each new image on a reference picture or a locked seed. Our generator exposes neither. Each request starts from noise and your sentence, and that is the whole of its memory.
So the practical method is the one this test measured. Write the person once, as a block of concrete markers, and paste that block unchanged into every shot. Change only the scene.
When you need tighter continuity than words can give, work from a single portrait instead and move it through the editor. An edit preserves the face by construction, which is exactly what we found while measuring outfit swaps: everything the instruction does not name comes through untouched. Scenes are harder that way than clothes, so expect to fight the background.
Ad
Markers that work, and markers that waste your words
Objects hold. Glasses, a jacket, a hat, a particular colour: these are things the model has seen labelled thousands of times, and they came back in every frame.
Skin detail does not hold. Freckles, a specific mole, a scar in a specific place. Ours were either ignored or relocated, and adding more of them lengthens the prompt without buying consistency.
Age and hair are somewhere in between. Early thirties and dark curly both survived in all twelve shots, but they are broad enough to describe thousands of different people, which is precisely why the loose set drifted.
When you want a camera instead
If the AI photoshoot is meant to be of you, this is the wrong tool. Nothing here makes a picture of a specific real person, our terms rule out that use, and a headshot that is not your face is worse than no headshot.
If the pictures are going on a product page, photograph the product. An invented model in an invented jacket is fine for a mood board and misleading in a listing.
Where this does earn its place is everything that needs a plausible person rather than a particular one: a slide deck, a wireframe, a blog header, a storyboard. For that, twelve shots costs nothing and a photographer costs a day.
Run the marker block yourself
This is the exact description we repeated across six scenes. Change the last clause to move her somewhere else and the face should follow.
How do I keep the same face across several AI images?
Write the person as a fixed block of concrete markers and repeat it word for word in every prompt, changing only the scene. In our twelve-shot test on 2026-08-12 the six shots using five repeated markers read as one person, while six shots with a loose description produced six different women.
Why does the face change every time I generate?
Because a text-to-image model keeps nothing between requests. Each image starts from noise and your prompt, so any detail you leave unspecified is re-invented. Hair colour alone will not hold a face together.
Which identity details actually survive?
Objects do: glasses, a jacket, a colour, all six out of six in our set. Small skin detail does not. The mole we asked for moved around the face and the freckles never appeared in any of the six shots.
Can I make an AI photoshoot of myself?
Not on this site. We do not generate specific real people, and our terms rule out using the service on someone in a deceptive or non-consensual way. Products that do this train on your uploaded selfies, which is a different transaction with different privacy consequences.
How long does a twelve-shot set take?
Ours took a median of 11.5 seconds per image on 2026-08-12, with the slowest at 42 seconds when the queue was busy. There is no credit counter, so re-running a scene that came out wrong costs only the wait.