The surest way to fail at a full set of visuals from a generative model is to carefully rewrite the prompt for every image.
However good each prompt is, the images drift from one to the next. You’re describing the same thing again every time, and no description comes out the same twice.
Two things that actually work
① Lock the style block. Lighting, color, material, contrast and camera language go into one fixed paragraph that every image shares, word for word. Only the line about the subject changes from image to image.
② Condition on reference images. Feed the palette and existing assets in as part of the input instead of describing colors in words. Color described in text is the least stable part of any prompt.
With those two changes, the drift went from “every image is different” to “now and then one gets thrown out.”
Repeat the qualities, not the shapes
Our first version wrote the rules for light very specifically. Out of a few dozen images, half ended up using the same trick.
It didn’t look like a set of assets. It looked like a filter.
The fix was to split the rule into two layers. What the light is like (cold, dim, always leaving a visible physical effect) stays locked. What carries the light (leaking through a gap, falling on a wall, passing through frosted material, reflecting off metal) rotates from image to image.
⭐ Identity comes from repetition, but the thing to repeat is the quality. Fix the shape and by the fifth image people read it as a template.
One more hard line
⛔ We don’t generate faces to stand in for our team or clients, and we don’t generate interface screenshots to pass off as real work.
Generated images handle mood and material. Evidence is always the real thing. This is about credibility more than taste. One “client case study” that someone spots as generated drags down everything real sitting next to it.


