J.K. Rowling kept notes. Robert Jordan kept a manual file for every person, nation, city, and culture across 14 books and 4.4 million words. George R.R. Martin has admitted he had to rely on fans to remind him what color his characters’ eyes are, that he forgot Bran Stark entirely in a Jon Snow POV chapter, and that his mistakes bother him not because they’re embarrassing but because they undermine the deliberate unreliable narrator moments he wrote on purpose. Stephen King wrote the left arm broken in one chapter of It and the right arm broken in another. Conan Doyle moved Watson’s war wound from his shoulder to his leg between books.
These are not careless writers. They are writers who had to remember. The imageprompt field means you don’t.
The standard advice for character image prompts goes something like this: describe the hair color, the eye color, the height, the build, maybe a few clothing details. Load that into GPT image gen and see what comes back. Adjust if it’s wrong.
That’s not the architecture approach. That’s the accent layer approach, and it produces the same result as writing dialogue phonetically to indicate a Southern character. You get the surface. The surface drifts.
The difference between a description and an architecture prompt is the same difference Post 4 established for dialogue voice: one tells the AI what your character looks like, the other tells the AI who they are. A description produces an image that is probably correct on the first generation and possibly correct on the fifteenth. An architecture prompt produces an image that is recognizably the same person every time, because the generator is pulling from the same set of load-bearing facts rather than re-interpreting a list of physical attributes.
Here is what that looks like in practice.
What architecture looks like
This is Betsy Bishop’s imageprompt field, pulled directly from her character sheet in the Vampires of Tucson writing room:
A hauntingly regal woman with English ancestry, in her early 20s, stands alone on the balcony of a desert estate beneath a full moon. Tall, approximately six feet. Her features are elegant and severe: caucasian skin, high cheekbones, long black hair, storm-dark eyes that have seen centuries, a carved cane in hand with its handle shaped as a serpent devouring a coin. Every part of her evokes a woman once condemned, now ascendant.
A faint brand on her left wrist, half-hidden beneath fabric, suggests her past as both survivor and executioner. Her expression is calm, but unmerciful. A woman born of gallows smoke and reborn in blood. She does not forget. She does not kneel.
Note what is not in that prompt: a specific eye color. “Storm-dark eyes that have seen centuries” is not an eye color. It is 327 years of arc compressed into a visual descriptor. The generator will produce eyes that are dark and weighted and ancient every time, because that is what the field encodes. Not a color to remember. A truth to reproduce.
“Born of gallows smoke” is not an atmosphere note. It is her origin story. The serpent cane and the branded wrist are not decoration. They are checkpoints the generator uses to identify her across every image, every book, every generation. They are load-bearing history rendered as visual fact.
Compare that to “dark eyes, high cheekbones, pale skin, carries a cane.” Both descriptions generate the same first image. Only one generates the same person in Book 6.
Before you use Betsy as your template, one important note: her prompt is an architecture proof case, not a complete checklist. She is intentionally missing explicit height, build, and a direct eye color specification. That is a craft choice for her character. Your checklist needs all of those. I will get to that.
The field does two jobs
Here is the part the image generation tutorials never cover: the imageprompt field is not a file you hand to an image generator. It is a load-bearing field in the writing room.
In the Vampires of Tucson workflow, the Chapter Delivery prompt auto-pulls the imageprompt field when environmental grounding scores below a threshold during quality review. The field is not just for image generation. It is the reference the prose AI uses when it needs to know how this person occupies physical space. When the AI writes Betsy entering a room, it is reading from the same field that produced the image. The prose and the image are generated from the same source of truth.
That means the field has to do both jobs. It has to be specific enough to generate a consistent image. And it has to be written in a way that tells a writing AI how this character moves through the world.
Write it thin and your images come back inconsistent and your scenes go generic. Write it from the architecture and every generation, whether it is producing a chapter header image in GPT image gen or grounding a scene in prose, gets the same person.
One field. Three consumers. One source of truth.
The contract argument
GRRM has enough characters in A Song of Ice and Fire that he has admitted he cannot track them all himself. He leans on fans who run Westeros.org to verify details, and he has said that the mistakes that irritate him most are the ones that accidentally contradict moments he wrote deliberately. An inconsistency he intended looks indistinguishable from one he didn’t catch. The reader cannot tell which errors are his and which are the character lying. The unreliable narrator move he constructed gets lost in the noise of the actual mistakes.
Robert Jordan built the manual version of the solution: a detailed reference file for every character, every location, every cultural detail across 14 books written over 23 years. Fans who catalogued the errors in the Wheel of Time noted it was remarkable there weren’t more, given the scale. Brandon Sanderson, who completed the series after Jordan’s death, found additional errors and corrected them for paperback. Jordan did everything right. He still had errors. And he still died before the series was finished.
The imageprompt field is the automated version. It does not get tired. It does not forget. It does not die before Book 14. (It also does not need Brandon Sanderson to finish its series.)
You write it once. Every image generation loads it. Every scene-grounding pass loads it. The eye color does not change in Book 6 because you never had to remember it.
When a reader flags an inconsistency, you know it is yours, not a drift artifact. That distinction matters. Martin knows this. It is why the accidental mistakes irritate him.
The Beetlejuice principle and the variable fields
Tim Burton made a deliberate worldbuilding decision for Beetlejuice: the dead always appear in the clothes they died in. Every other ghost production of the period changed costumes scene to scene and got visual drift. Burton made it a rule, locked it in, and every scene respects it. Your imageprompt field is that rule. The clothing signature, the props, the physical markers: these are Beetlejuice decisions. You make them once, you encode them, and every image generation that follows respects them.
What you do not do is let the generator decide for you.
There is one place the Beetlejuice rule has a deliberate override, though. Clothing falls into two categories. If the outfit is who the character is (Betsy’s ivory colonial coat, Burton’s ghosts), it belongs in imageprompt. It is load-bearing identity, not wardrobe. But if the outfit varies by scene, arc, or book, it gets its own field: outfitprompt.
The outfitprompt field is an indexed wardrobe, not a single alternate. Each entry is a numbered variation you call by reference:
- Spring - light jacket, jeans, sneakers
- Summer - shorts, t-shirt, sandals
- Fall - heavy jacket, flannel, jeans, boots
- Winter - parka, sweatshirt, jeans, boots
“Use outfit 3” is unambiguous. “Use the fall outfit” requires interpretation. Number everything.
The same logic applies to hairstyleprompt and any other element that varies. A contemporary character might have “1. Working: hair up, tight bun. 2. Private: hair down, off duty.” A period character might have one hairstyle per social context. The governing principle is the same: separate the permanent from the variable. imageprompt holds what is always true. Every field beyond it holds indexed, deliberate variation.
Here is why the indexed fields matter beyond consistency: they make change carry narrative weight. If the AI has been rendering your character with hair up in every chapter header and grounding every scene with hair up for thirteen chapters, then when she lets her hair down in Chapter 14, the reader feels it. The contrast is visible because the baseline was locked. Without the field, the AI might have rendered her with hair down in Chapter 6 because it was a relaxed scene and it made a choice. The moment in Chapter 14 is gone. It already happened accidentally.
Locked fields do not just ensure consistency. They make intentional change mean something.
One more thing the outfit field encodes that most writers do not think about: cultural markers. A Wisconsin character wearing a sweatshirt with shorts in October is not making a fashion choice. That is a regional marker as specific as “bubbler” or “ope, sorry.” The outfit entry encodes it: “Fall - sweatshirt, shorts, sneakers.”
Now consider what happens when that same character moves to Tucson. October in Wisconsin is 50 degrees. October in Arizona is 85. Their internal thermostat was calibrated somewhere else. They walk outside into 85 degrees in a sweatshirt and shorts because their brain still says sweatshirt weather, and they do not think it is strange. The Arizona locals do.
That outfit entry is not clothing. It is Wisconsin still living in their nervous system 2,000 miles away. Same mechanism as Betsy’s ivory coat. Same mechanism as “might could” from a Memphis grandmother. The clothing is the artifact. The geography built the clothing choice. Post 4 established the same principle for dialogue voice. The writing room draws from the same row in both directions.
The AI co-authorship problem
Everything above applies to solo authors. It applies to AI co-authors with a specific additional consequence.
The Expanse series (nine mainline novels plus novellas) is written by Daniel Abraham and Ty Franck under the pen name James S.A. Corey. Two meticulous, coordinated writers sharing one character bible across years. Even with that level of intentional coordination, maintaining visual consistency across two separate mental images of every character is a structural problem, not a discipline problem. Two experienced professionals, working deliberately, still require explicit shared documentation to stay consistent.
Now consider the AI version: you and your AI are James S.A. Corey, except your co-author has no memory between sessions. None. Abraham at least remembers what Holden looks like when he sits down to write. Claude does not. Every session is a new co-author who has never read the previous books and has never seen the previous images.
Without the imageprompt field loaded, you are not co-writing with a partner who knows your character. You are co-writing with a partner who is guessing. Every failure mode in this article (King’s fog, Jordan’s scale, Martin’s cast, Watson’s wound) is present in every AI writing session by default. The field is the shared bible that makes the collaboration coherent.
This is not a metaphor. It is the mechanism. Load the field, and the AI knows the person. Do not load the field, and the AI invents one. It will probably be close. It will not be right. And it will not be the same close next session.
How to build it
Two steps.
Step 1: define your visual anchors. These are the things you see when you close your eyes and picture this character. Work through the checklist:
- Hair color and texture
- Eye color
- Skin tone
- Height and build
- Apparent age
- Clothing signature (if load-bearing)
- Signature items or props
- Distinguishing marks, scars, or physical details
These are your consistency checkpoints: the things that would be visibly wrong if the image came back different. Do not skip any of them. Gaps here are the generator’s invitation to invent.
Step 2: load the character’s full architecture (wound, ambition, voice, history, the whole row from the spreadsheet) and ask the AI: “Based on everything you know about this character, how would you describe them visually?” Let the AI synthesize the architecture into appearance language. Then merge the two outputs: your visual anchors as the skeleton (what must be correct), the AI synthesis as the flesh (what the character’s history made visible on their face, in their posture, in what they carry).
Neither step alone is sufficient. Generic image prompts are all Step 1. “Dark eyes, brown hair, medium build” tells the generator what to draw and nothing about why it matters. The synthesis step is what produces “storm-dark eyes that have seen centuries” instead of “dark eyes.” The architecture shows up in the image because you put it in the prompt.
End the merged result with the thesis sentence: the one-liner that names who this person is, not just what they look like. “She does not kneel.” That line tells the generator the register of every image. It tells the prose AI the register of every scene. It is not decoration. It is the anchor.
Two notes on apparent age:
For immortal or ageless characters, state apparent age explicitly. “In her early 20s” is in Betsy’s prompt because the generator does not know she has been 21 for three centuries. It will guess an age if you do not give it one. Apparent age is a separate concept from actual age, and the prompt needs both: the number the world sees, and the weight behind it.
For aging characters, the “write it once” rule has one deliberate override: write one imageprompt per era. A character at 25 in Book 1 who is 45 by Book 6 needs two prompts. Same base architecture. You add to it: the years, the injuries, the weight of what happened between books.
This is not a contradiction of the consistency argument. It is the advanced version of it. Deliberate change, locked per volume. Accidental drift is the failure mode. Intentional evolution is the craft.
Where do these fields live?
The variable fields (outfitprompt, hairstyleprompt, and any others you create) need a home. Two defensible answers.
Character sheet: clothing and hairstyle are properties of the person. They belong with everything else that defines who the character is. One place to look up everything.
Chapter sheet: clothing and hairstyle are scene-specific choices. What she’s wearing in Chapter 3 is a function of what’s happening in Chapter 3, not a permanent attribute of who she is.
For the Vampires of Tucson writing room, these fields live in the character sheet. That is the decision I made. This is a choice each author makes for themselves.
The failure mode is not choosing the wrong sheet. The failure mode is never deciding: ending up with the information in neither place, or both, or scattered across notes in three locations. Make the decision. Apply it consistently. There is no wrong answer.
Adding the field
Add the imageprompt column to the character spreadsheet from Post 3. Run the two-step build. Generate. Look hard at what came back. Revise the prompt, generate again, look again.
Do not rush this. Take multiple days. Take multiple weeks if the character warrants it. This is a write-once, reused-forever field. The investment asymmetry is real: a few hours of iteration now pays back every time you generate an image, write a scene, or open a document two years from now that needs to know what your character looks like.
The field is done when the image that comes back is the person you see in your head. Not close. Not mostly right. Them.
Blank template
imageprompt: [Visual anchors: hair color and texture, eye color, skin tone, height and build, apparent age, clothing signature if load-bearing, signature items or props, distinguishing marks.] [Architecture synthesis: what their history, wound, or ambition made visible in how they look and carry themselves.] [Thesis line: one sentence naming who they are, not just what they look like.]
outfitprompt:
1. [Context 1]: [clothing description]
2. [Context 2]: [clothing description]
3. [Context 3]: [clothing description]
hairstyleprompt:
1. [State 1]: [hair description]
2. [State 2]: [hair description]
Posts 3, 4, and 5 have been building the same thing from three angles: the spreadsheet (Post 3), the voice (Post 4), the image (Post 5). All three pull from the same row. All three encode architecture, not surface. The writing room knows your character the same way in every direction, and it will keep knowing them in Book 9 the same way it knew them in Book 1, because you made the decisions once and locked them in.
Martin had to rely on fans to remember the eye colors. You have a field.
You may also like: - Defining Character Voice - Building Believable Characters - “Wait, What Was I Doing?” or Why Your AI Starts Acting Like an Alzheimer’s Patient