Scene-to-image prompts that stop fighting your novel
Four copy-ready prompt templates for illustrating fiction — establishing shot, character close-up, action freeze-frame, quiet interior — plus the documented quirk of each major image model and the fix for it: GPT Image's fixed size grid, Nano Banana's preference for narrative prompts over keyword lists, Seedream's front-loaded attention. Checked against the official docs on 2026-07-18. Ends with the one thing no prompt solves: keeping the same face across fifty images.

The first illustration most readers generate from a novel scene looks like a stock fantasy book cover: centered subject, teal-and-orange grade, nothing from the actual paragraph. That failure has a boring cause. The prompt asked the model to invent a scene instead of binding it to one. Your passage already contains the subject, the action, the weather and the hour — what it lacks is a style contract. This post is the resource-manual half of that fix: four templates you can copy, then the documented quirk of each major model, because the same template lands differently on different engines.
What makes a fiction illustration prompt different?
It has two layers with different owners. The scene layer is your passage, pasted whole — not summarized, because summaries shed the concrete details (the wet cobbles, the second lantern) that make an illustration feel like it belongs to the book. The style layer is a reusable block you write once: art style, lighting plan, composition, and a negative list. Keep the style block byte-identical between scenes and your book stops looking like an anthology by twelve artists. In the Foreverse reader the highlight-to-image flow does the pasting for you and anchors the result back into the page; the templates below work there, and they work anywhere you can paste text.
The working set
Four templates cover most of what long fiction asks for. Replace the bracketed slot, keep the rest — including the negative lists, which look paranoid until you skip one and get a watermark baked into a scene you liked.
The single-beat rule in template 3 is stolen from video generation, where compound actions are the top failure source; it turns out still images fail the same way, just less visibly — two overlapping actions give you three arms.
What binding actually changes: one worked example
Take a plain passage: "She waited under the station clock, coat soaked through, while the last train's lights slid away into the fog." Prompted bare, most models return a generic woman under a generic clock in golden hour, because golden hour is the statistical default of pretty. Run it through template 1 and three things change. The stated facts become constraints: soaked coat, fog, night, one departing train — the "invent nothing that contradicts" line forbids the sunset. The composition line moves her to the middle distance, so the scene reads as loneliness rather than a portrait. And the negative list quietly removes the failure noise: no text on the clock face, no watermark, no second figure. None of that required prompt-craft talent. It required a template that treats the passage as a contract instead of a mood board, which is the whole argument of this post in one image.
The same passage also shows where the two layers split responsibility. Swap the style block from painterly editorial to flat woodblock print and you get a different book on the shelf, but she is still under that clock in the fog. Change the passage and keep the style, and the gallery still hangs together. When an output disappoints, you can now tell which layer to fix: wrong facts means the scene layer got summarized, wrong vibe means the style block drifted.
Each model's quirks, and the fix
We checked the vendors' own documentation on 2026-07-18; dated claims below come from those pages, not from memory.
GPT Image (OpenAI lists gpt-image-2 as the current default for new work) is the instruction-follower of the three, and the one to pick when the image must contain readable words — a shop sign, a headline, a handwritten note. The official prompting guide says to put literal text in quotes and spell tricky words letter by letter — advice worth taking, because misspelled signage ruins an otherwise good plate. The quirk: output sizes come from a fixed grid (square, portrait, landscape), so your composition instruction should agree with the size you request — asking for a sweeping panorama inside a portrait canvas is a self-inflicted wound. Landscape for template 1, portrait for template 2, and the conflict never comes up.
The Nano Banana family — Google's name for Gemini's native image models, with the Pro tier supporting 2K and 4K output — is the strongest at conversational editing and multi-reference work. Its image generation docs make one point that matters enormously for fiction: describe the scene narratively instead of stacking keywords. A keyword list ("rain, alley, neon, cinematic, 8k") underperforms a flowing description on this family. The fix costs nothing here, because your passage already is a flowing description — resist the urge to compress it into tags.
Seedream, ByteDance's image line, is the batch workhorse: the official release notes document 2K generation in a few seconds — a tenfold speedup over its predecessor — and up to a dozen reference images per request. Third-party prompting guides converge on the same two observations: it weighs the start of the prompt more heavily, and it likes concise prompts. The fix: lead with the subject sentence, put the style block after, and trim ruthlessly. One honest note from our own template packs — on fast batch models we tell users to check outputs for stray logos and watermark-like artifacts before saving, and that advice stands regardless of engine.
The thing no prompt solves
Style pins with a template. A face does not. If the same character must survive fifty illustrations, the mechanism is reference images shown to the model on every generation — and the discipline that matters is who controls the reference set. In Foreverse, the set contains only images you confirmed by hand; generated images never promote themselves into it, which is what stops small deviations from compounding into a stranger. The measured results and the two-person-photo case are in the selfie piece. And if you want scene art without writing prompts at all, the theater's stage backdrops render from the passage on stage — behind an explicit cost confirmation, reused free when you revisit the scene.
Start with template 1 on the chapter you're reading tonight. If the second image comes back in a different style than the first, you paraphrased the style block — paste, don't retype. That one habit ends most of the drift people blame on the model.
FAQ
Can I describe my character in text and get the same face every time?
No. Text prompts converge style — palette, brushwork, lighting — but not identity. Two renders from the same 60-word character description will give you two different people who dress alike. The reliable mechanism is reference images: a confirmed set the model is shown on every generation. That is how Foreverse handles it, with a strict rule that generated images never auto-join the reference set.
Which image model should I default to for fiction scenes?
Depends on the job. Batch scene plates on a budget: a Seedream-class model, which ByteDance documents as producing 2K images in a few seconds, a tenfold speedup over its predecessor. Scenes that need legible in-image text, like a shop sign or a letter: GPT Image, whose prompting guide covers exact text rendering. Iterative editing and reference-heavy work: the Nano Banana family, which takes multiple input images and supports conversational edits.
Do I have to write a fresh prompt for every scene?
You shouldn't. Split the prompt in two layers: the passage itself supplies subject, action and setting; a reusable style block supplies palette, lighting, composition and the negative list. Keep the style block verbatim between scenes and only the passage changes. In Foreverse the selection you highlight fills the scene layer automatically, so the style block is the only part you maintain.
Why does my art style drift between chapters?
Usually because the style is being re-described from memory each time — paraphrasing your own style block is enough to shift the output. Pin one block and paste it unchanged, negative list included. If you generate inside a reader app, pick one wrapper template and stay on it for the whole book; changing wrappers mid-book is the style-drift equivalent of switching narrators.
Questions or ideas? Join our Discord →