Which LLM continues a novel best? We tested two genres — the winners don't overlap
Nine LLMs each continued two Chinese novels for 20 consecutive rounds — 360 rounds total, ranked by double-blind review. DeepSeek V4 Flash won the fantasy epic; DeepSeek V4 Pro and GPT-5.6 Terra took the top tier on the palace-intrigue novel; the fantasy champion dropped to sixth on the second book. Full ranking tables, a checklist of three long-run failure modes, and how to actually use the results.

Short answer: there is no all-genre champion. We had nine LLMs continue two novels for 20 consecutive rounds each — 360 rounds total, ranked by double-blind review. DeepSeek V4 Flash won the fantasy epic. DeepSeek V4 Pro and GPT-5.6 Terra took the top tier on the palace-intrigue novel. The fantasy champion fell to sixth place on the second book. Pick the model for the book you’re reading, not from a single leaderboard.
This page is the conclusions layer. Every placement and number below comes from experiments we’ve already published in full: the six-model fantasy benchmark (120 generations) and the follow-up runs documented in our long-run failure study. The protocol in one sentence: same opening passage, same prompt, each round’s output appended back into the context before the next round, 20 rounds per chain; outputs anonymized under two random letter mappings and ranked by two reviewers unaware of each other. Read those posts for transcripts and reviewer notes.
Which model best continues a fast-paced fantasy novel?
Six models continued an 8.9-million-character Chinese fantasy epic. DeepSeek V4 Flash won the blind review — first place in nearly every review window, with the note “the only system that reads like a chapter of the original”. It is also one of the cheapest models in the field. Expensive did not mean faithful.
| Rank | Model | One-line profile |
|---|---|---|
| 1 | DeepSeek V4 Flash | The only output that reads like the original author: vocative dialogue, standalone onomatopoeia lines, fast plot advancement |
| 2 | DeepSeek V4 Pro | Sharpest spoken dialogue in the field, but consistently over-describes — "like a finer-penned author writing the same story" |
| 3 | GLM 5.2 | Real content-level fit (tactical exchanges, one-word replies), but half-width quotation marks throughout break the surface |
| 4-5 | Claude Opus 4.8 | Deep-V trajectory: heavy literary flourish early, strong middle, punctuation corruption late — most volatile of the six |
| 4-5 | Qwen 3.7 Max | A consistently polished sensory stream — smells and textures the original never writes; consistently competent, consistently unlike |
| 6 | Gemini 3.1 Pro | Disqualified: rewrote the same opening for all 20 rounds (a prompt-wording misread; fixed, see below) |
Two footnotes. Gemini’s last place was not a capability problem: our context labels said earlier AI continuations were “for plot continuity reference”, and Gemini read “reference” as “not canon”, restarting from the original text’s ending every round. Rewording the label brought regressions to zero, and that fix already ships inside Foreverse’s continuation pipeline. Second, this table covers six models; GPT-5.6 Terra, Grok 4.5 and Kimi K2.6 ran the same fantasy protocol later, and their behavior is recorded in the failure section below. For price context, DeepSeek’s official rate card lists V4 Flash at $0.14 per million input tokens.
Which model is steadiest on ornate, dialogue-driven historical fiction?
Nine models continued Empresses in the Palace — a palace-intrigue classic written in first-person limited perspective and a formal, ornate register — from the same plot crisis. Two systems formed the top tier: DeepSeek V4 Pro (range 1-2), the only one that reliably reproduced the book’s two-layer structure of courteous surface dialogue decoded by the narrator’s inner voice, and GPT-5.6 Terra (1-3), which ran the full 20 rounds with zero formatting, honorific, or continuity incidents.
| Rank | Model | One-line profile |
|---|---|---|
| 1-2 | DeepSeek V4 Pro | Only system to reliably reproduce the "polite words, decoded subtext" double layer; wrote the single best scene in the whole sample |
| 1-3 | GPT-5.6 Terra | Zero-incident run; its evidence-chain plotting (stitch patterns, aged spices, baited traps) closest to the original author |
| 2-3 | GLM 5.2 | Purest limited-perspective observation, but the half-width-quote habit carried over unchanged from the fantasy run |
| 4-5 | Claude Opus 4.8 | Sharpest verbal sparring in the field, but late-run punctuation corruption and drifting names and ranks |
| 4-6 | Gemini 3.1 Pro | Back to normal mid-table after the wording fix; still the most literary-flourished voice of the nine |
| 5-7 | Kimi K2.6 | The only system that improved as it went: opened in a shouty webnovel register, settled into palace decorum |
| 6-7 | DeepSeek V4 Flash | The fantasy champion, down to mid-low: plain-spoken instincts fight the ornate register |
| 8 | Qwen 3.7 Max | Both reviewers independently ranked it eighth; the polished sensory stream reads even more foreign here |
| 9 | Grok 4.5 | Both reviewers ranked it last: verbatim repetition late in the run, and twenty rounds of story time never left one afternoon |
Read the two tables against each other and the reversal is systematic. V4 Flash won fantasy on plain, fast, clipped prose; the same instincts turn into a liability against this author’s ornate sentences. Its sibling V4 Pro flips the other way — dinged in fantasy for writing “like a finer-penned author”, it wins here because a finer pen is exactly what this book is written with. Same trait, opposite sign, different genre. GLM 5.2 is the only model in the top three of both tables, and its half-width quotation marks survived both books unchanged.
Update, July 21: Kimi K3 and Qwen 3.8 are in — does the board change?
Two trillion-parameter-class models shipped three days apart in mid-July, and we ran both through the same protocol as paired double-blind duels. The short version: neither top tier moves, but the mid-table is now stale — the new Kimi deserves a promotion, and the new Qwen only cashed in half of its upgrade.
Kimi K3 (released July 16) against its predecessor K2.6 was a rout: 10:0 on the fantasy epic, 11:1 on the palace novel, with K2.6’s metaphor pile-ups and verbatim self-copying gone. On structural metrics its fantasy proximity now lands between the two DeepSeek tiers — so K2.6’s “5–7” placement above no longer describes K3 — at the cost of half-width quotes on roughly 85% of dialogue and reasoning latency you cannot switch off.Qwen 3.8 Max Preview (July 19) against Qwen 3.7 Max was half an upgrade: 11:1 on the palace novel, where the old “polished-sensory-stream” tic genuinely receded, but only 7:5 on fantasy, opening at 6:6 — not a generational gap.
Put the two newcomers in the same ring and the duel ends 3:20 — K3 takes both genres, and Qwen 3.8 still doesn’t clear this generation’s top-tier bar. So the practical advice stands: DeepSeek V4 Flash for fantasy, V4 Pro and GPT-5.6 Terra for ornate historical fiction; if you want to try something new, K3 is worth the queue, and Qwen 3.8 can wait for its stable release. One more caveat: continuation rankings do not transfer to companion roleplay. We tested the memory axis separately over a 400-turn roleplay corpus, and neither a version bump nor a vendor switch cures bare-context amnesia — that is a different scoreboard entirely.
Before you commit: three long-run failure modes to check for
Every failure we observed across 360 rounds fits one of three patterns. All three are cumulative — invisible in a one-shot demo, unmistakable after enough consecutive rounds — which makes them a better shopping checklist than any single placement.
The restart loop. The model ignores everything generated so far and restarts from the original text’s ending, every round. Gemini 3.1 Pro did this for 20 out of 20 fantasy rounds. The root cause was one ambiguous sentence in our prompt; after rewording, zero restarts in 28 rounds across both books.
The mid-run freeze. The model advances, then locks onto its own recent output. Grok 4.5 froze in both genres: one fantasy stretch repeats the same two paragraphs three times word for word, and on the palace novel both reviewers independently wrote “plot rewind” — across twenty rounds the story clock never left the afternoon of the inciting incident. Its first-person density there was half the original’s (9 uses of “I” per thousand characters in the source, 5 in Grok’s output). Kimi K2.6 showed a milder version, copying whole passages from its own earlier rounds.
The ending rewind. The strangest one, and it hit the best prose mimic in the field. GPT-5.6 Terra’s early and middle fantasy windows were, by both reviewers’ judgment, the closest thing to the original author anyone produced. Then in rounds 18 through 20 it rewound the plot back to the anchor point and replayed chapter one, reusing lines from its own round-2 output. Surface mimicry and long-range coherence are separate capabilities; on the palace novel Terra never rewound and took a zero-incident top-tier run.
How do you actually use this ranking?
Pick per book, switch per scene, and don’t lock a whole novel to one model. The failures above build up over rounds; switching models mid-book, or regenerating one segment, resets the loop before it locks in. That is a more realistic plan than hunting for a perfect model up front.
Foreverse’s reader is built around exactly that: continuation can switch models per paragraph, with 62 preset BYOK providers plus any OpenAI- or Anthropic-compatible endpoint as a custom entry. Battle chapters on the fantasy champion, court scenes on the palace champion — inside the same book. Signup grants 5,000 credits, roughly 260 continuation segments on the official channel; both DeepSeek tiers are available there directly, while models outside the official catalog, like GPT-5.6 Terra or GLM, plug in with your own key. The full palace-run write-up is currently published in Chinese, with reviewer quotes per model.
Method and limits
The full protocol: each chain starts from a fixed anchor (a battle scene at the 55% mark of the fantasy epic; the plot crisis where contraband is discovered in the palace novel), runs 20 consecutive rounds, and feeds every round’s output back into a fixed 16k-token context window. Double-blind means outputs were anonymized under random letters with two different mappings and ranked by two reviewers unaware of each other; only mapping-consistent placements are reported as settled, and disagreements appear as ranges — that’s why the palace table shows 4-6 and 5-7. We used human judges rather than an LLM judge for a reason: our judge-reliability experiment showed LLM reviewers confidently misclassifying human prose.
Limits, stated plainly: every run shared one prompt framework; we tested two genres, both Chinese; and model versions will age the table. The rank ranges are not hedging — they record real reviewer disagreement. This page is updated as we re-run the benchmark; the date at the top is the truth source.
FAQ
Which free or cheap model holds up for continuation?
We didn't benchmark free tiers, so no ranking there. The closest data point: DeepSeek V4 Flash won the fantasy run outright while sitting in one of the lowest price tiers of the field — $0.14 per million input tokens at the official rate. Foreverse grants 5,000 credits on signup, roughly 260 continuation segments on the official channel, which is enough to try the top of both tables against your own book before paying anyone.
Models update constantly — is this ranking still valid?
Placements are tied to the versions we tested in July 2026: DeepSeek V4 Flash and Pro, Claude Opus 4.8, Gemini 3.1 Pro, Qwen 3.7 Max, GLM 5.2, GPT-5.6 Terra, Grok 4.5, Kimi K2.6. Kimi K3 and Qwen 3.8 Max Preview, both released in mid-to-late July, have since been retested head-to-head under the same protocol — see the update section in the article. New versions can reshuffle names, and this page gets updated when we re-run — check the date at the top. What outlasts any version bump is the method: pick the model per genre, and use a tool that lets you switch mid-book.
Should English and Chinese novels use different models?
Both benchmark books are Chinese (a fantasy epic and a palace-intrigue classic), so we have no English-novel ranking and won't invent one. What the data does support: switching genre within one language was enough to knock each champion to mid-table, so genre matters before language does. For an English book, replicate the protocol at small scale — same opening passage, a few candidate models, several rounds each, then read the outputs without knowing which is which.
How does the double-blind review actually work?
Each model's output is anonymized under random letters, with two different random mappings, given to two reviewers who don't know each other exists. Only placements that agree across both mappings count as settled; anything the reviewers flag as low-confidence or disagree on is reported as a range. In the fantasy run, the top three and last place matched exactly across both mappings; on the palace novel, the top and bottom tiers matched.
Questions or ideas? Join our Discord →