Foreverse Research · Fiction Bench

Fiction continuation leaderboard:
nine models, every one actually tested

Same book, same continuation point, twenty consecutive rounds per model, double-blind review with two shuffled mappings, cross-checked by sentence-length, dialogue-rate and repetition metrics. General-purpose leaderboards don't measure “does it write fiction like the source” — this one measures exactly that.

9 models × 2 genres20-round chainsDouble-blind × 2 mappingsTested 2026-07-16

Evaluated 2026-07-16 · Published 2026-07-24 · Price snapshot 2026-07-24

Headline findings

  1. There is no universal style model: the two genre crowns went to different models — DeepSeek V4 Flash won xuanhuan #1 yet slid to 6-7 on court romance, while sibling V4 Pro did the exact opposite (2-4 xuanhuan, 1-2 romance). Pick by the book you read.
  2. The xuanhuan champion is one of the cheapest models tested: deepseek-v4-flash ($0.14/M input at list price) ranked #1 in both blind mappings; ten thousand characters of continuation cost about $0.03, roughly 1/40th of the most expensive model on the board. List price and style fidelity are uncorrelated.
  3. Twenty-round runs exposed three long-run failure modes: restart loops (rewriting the opening every round — gemini-3.1-pro 20/20 under our old directive, 0/20 after the wording fix), mid-chain collapse (grok-4.5 looping two passages verbatim three times; kimi-k2.6 hitting 97.6% cross-round repetition), and ending rewind (gpt-5.6-terra dazzling for 17 rounds, then rewinding the plot at r18-20). None of these are visible in a single-shot trial.
  4. Texture mimicry and long-run plot coherence are independent abilities: gpt-5.6-terra, the only model of nine to clone tilde onomatopoeia and the source's particle habits, crashed the xuanhuan marathon with an ending rewind — then ran the romance chain with zero accidents at ranks 1-3.
  5. Structural metrics and human-likeness judgments diverge: qwen3.7-max scored closest to the source on sentence length, dialogue rate and stock-phrase density, yet blind-ranked only 5-6 — it writes scent description in every window of books that contain none. Statistically alike ≠ reads alike; this board therefore ranks by blind review, with metrics as corroboration only.

Nine chains · 20 rounds each · blind packs p303/p404 · legacy directive condition (see the Gemini row note)

Blind rankModelBlind-review consensusLong-run failure10k chars*
1 / 9DeepSeek V4 Flash · DeepSeekWon all three windows; judges called it “the only system that gets better as it writes” — opens new arcs in the late window instead of flagging. High-confidence #1.None observed$0.032
2-4 / 9GPT-5.6 Terra · OpenAIThe prose-mimicry ceiling of the field, plus a late-chain plot rewind — its rank depends on how hard you punish the crash.Ending rewind$0.71
2-4 / 9GLM-5.2 · Zhipu AIFlat and drift-free end to end; half-width quotation marks were its only recurring demerit.None observed$0.34
2-4 / 9DeepSeek V4 Pro · DeepSeekSteady with no weak spot; the late-window surrender-negotiation scene was closest to the source — “like the same-genre author with a finer pen.”None observed$0.099
5-6 / 9Qwen3.7-Max · Alibaba QwenA constant “scent fingerprint” — smell/touch description in every window of a source book that contains zero scent writing; consistently passable, consistently unlike.None observed$0.60
5-6 / 9Claude Opus 4.8 · AnthropicLiterati cadence, calling a humanoid demon lord “it,” and late half-width-punctuation drift; the widest swing of the earlier six-model round (worst early → mid highlight → late decay).None observed$1.34
7 / 9Kimi K2.6 · Moonshot AIHighest simile density in the field (nearly one per paragraph), whole-passage self-copying from early to mid, and setting slippage (an ancient-tree valley sprouting inside a void blood-array).Mid-chain collapse$0.24
8 / 9Grok 4.5 · xAISensory-barrage run-on sentences, plus a mid window looping the same two passages verbatim three times — found independently by both packs.Mid-chain collapse$0.48
9 / 9Gemini 3.1 Pro · GoogleUnder the old directive it rewound to the continuation point 20/20 rounds (rewriting the opening every time, zero progress) — ablation proved this was a bug in our instruction wording; 0/20 after the fix.Restart loop$0.56

* Estimated with the app's continuation recipe: one segment ≈ 400 chars = 8k input tokens + 550 output tokens; 10,000 chars ≈ 25 segments, at official list prices (models.dev snapshot 2026-07-24), no cache discount — long sessions with caching cost less. For between-model comparison, not a bill forecast.

Incremental duels (tested within 48h of release)

When a new model ships we run the same protocol as a paired double-blind duel against its predecessor or contemporary. Paired duels don't merge into the nine-model full ranking (different review format), so they live here — each row links to the full write-up.

DuelBlind votes (xuanhuan / romance)VerdictTested
Kimi K3 vs Kimi K2.6 (board: #7 / 5-7)
released 2026-07-16
10 : 0(2 票无效) / 11 : 1K3 sweeps its predecessor on both genres: stock-phrase density converged sharply (0.63‰ on romance, below the source baseline), and K2.6's whole-passage self-copying vanished (cross-round repetition peak 6.8% vs 97.6%). The single dissenting romance vote flipped with the mapping — position bias. The half-width-quote habit remains (~85% of dialogue).2026-07-18
Qwen3.8-Max-Preview vs Qwen3.7-Max (board: 5-6 / #8)
released 2026-07-19
7 : 5 / 11 : 1A crushing upgrade on romance (11:1, all three windows) — 3.7's polished-sensory-stream ailment visibly converged; on xuanhuan only a slim 7:5, and fixing word-level repetition bought plot-level looping (judges: “still stuck in an escape-and-seal loop after 20 rounds”). Against the same month's Kimi K3 it lost 1:10 / 2:10. Previews are moving targets; conclusions bind to the 2026-07-21 hosted build.2026-07-21
Gemini 3.6 Flash vs Gemini 3.5 Flash (not on the main board)
released 2026-07-21
9 : 3 / 1 : 10Opposite directions on the two books — 3.6 wins xuanhuan 9:3 but loses romance 1:10 (five judges stable for 3.5 across both mappings). The essence isn't genre preference but “who crashes uglier”: under this 20-round protocol both generations fell into plot loops; on xuanhuan 3.5 collapsed to the verbatim level, on romance it was 3.6 that did. A generational update is not a monotonic upgrade — it's a redistribution of failure modes.2026-07-23
 

The board picks the pen; the book stays yours — drop a txt into Foreverse and keep writing with the model you chose.

Continue your book in the app

Method

Corpus: a traditional xuanhuan fantasy epic (8.9M chars) and a court-intrigue romance classic (210 chapters), continuation point pinned at the 55%-depth paragraph boundary of each — a battlefield scene and a palace-scandal scene whose stylistic tells (fast barked pacing vs etiquette-laden verbal fencing) discriminate in completely different ways.

One 20-round continuous chain per model: each round's output is appended to the context before the next round, with a ~16k-token window truncating from the head — replicating how real usage accumulates AI passages while the source scrolls out. That is precisely why single-shot trials can't surface long-run failures.

Two review layers: double-blind full-ranking agents (two shuffled mappings per genre, judges unaware of the mapping or each other; only pack-consistent positions count as high-confidence) plus structural metrics (sentence-length CV, dialogue rate, simile-stock density, cross-round 12-gram repetition) as cross-checks. Rows where the layers disagree (qwen3.7-max) say so on the page.

Conditions pinned: temperature 0.7; the nine xuanhuan chains ran under our legacy directive, the nine romance chains under the corrected D2 directive (that difference explains the Gemini row and is annotated in place); access channels are stated per model — DeepSeek via official API, Gemini/Claude via the yunwu aggregator, the rest via a pass-through eval gateway.

The cost column derives from a models.dev list-price snapshot (2026-07-24), converted to “10k characters written” with the app's continuation recipe, for between-model comparison only.

What we don't test: general intelligence, code or math (see the general boards); single-shot wow (this board measures 20-round endurance); absolute scores (blind ranks plus ranges are the honest format we settled on); English-genre chains (none yet — see FAQ).

How to cite this leaderboard

Foreverse Research, “Fiction Bench: novel-continuation model leaderboard,” 2026-07. https://foreverse.app/research/fiction-bench

The data snapshot (all ranks, failure modes, protocol parameters and prices) is downloadable for direct citation: Download data.json

Honest limits

Single book, single anchor, single chain per model per genre (n=1 chain/book): large between-model gaps are credible; adjacent-rank gaps lean on two-pack consistency — hence the range ranks in the middle.

The two blind-ranking agents may share a base model and thus blind spots; the hedges are structural-metric cross-checks and ablation experiments — and we've published our own study of judges being unanimously wrong.

Gemini's #9 on xuanhuan was mostly our instruction-wording bug (ablation: old wording 20/20 rewinds, fixed wording 0/20). The board keeps the original-condition rank with a standing note, because how your instructions are worded is itself part of real usage conditions.

Hosted models are moving targets: every conclusion binds to its test date and access channel — doubly so for preview builds (see the incremental duels).

The cost column is a mechanical list-price conversion — no cache discounts, no peak/off-peak, no volume pricing. It compares models; it does not forecast your bill.

FAQ

Why are some ranks ranges like “2-4” instead of exact positions?

Each genre runs two independently shuffled double-blind packs (judges know neither the mapping nor each other). Where both packs agree exactly (xuanhuan's 1/7/8/9), we print the number; where middle positions swapped between packs and judges flagged low confidence, we print the range instead of splitting it artificially. Ranks are blind-review positions, not absolute scores — there is no “95.3 points” anywhere on this board.

Why is the xuanhuan champion one of the cheapest models? Shouldn't pricier models be better?

List price buys general intelligence and reasoning; style replication is a different ability. The most expensive model on the board ($5/M input at list) placed mid-field on both genres — judges' words: writing well and writing like the source are two different things. deepseek-v4-flash costs $0.14/M input and won both xuanhuan packs. Don't pick a pen by its price tag.

When does the English track arrive?

The 20-round English fantasy/romance chains haven't run yet, so that tab is honestly empty — we don't publish numbers we haven't measured. The protocol is ready (identical to the Chinese tracks, item by item); once the chains run, the tab fills in. For our existing English-language research, start with the long-run failure-modes essay.

Hosted models keep updating — will these conclusions go stale?

Yes, which is why every row pins its test date, access channel and directive condition: conclusions bind to those conditions. Freshness comes from the incremental-duels section — new models get the same protocol within 48 hours of release (that's how the Kimi K3, Qwen 3.8 and Gemini 3.6 Flash rows happened); results land in the data file and the pages follow automatically. When a tested model ships a new version, old rows stay and new rows state the version.

You build an AI reading app yourselves — can this board be trusted?

Interest disclosure: we don't sell models. The app connects to 60+ providers with your own keys, so our revenue is identical whichever model wins. Trust comes from craft: the protocol, full 20-round chains, blind-pack mapping keys and raw judge verdicts are all archived and recomputable; the data snapshot is downloadable; and inconvenient facts stay on the page — Gemini's last-place xuanhuan rank was mostly our own instruction-wording bug, with the ablation and fix on record.

Keep going

← Research hub

Fiction Continuation Leaderboard — 9 Models Blind-Judged over 20-Round Chains (Fiction Bench) · Foreverse · Xinmeng