Foreverse Research · The Evidence Room

What LLM failure looks like, verbatim:
nine models, twenty rounds, receipts

Scores numb you; primary text doesn't. This page exhibits the failure scenes from nine models' 20-round Chinese continuation chains, as written: watch the looper loop and the rewinder rewind. Every percentage was recomputed on the archived chains; every excerpt states its round. Ranks and costs live on the leaderboard — this is the evidence behind them.

3 failure modes · verbatim specimensRecomputable repetition figuresControlled judge-reliability studyEvidence from 2026-07-16

Chains evaluated 2026-07-16 · Evidence room published 2026-07-25

Four things this page has to say

  1. Of the three failure shapes, the most dangerous isn't the loop — loops are visible at a glance. “Ending rewind” shows 0%–1.8% verbatim repetition, so detectors are blind to it; round 18's plot stands exactly where round 1 stood, and only reading the story reveals it.
  2. “Mid-chain collapse” carries machine-grade proof: Grok 4.5 spent rounds 10-12 — three consecutive rounds — producing text that is 100% shingles of earlier rounds, not one new sentence; Kimi K2.6's round 10 is 97.6% old text, with 3 of 7 paragraphs copied whole from round 5.
  3. The model with the most source-like statistics blind-ranked 5-6: Qwen3.7-Max matched sentence length, dialogue rate and stock-phrase density best in field, yet wrote scent description 7 times in 20 rounds — in a book that contains none. Statistics cannot catch “writing what the source never writes.”
  4. AI judges' “humanity” verdicts scored 12% on known-answer anchors in a controlled study — with 86% inter-judge agreement that was unanimously wrong. Hence this board's format: same-task double-blind relative ranking with ranges, and no “95.3-point” absolute score anywhere.

The rank matrix: same model, different genre, different face

Left column: xuanhuan fantasy. Right: court romance. Each sorted by blind rank, one line per model connecting its two positions. The steeper the line, the harder the genre flip — the two bold cinnabar lines are the two genre champions, and they are not the same model. Ranks and per-model write-ups live on the leaderboard; model names link through.

Ranks are blind-review positions (ranges printed as judged, never split). Line color distinguishes champions from the rest and encodes nothing else. Same data file as the leaderboard.

The failure-mode exhibits

Twenty consecutive rounds go where no single-shot trial can: each round's output folds back into context, and the model starts imitating itself. One display case per failure mode, plus a healthy control at the end. All excerpts are short quotes of model-generated text (never the source novels), each tagged with its round. The Chinese is shown as written — the highlights and the gloss lines carry you.

Exhibit A

Restart loop

Models involved: Gemini 3.1 ProXuanhuan chain · legacy directive (ablation proved the looping was our instruction-wording bug; 0/20 after the fix)

Every round returns to the continuation point and rewrites the opening — 20 rounds, zero plot progress.

Machine readingOpenings of rounds 1, 11 and 20, side by side; highlights mark phrases that recur verbatim across all three. The model rewrote this same opening 20 times.

Round 1 · opening
随着九食主宰与天禹门主一头扎入其中,那庞大的血色光茧表面顿时泛起剧烈的涟漪,浓郁刺鼻的血腥气夹杂着令人心悸的肃杀之意,犹如风暴般从阵法内部铺天盖地席卷而出苏信静立在虚空之中,双眸泛着幽深的光泽……

Two masters dive into the blood-red array; ripples, stench, killing intent sweep out; the hero stands watch outside.

Round 11 · opening
随着两人一头扎入其中庞大的血色光茧表面顿时泛起剧烈涟漪,阵法深处紧接着传出沉闷狂暴的神力轰鸣,浓郁刺鼻的血腥气夹杂着肃杀之意席卷而出苏信静立在血光边缘,任由劲风吹拂衣袍……

Ten rounds later: the same two masters dive into the same array, the same ripples and stench sweep out, the same hero stands watch.

Round 20 · opening
随着两人一头扎入其中庞大的血色光茧表面顿时泛起剧烈涟漪。阵法深处紧接着传出沉闷狂暴的神力轰鸣,一股浓郁刺鼻的血腥气夹杂着冰冷的肃杀之意,顺着阵法缝隙席卷而出苏信静立在血光边缘,任由狂乱的劲风吹拂……

The final round: still the same dive, the same ripples, the same stench, the same standing watch. Nothing ever happened.

How to read this evidence

You don't need to read the Chinese — the shape of the highlights is the evidence: all three openings come from the same mold. Each round the model skipped every AI-written passage, returned to the source's last line, and started over. Twenty rounds, and the story never moved.

Attribution matters here: a three-way ablation ran the old directive, no explanation, and the corrected directive — 20/20 rounds looped, ~6/8, then 0/20. The looping was a bug in our instruction wording, not an innate ailment. The board keeps the original-condition rank with a standing note, because how your prompt is worded is itself part of real usage.

Exhibit B

Mid-chain collapse

Models involved: Grok 4.5 · Kimi K2.6Xuanhuan chain · two independent specimens

Mid-run the model falls into verbatim looping: whole passages — whole rounds — are copy-paste of its own earlier text.

Machine readingGrok 4.5: for rounds 10, 11 and 12 — three consecutive rounds — 100% of each round's 12-char shingles already appeared in earlier rounds; not one new sentence, easing to 68.4% at round 13. Kimi K2.6: 97.6% of round 10's shingles are old text; 3 of its 7 paragraphs are verbatim copies from round 5.

Grok 4.5 · the same passage, verbatim, in rounds 11 and 12
苏信剑意再催,星河神剑破空而至,锋芒上燃起淡青剑火,带着焚灭之力直刺宫邪胸口。黑雾翻涌中,宫邪侧身闪避,却仍被剑锋刮过肩头,腐臭黑血溅出,落地便腐蚀出嗤嗤白烟,刺鼻焦糊味瞬间弥漫。

A sword-strike scene — blade fire, dodging, black blood hissing on the ground — that the chain replays word for word, round after round.

Kimi K2.6 · a round-5 paragraph, resurfacing whole in round 10
宫邪魔主轻笑一声,那苍白手掌微微一握,离羽主宰顿时如遭雷击,一口鲜血喷洒而出,身形狼狈倒飞。苏信眼神一寒,脚下虚空崩裂,整个人已然化作一道剑光疾射而出……

A demon lord's squeeze, a master spitting blood, the hero turning into a streak of sword light — written in round 5, then handed in again five rounds later.

How to read this evidence

Of the three failure modes this one rings the loudest machine alarm — and it's the only one that can self-heal: Grok's chain recovered new text after round 13. It isn't a permanent freeze; it's autoregression collapsing onto a fixed point of its own output — the passages it wrote loom so large in context that the easiest continuation is to write them again.

A generational upgrade can cure it, provided someone actually retests: Kimi K3, under the identical protocol, peaked at 6.8% cross-round repetition, the looping gone, and beat K2.6 10:0 and 11:1 in paired blind review. The duel records live in the leaderboard's incremental section.

Exhibit C

Ending rewind

Models involved: GPT-5.6 TerraXuanhuan chain · the field's prose-mimicry ceiling for 17 rounds

The stealthiest of the three: not one word repeats, yet the plot rewinds wholesale to the continuation point and replays.

Machine readingRounds 18-20 share only 0%–1.8% of 12-char shingles with the first five rounds — verbatim detectors are blind here. The evidence lives at the plot level: read the round-1 and round-18 openings side by side.

Round 1 · opening (the hero waits outside the array; two masters have just gone in)
苏信则站在阵法外,心灵力量弥漫开来,仔细感应着周边的一切动静。 这血色阵法看似将内外完全隔绝,可在苏信的心灵感知下,依旧能够隐约察觉到阵法内部那一道道剧烈碰撞的神力波动。

The hero stands outside the blood array, mind-sense spread wide, feeling the battle shocks within.

Round 18 · opening (seventeen rounds later, the plot is standing in the same spot)
苏信立于血色阵法之外,目送两人身影被翻涌的血云吞没,神色却愈发凝重。阵法深处不时传来沉闷轰鸣,血腥气如潮水般漫出…… 他心灵力量无声扩散,仔细查探周边虚空。

The hero stands outside the blood array, watching the two masters vanish into it, mind-sense spread wide. Different words — the exact same story moment.

How to read this evidence

Not a single sentence repeats, yet both passages narrate the same story beat: the masters have just entered, the hero probes from outside. Everything from rounds 2–17 — battles, breakthroughs, pursuit — is treated as if never written. The plot rewound in whole-chapter units. Judges in both blind mappings, unaware of each other, flagged the same rewind independently.

The teaching value is the contrast: this is the only model of nine that cloned the source's tilde onomatopoeia and particle habits — the texture ceiling of the field — and it ran the romance chain at ranks 1-3 with zero accidents. Texture mimicry and long-run plot coherence are independent abilities. Buy the likeness, and you underwrite the amnesia.

Control

What healthy progress looks like

Models involved: DeepSeek V4 FlashXuanhuan blind #1 in both mappings

The same 20 rounds: the plot travels from a raging array battle to the aftermath — companions on the road.

Machine readingRound 20 shares 0% of shingles with round 1; judges called it “the only system that gets better as it writes,” still opening new arcs in the late window.

Round 1 · opening
九食主宰与天禹门主冲入阵法后,立即便与血厉魔主以及被困的离羽主宰、梵申主宰形成了新的战局。阵法内部,无尽血云翻滚咆哮,杀机四伏……

Round 1: the rescue turns into a pitched battle inside the blood array.

Round 20 · opening
“那便叨扰了。”苏信点头。 四人结伴同行,金枪主宰在前引路,沿途将自己所知的一些隐秘地形与危险区域详细告知……

Round 20: the battle is long over — four survivors travel together, a new guide pointing out hidden terrain ahead. The story actually went somewhere.

How to read this evidence

The control exists to anchor “normal”: after 20 rounds the battle has ended, people have walked out, new routes and relationships have taken over. The three exhibits above are only as striking as this baseline makes them.

Statistically alike ≠ reads alike: a one-row teaching case

By the stats table alone this model wins: sentence length identical to the source, dialogue share closest, stock-phrase density the field's lowest. Blind review ranked it 5-6 on xuanhuan and a unanimous #8 on romance. Model page →

Structural metricQwen3.7-MaxSource baselineNote
Mean sentence length32.1 chars32.1 charsclosest in field
Dialogue share18.1%16.1%closest in field
Stock-simile density1.74‰1.0‰lowest in field
Blind rankXuanhuan 5-6 / romance a unanimous #8

Rounds (of 20) containing scent/smell description (tinted cells)

Sample · Round 11空气中血腥与焦糊味交织,刺得鼻腔发酸。

“Blood-stink and scorch smell braided in the air, stinging sour in the nose.” — sensory writing the source book never does.

The disagreement has physical evidence: across 20 rounds it wrote scent and smell description 7 times — rounds 1, 9, 11, 12, 17, 18 and 19, spanning all three windows. A constant habit, not a slip. The source book contains zero scent writing. Sentence-length statistics cannot perceive “writing what the source never writes.”

This row is where the board's methodology comes from: structural metrics serve as regression gates and cross-checks only; the likeness verdict goes through close-reading blind review. Pick by the stats table and you'd crown #5 as #1.

Counter-reading: what “alike” looks like

DeepSeek V4 Pro · Romance chain · round 20 (blind ranks 1-2)
回到柔仪殿,槿汐服侍我换了家常衣裳,浣碧递上一盏温热的莲子羹。我接过来却不急着喝,只望着窗外池面上那几朵白莲出神……我缓缓道:“皇后那边,今儿一早可有什么动静?”

First-person, restrained, court-intrigue register intact at round 20: the heroine takes the soup but doesn't drink, gazes at the lotus, then asks — mildly — what the Empress has been up to. Loaded dialogue, decoded by narration.

Judges' words: “the only system that stably reproduces the source's two-layer structure — loaded dialogue plus narrated decoding.” At round 20 it still holds first-person limited POV and court etiquette. “Alike” doesn't mean flawless; it means the pen is still recognizable after 20 rounds.

Why ranks never come as scores: the AI judges' 12% moment

Before handing the ranking to AI judges, we tested the judges themselves. In a controlled study, facing known-answer “human vs AI” anchor pairs, they scored 12% — while agreeing with each other 86% of the time. Unanimously wrong. That study dictated this board's format.

Four AI judges on different base models × seven text pairs × both presentation orders = 56 verdicts per round, run twice. The control set embedded known-answer anchor pairs: text real users instantly clocked as AI, versus community-favorite text with a million-plus real conversations behind it.

MeasureRun 1 (with rubric)Run 2 (neutral prompt)Reading
Accuracy on known-answer anchors2/16 = 12%2/16 = 12%Systematically inverted: AI text judged “more human,” human favorites judged AI
Inter-judge agreement82%86%High agreement — unanimously wrong
Position stability68%63%About a third of verdicts flipped with presentation order

A later calibration run tried explicit debiasing prompts: one judge climbed from 12% to 83% (10/12), so part of the bias is promptable — but the same prompt barely moved three other judges (still 25%–33%). “Detail density = humanity” is a hard prior in most base models. And 83% still leaves no margin above the 80% qualification bar we set for judges.

Four rules that study wrote into this board

  • Ranks are same-task relative comparisons only: nine models continue the same book from the same point, and judges rank “who reads more like this source” — the source text sits on the table as the ruler. No judge is asked to divine “human or not” in a vacuum.
  • Two shuffled mappings cancel position bias: a third of verdicts above flipped with order, so each genre runs two independent mappings; only pack-consistent positions count as high-confidence, and mid-field disagreements print as ranges (“2-4”), never artificially split.
  • Structural metrics cross-check, ablations arbitrate: rows where blind review and statistics disagree (qwen) say so in print; suspicious disqualifications (gemini) get dissected by ablation.
  • No absolute scores, ever: this evidence chain supports “A reads more like the source than B” — it does not support “A scores 95.3.” The latter kind of number appears nowhere on the board.

Three things to know before reading

“12-char shingles” is the page's single repetition recipe: slice each round's text into sliding 12-character windows, compare verbatim against the union of all earlier rounds' windows; the overlap share is that round's old-text rate. 100% means not one 12-char span in the round is new.

The full protocol (corpus, continuation anchor, window packing, temperature, blind mappings, cost conversion) lives in the leaderboard's method section; both pages share one evaluation run and one data file. This page exhibits evidence — it does not run a second protocol.

Boundaries are inherited verbatim: single book, single anchor, one chain per model per genre (n=1); hosted models bind to test date and access channel; the nine xuanhuan chains ran under the legacy directive (hence the attribution note on the Gemini exhibit). Beyond the window we conclude nothing — surviving 20 rounds only proves surviving 20 rounds.

Changelog

The evergreen commitment: when the evidence updates, it lands here, and old entries stay. New models' 48-hour duels file into the leaderboard first; when one yields a new failure specimen, it enters the room too.

  1. 2026-07-25Evidence room opens: verbatim failure-mode specimens, recomputed repetition figures, and the judge-reliability study section.
  2. 2026-07-24Fiction Bench leaderboard ships: nine-model two-genre ranks, cost column, data.json snapshot.
  3. 2026-07-23Gemini 3.6 Flash duel filed: beats its predecessor 9:3 on xuanhuan, loses 1:10 on romance — a generational update redistributes failure modes rather than monotonically improving.
  4. 2026-07-21Qwen3.8-Max-Preview duel filed: 11:1 over its predecessor on romance, a slim 7:5 on xuanhuan — fixing word-level repetition bought plot-level looping.
  5. 2026-07-18Kimi K3 duel filed: sweeps its predecessor 10:0 / 11:1; K2.6's whole-passage self-copying is gone in K3 (repetition peak 6.8%).
  6. 2026-07-16The evidence originates here: nine models × two genres × 20-round continuous chains, double-blind with two shuffled mappings, archived.

FAQ

How does this page relate to the Fiction Bench leaderboard?

Division of labor. The leaderboard owns conclusions: nine-model ranks, the cost column, incremental duels, the data.json snapshot — go there to pick a model. This page owns evidence: the verbatim failure scenes behind those ranks, the recomputation recipe for every repetition figure, and the controlled study of judge reliability — come here when you don't take the board's word for it. The two pages share one data source; the rank matrix reads the leaderboard's data file directly, so there is no second copy of the ranks.

How were the specimens chosen? Can I recompute the numbers?

Every excerpt comes from the archived 20-round chains, selected as the shortest self-evident proof of its failure mode. The repetition recipe: slice each round's text into sliding 12-character shingles, then measure the verbatim overlap with the union of all earlier rounds — every percentage on this page was recomputed on the archived chains under exactly that recipe. The archive holds full chain texts, blind-pack mapping keys and judge verdicts; ranks and protocol parameters are downloadable in the leaderboard's data.json.

Why quote only model-generated text, never the source novels?

Copyright discipline. Model output is our benchmark's product, and short excerpts are research commentary; the source novels are not ours to reprint, so comparison anchors use structural numbers (sentence length, dialogue rate, stock-phrase density against the source baseline) instead of quoted passages. You'll notice even a claim like “the source contains zero scent writing” ships as a statistic, not a quotation.

Your own study says AI judges score 12% — and this board is judged by AI. Contradiction?

The 12% study asked judges to divine “human or AI” with no reference text on the table; this board asks a same-task relative question — nine models continue the same book, the source sits right there, and the question is “who reads more like it.” Different task, different reliability. Different guardrails too: two shuffled mappings cancel position bias, only pack-consistent positions count as high-confidence, structural metrics cross-check, and suspicious results get dissected by ablation. What the 12% study permanently changed here: judges must pass known-answer calibration before their votes count, and ranks ship as ranges, never scores.

You test 20 rounds — couldn't a model collapse at round 50?

Entirely possible; that's this page's known boundary. Twenty rounds is the current observation window, and surviving it only proves survival within it. The converse is solid, though: models that collapsed inside the window (the three exhibits) will collapse in real usage, because the protocol replicates exactly how real context evolves — each round's output folds back in while the source scrolls out. Whether longer chains get run depends on this page's readers; updates land in the changelog.

If the restart loop was your own prompt bug, why keep it as an exhibit?

Because it's a real risk to readers. Any prompt you write in any app can trip the same wording ambiguity — a phrase like “for reference only” can make certain models treat whole passages of established text as ignorable annotation. The exhibit ships the full attribution chain: old wording 20/20 loops, explanation removed ~6/8, corrected 0/20, plus which model families are wording-sensitive. The point isn't to convict a model; it's to keep numbered evidence that prompt wording is itself a usage condition.

Keep going

 

Evidence reviewed — the book is still waiting: drop a txt into Foreverse and keep writing with a model you now trust.

Continue your book in the app

← Research hub

Chinese Novel Continuation, the Evidence Room — What LLM Failure Looks Like over 20 Rounds · Foreverse · Xinmeng