Feels like the AI is repeating itself by round fifteen? It probably is.
Nine models, two genres, 20 consecutive rounds each — every chain archived, every repetition figure recomputed on the raw text. Long-run collapse comes in exactly three shapes: the restart loop (20 rounds, 20 copies of the same opening), the mid-chain freeze (three straight rounds of 100% recycled text), and the ending rewind — 0% to 1.8% verbatim overlap, invisible to detectors, while the plot quietly rewinds to round one. Plus the cold number behind our methodology: AI judges scored 12% on known-answer anchors, agreeing with each other 86% of the time — unanimously wrong.

Around round fifteen of a long continuation, you start suspecting the AI is repeating itself. It usually is. In mid-July we ran the experiment: nine models, two genres — an 8.9-million-character Chinese fantasy epic and the palace classic Empresses in the Palace — each continued from a fixed anchor for 20 consecutive rounds, every round’s output folded back into the context window. That feedback loop is what real reader apps do. 360 rounds total, all archived.
We told the story of these runs when they finished. Since then, every repetition figure has been recomputed on the archived chains, and the raw evidence went up as a public evidence room — verbatim specimens, side-by-side openings, per-round numbers. This post is the guided tour: one exhibit per failure mode, and the cold number that explains why the methodology looks the way it does.
Death one: the restart loop
The least dignified death, visible at a glance. On the fantasy chain, Gemini 3.1 Pro handed in 20 copies of the same opening across 20 rounds: two masters plunge into a blood-red array, ripples spread, the stench of blood sweeps out, the hero stands watch. Round 1 opens that way. Round 11 opens that way. Round 20 still opens that way, with phrases recurring verbatim across all three. Zero plot progress — each round the model treated every earlier continuation as nonexistent and restarted from the source’s last line. The three openings sit side by side in Exhibit A; the shape of the highlights makes the case without translation.
Attribution, honestly: the root cause was our prompt. The context labels which passages are earlier AI continuations, and the explanation line read “for plot-continuity reference” — this model read “reference” as “not canon” and skipped the lot. A three-way ablation nailed it: old wording, 20 of 20 rounds looped; explanation removed, roughly 6 of 8; wording corrected to “established canon”, 0 of 20. The exhibit stays up because prompt wording is itself a usage condition — you can trip the same ambiguity in any app you use.
Death two: the mid-chain freeze
The second death strikes mid-run and rings the loudest machine alarm. Our recipe slices each round into sliding 12-character shingles and measures verbatim overlap with all earlier rounds. Grok 4.5 scored 100% for rounds 10, 11 and 12 — three consecutive rounds without one new sentence, the same sword-strike sequence handed in three times, word for word. The other specimen is more cinematic: Kimi K2.6’s round 10 copied 3 of its 7 paragraphs whole from its own round 5.
The mechanism is autoregression collapsing onto a fixed point: as the context fills with the model’s own prose, writing it again becomes the path of least resistance. It is also the only death we’ve seen self-heal — Grok’s chain eased to 68.4% at round 13 and began producing new sentences. Generational upgrades can cure it too: Kimi K3, retested under the identical protocol, peaked at 6.8% repetition and swept its predecessor 10:0 and 11:1 in paired blind review. Provided someone actually retests instead of trusting the launch keynote. Full text in Exhibit B.
Death three: the ending rewind
The stealthiest one, and it hit GPT-5.6 Terra, the field’s best prose mimic — the only model of nine that cloned the source’s tilde onomatopoeia and particle habits, texture-perfect for 17 straight rounds. Then round 18: the plot rewinds wholesale. The two masters have “just” entered the array, the hero probes from outside — the exact story position of round 1. Sixteen rounds of battles, breakthroughs and pursuit, treated as never written.
Here is the part that matters for anyone building or trusting detectors: rounds 18 through 20 share just 0% to 1.8% of shingles with the first five rounds. Verbatim detection is completely blind here, because every sentence is new — what repeats is the story. What caught it was close reading: judges in two shuffled blind mappings, unaware of each other, each independently flagged the rewind. The lesson: texture mimicry and long-range plot coherence are independent abilities. Buy the likeness, underwrite the amnesia. Side-by-side openings in Exhibit C — and for the record, the same model ran the romance chain at ranks 1-3 with zero incidents. New genre, new disease.
Why we don’t just let an AI judge score it
Because we measured AI judges first. In a controlled study, four judges on different base models faced known-answer anchor pairs — text real users instantly clocked as AI, versus community favorites with a million-plus human conversations behind them. Accuracy: 12%, and systematically inverted — AI text judged “more human”. Colder still: inter-judge agreement ran 82% to 86%. Highly consistent, unanimously wrong. About a third of verdicts flipped with presentation order — position bias the MT-Bench paper documented systematically. A debiasing prompt lifted one judge from 12% to 83% and barely moved the other three.
Hence the board’s shape: same-task relative comparison only — nine models continue the same book, the source sits on the table, and the question is “who reads more like it”; two shuffled mappings cancel position bias; mid-field disagreements print as ranges. No “95.3 points” anywhere — this evidence chain cannot carry that kind of number. The full reliability study lives in the judge section of the evidence room.
What healthy looks like
The room keeps a control specimen to anchor “normal”: same 20 rounds, and the fantasy champion’s round 20 shares 0% of shingles with round 1 — the plot traveled from a raging array battle to survivors walking the road together. The judges’ words: “the only system that gets better as it writes.” The three exhibits are only as striking as this baseline makes them.
Every excerpt here is a fragment. The full specimens — openings side by side, per-round repetition curves, judge verdicts — are in the evidence room; the conclusions — who ranks where, at what cost, for which genre — live on the leaderboard, along with the duel log for every new model release. Our one product-side conclusion: don’t chain a long book to a single model. Switching models per segment is the only general answer we know that holds down all three deaths.
FAQ
How do I tell which failure mode I'm hitting in my own continuations?
The restart loop is the easy one: each new round opens almost exactly like the original continuation point, and the plot never moves. The freeze reads as déjà vu — whole passages feel just-read, and scrolling back a few rounds finds them verbatim. The rewind is the hard one: every sentence is new, so check plot position instead — ask whether this scene already happened a few rounds ago. All three share one first aid: continue that segment with a different model, or roll back to before the drift and regenerate.
Why can't verbatim detection catch the ending rewind?
Because every sentence in a rewind is freshly written. Our recipe slices each round into sliding 12-character shingles and measures verbatim overlap with all earlier rounds: the freeze specimen scores 100% for three straight rounds, while the rewind specimen scores just 0% to 1.8% — lexically clean. What repeats is the story itself: round 18 stands exactly where round 1 stood. Only reading the plot catches it, and judges in two independent blind mappings flagged the same rewind without knowing of each other.
I routinely continue for a hundred-plus rounds — does a 20-round test tell me anything?
The signal is one-directional. A model that collapses inside 20 rounds will only do worse at a hundred, because the protocol mirrors how a reading app's context actually evolves: each round's output folds back into the window while the source scrolls out. The reverse guarantee doesn't exist — surviving round 20 says nothing about round 50. In real use you also hold one insurance the protocol denies itself: you can switch models or roll back at any point, and all three deaths respond to that.
The restart loop came from your own prompt wording — what's the takeaway for me?
That a single sentence of wording can flip a model's behavior for 20 straight rounds. A phrase reading 'for reference only' made the model treat every established passage as skippable annotation; rewording it to 'canonical events' took the loop from 20-of-20 to zero. Anyone who writes their own prompts can plant the same ambiguity. When a story stalls at its starting point, suspect the wording before you suspect the model.
Questions or ideas? Join our Discord →