
Grok 4.5's case file, retried under the original protocol
Grok 4.5 holds the worst record on our nine-model continuation board: #8 on the fantasy epic, a unanimous #9 on the palace novel, convicted of mid-chain verbatim looping, a stalled story clock and collapsed first-person discipline. Grok 4.6 reached the xAI API on August 12, 2026; within roughly 24 hours we retried the case under the identical protocol — two Chinese novels, 20 consecutive continuation rounds each, paired double-blind. All three old charges dropped: 12:0 against the predecessor on fantasy, 12:0 on the palace novel, twelve mapping-stable verdicts all at high confidence — the strongest generational signal this pipeline has produced. Verbatim repetition peaked at 2.8% and 2.1%; first-person density recovered from half the source's rate to the champion's level. But the court filed one new charge: outline-itis. Against the reigning romance champion it lost 2:10, five judges independently citing anemic narrative density — about 150 characters per round against the champion's 475. The fantasy bout against the reigning champion went 8:4 raw, 4:0 among stable ballots; per our own rules, no crown changes hands. Metered cost: $2.10.
Read the post →






