Grok 4.5's case file, retried under the original protocol

Grok 4.5 holds the worst record on our nine-model continuation board: #8 on the fantasy epic, a unanimous #9 on the palace novel, convicted of mid-chain verbatim looping, a stalled story clock and collapsed first-person discipline. Grok 4.6 reached the xAI API on August 12, 2026; within roughly 24 hours we retried the case under the identical protocol — two Chinese novels, 20 consecutive continuation rounds each, paired double-blind. All three old charges dropped: 12:0 against the predecessor on fantasy, 12:0 on the palace novel, twelve mapping-stable verdicts all at high confidence — the strongest generational signal this pipeline has produced. Verbatim repetition peaked at 2.8% and 2.1%; first-person density recovered from half the source's rate to the champion's level. But the court filed one new charge: outline-itis. Against the reigning romance champion it lost 2:10, five judges independently citing anemic narrative density — about 150 characters per round against the champion's 475. The fantasy bout against the reigning champion went 8:4 raw, 4:0 among stable ballots; per our own rules, no crown changes hands. Metered cost: $2.10.

A case folder with a broken cinnabar seal spills a paper ribbon that coils in tight loops before running straight past an open pocket watch

The ugliest file on our nine-model continuation board belongs to Grok 4.5: #8 on the fantasy epic, a unanimous #9 on the palace novel, bottom of both genres. Three counts on the charge sheet: mid-chain it looped the same two passages verbatim, three times over; its palace chain spent twenty rounds unable to leave the afternoon of the inciting incident; and its first-person density ran at half the original author's, losing the narrator of a first-person book. That file has sat in the archive for almost a month.

On August 12 the Grok 4.6 API opened — same base model, post-training upgrade, per xAI. Grounds for a retrial. Within roughly 24 hours we reconvened: identical protocol, each old charge re-examined, plus the customary bout against the reigning champions. The verdict up front: all three old charges dropped, 12:0 and 12:0 against the predecessor — and one new charge filed, which cost it the romance bout 2:10.

Terms of the retrial

The protocol matches the K3, Opus 5 and Gemini 3.6 Flash rounds item for item: the same two books (an 8.9-million-character Chinese fantasy epic, and Empresses in the Palace), the same anchor at roughly 55% depth, the same 16k-token rolling window, 20 rounds per chain at temperature 0.7, the model as the only variable. The defendant ran on xAI's official API (api.x.ai/v1; the model ID is grok-4.6 with a dot — the hyphenated form 404s), tested 2026-08-13, no protocol deviations. The aggregators hadn't picked the model up yet, which conveniently forced us onto the cleanest possible channel.

Every opponent is archive footage, re-run zero times: Grok 4.5's two chains, fantasy champion DeepSeek V4 Flash's chain and romance champion V4 Pro's chain, all from the July 16 benchmark archive. Two condition gaps, disclosed in the open: the fantasy opponents ran under our older context-labeling wording (the palace-side chains all share the current wording, so those pairings are clean), and the opponents are snapshots rather than today's live weights — a caveat that acquired new weight the very same day; we'll get there at the end. Judges: six models across five vendors; deepseek-v4-pro recused itself as a contestant's stablemate and gemini-3.5-flash took the seat per precedent; no xAI model judges. All 48 verdicts returned clean.

Charge one: mid-chain verbatim looping — dropped

Old evidence first. Grok 4.5's fantasy rounds 9, 10 and 11 are an identical 271 characters — the same two passages, three times — re-verified byte for byte before this retrial. The palace chain's image swamp was also measured to the bottom: calamus 21 times across 19 rounds, “musk mingled with” 13 times. Six judges flagged both without knowing whose chains they were reading.

New evidence: Grok 4.6's cross-round 12-gram repetition peaks at 2.8% (fantasy, round 14) and 2.1% (palace, round 13) — far under the 30% alarm line, not one round flagged. A round-by-round human read finds linear plot on both books: the fantasy chain runs from blood-cocoon siege through breakout to regrouping without one backtrack. Charge dropped.

Charge two: the stalled story clock — dropped

Grok 4.5's palace chain never left the afternoon of the incident; both blind reviewers independently wrote “plot rewind” in the original benchmark. Grok 4.6 moves the clock: the accusation scene, house arrest, a demotion edict, aftermath maneuvering, a fresh plot hook planted for the next arc — twenty rounds that actually leave the incident day. gemini-3.1-pro's line from the predecessor bout reads like a closing statement: “B's pacing is composed, its sentences fall in varied lengths, and scenery like the brick seams rinsed with well water carries the original's cool, fine-grained spirit” — the quoted image exists verbatim in round 17. Charge dropped.

Charge three: first-person collapse — dropped

This one has a blunt number. The source writes “I” at 9.0 per thousand characters; Grok 4.5 managed 5.2, a first-person book slowly losing its narrator. Grok 4.6 reads 11.0 — level with the romance champion's archived chain, back in the source's range. Of the three convictions this is the cleanest repair; no commentary required.

The verdict: 12:0, and another 12:0

Against the predecessor: fantasy 12:0 (windows 12:0 / 12:0 / 12:0), palace 12:0 (windows 10:2 / 12:0 / 12:0), and all twelve mapping-stable verdicts at high confidence. The combined 24:0 is the most lopsided generational result this pipeline has produced — the previous record was Opus 5's single-book 12:0. xAI's “post-training upgrade” claim is, in the fiction-continuation domain, measurably true.

Two verified compliments from the fantasy bench: gemini-3.5-flash — “the original is a minimalist, action-heavy, description-light fast webnovel; A fits this plain-spoken, high-action-density register perfectly”; claude-4.6-sonnet — “short-sentence attack, breath-level detail riding the action rhythm ('the blood cocoon's surface shivered', 'rust-iron tang'), dialogue brief and restrained” — both quoted phrases exist verbatim in round 1.

The new charge: outline-itis

The romance bout against reigning champion V4 Pro (0716 snapshot) went 2:10, and it is a strong signal: five judges mapping-stable for the champion, glm-5.2 alone stable for Grok 4.6 (a taste split), zero contradictory ballots. The five wrote variations of one sentence. gemini-3.1-pro: “too few words, too fast — the long-form texture of the original is entirely lost.” gemini-3.5-flash: “severe outline-style shrinkage late in the chain, anemic narration, no layered detail.” kimi-k2.6: “degrades into outline prose late, losing the winding delicacy.”

The mechanical exhibit pins the number: about 150 characters per round from Grok 4.6, about 475 from the champion. The irony deserves its own line — Grok 4.6 is the only chain of the six on the table that held the 80-220-character length spec every single round; the champion runs 475, the predecessor 239-267, all out of spec. The same 150 characters that fantasy judges praised as punchy rhythm, romance judges read as poverty. Rule-keeping is a virtue in one genre and a charge in the other — for anyone wiring a continuation model into a reader, that finding is worth more than the score.

The championship bout: unresolved

Fantasy, against reigning champion V4 Flash (0716 snapshot): raw ballots 8:4 for Grok 4.6. But this bout drew the heaviest position-bias contamination in the series: four of six judges cast eight ballots that contradicted themselves under the mapping flip, all voided. The surviving stable component runs 4:0 for Grok 4.6 — claude and kimi, with no judge stable for the champion. Direction favorable; confidence insufficient. Stack on the old-directive condition gap in the opponent chain, and our own rules say: no crown changes hands. Logged as “same tier as the champion, stable ballots 4:0 in its favor.” The champion's chain took its own verified hits — judges flagged late-window padding collapsing into log-keeping prose, with the recurrence counts checked one by one.

Grok 4.6's residual record, also in the open: image-level reuse. “Rust-iron tang” nine times across nine rounds, “chill” twelve times; on the palace side a metallic-sweet taste six times in eight rounds and the well-water motif across five. The two-sidedness is real — the same “brick seams rinsed with well water” that one judge praised as the original's spirit, two others filed under repetitive imagery. Downgraded from verbatim looping to a fondness for certain images: that is what remains after the generational repair.

Release terms, and the bill

List price (xAI's model page, verified 2026-08-13): $2 per million input tokens, $6 per million output; cached input metered at $0.5 in our runs; a double-priced faster tier we didn't touch; 500K context; reasoning on by default. The bill: $2.0955 across both chains as metered round by round, about $2.10 with probes — roughly ¥15, the cheapest run in this series despite a list price 3-10× the domestic models' (K3's round: ¥12.21; Opus 5's: $12.54). The reason is the new charge itself: 150-character rounds, with reasoning metered separately. That reasoning runs 95.5-95.6% of output tokens, the highest share we've recorded (K3 held the record at 81-86%); visible prose is about 130 tokens per round, the thinking about 21 times the prose; latency 39-44 seconds per round, and inside a reader that wait is real.

Boundaries: one prompt framework; two books, both Chinese; one chain per book, single anchor; the fantasy opponents carry the old-directive condition; Grok 4.6 was about 24 hours old and hosted behavior may drift; the outline-itis verdict's sensitivity to the anchor point is untested, and n=1 stays n=1. One boundary gained weight the same day: the opponents are July 16 snapshots, not live weights — and on August 13 DeepSeek hot-swapped V4 Pro to its GA build under the same model ID, a build that lost 1:11 to the very snapshot Grok couldn't beat. The romance winner is now as unreachable as the loser's old case file. The same day handed the industry both halves of the lesson: one vendor's new version swept its predecessor 12:0 twice; the other's lost to its own preview with zero stable ballots.

The practical ruling: for fast-paced fantasy, Grok 4.6 belongs on your shortlist today; for dense period prose it is one tier of narrative supply short — check the buying guide for the current callable list. Whether to switch is not a question for us: same opening, three continuations each, read them shuffled and flip the cards. This file reopens the day Grok 4.7 ships.

FAQ

Is Grok 4.6 good for creative writing?

Two answers. Against its own predecessor it is a rout: paired double-blind under the same protocol, 12:0 on the fantasy epic and 12:0 on the palace novel, all twelve stable verdicts at high confidence, with Grok 4.5's verbatim loops, stalled plot clock and first-person collapse all cured. Against the reigning champions it splits: on fantasy it beat DeepSeek V4 Flash 8:4 raw with stable ballots 4:0 in its favor — but only two judges were mapping-stable and the opponent chain carries an old-directive condition, so we don't declare a new champion; on ornate court romance it lost 2:10 to DeepSeek V4 Pro's archived snapshot, diagnosed with outline-itis. For fast-paced fantasy it now belongs on your shortlist; for dense period prose it is one tier short.

Did Grok 4.6 fix Grok 4.5's repetition loops?

Yes, with machine readings to show it. The old conviction: Grok 4.5's fantasy chain looped the same two passages verbatim three times (rounds 9, 10 and 11 are an identical 271 characters), and its palace chain spent 20 rounds stuck on the day of the inciting incident. Grok 4.6 under the identical protocol: cross-round 12-gram repetition peaked at 2.8% (fantasy, round 14) and 2.1% (palace, round 13), far below our 30% alarm line, and a round-by-round read shows linear plot advancement on both books — the palace chain finally moves past the incident day. What remains is an image-level fondness for reuse ('rust-iron tang' nine times, a well-water motif across five rounds): downgraded from disease to habit.

Is Grok 4.6 better than DeepSeek for fiction?

Pick by book. Fast-paced fantasy: 8:4 raw and 4:0 stable against reigning champion V4 Flash puts them in the same tier — Flash is an order of magnitude cheaper, Grok 4.6 is far more disciplined about length. Ornate court romance: a clear 2:10 loss to V4 Pro's 0716 snapshot, and the gap is narrative density — about 150 characters per round against 475. One same-day complication worth knowing: DeepSeek hot-swapped V4 Pro to its 0813 GA build under the same model ID, and that GA build lost 1:11 to the very snapshot Grok couldn't beat. The romance winner is now unreachable too — full accounting in our DeepSeek GA write-up.

How much does Grok 4.6 cost for novel continuation?

List price, verified 2026-08-13: $2 per million input tokens, $6 per million output, cached input metered at $0.5 in our runs, a double-priced faster tier we didn't use, 500K context, reasoning on by default. Our bill: both 20-round chains cost $2.0955 as metered per round, about $2.10 with probes — the cheapest run in this series despite a list price 3-10× the domestic models', because it writes 150-character rounds and its reasoning tokens are metered separately (K3's round cost ¥12.21; the Opus 5 round, $12.54). Reasoning is 95.5-95.6% of output tokens — the highest share we've recorded — and each round takes 39-44 seconds.

Why did Grok 4.6 lose the romance duel?

Density, not machinery. Five judges were mapping-stable for the champion with zero contradictory ballots, and their reasons converge on one word: outline. 'Too few words, too fast, the long-form texture of the original is entirely lost.' 'Severe outline-style shrinkage late, anemic narration, no layered detail.' The numbers: about 150 characters per round versus the champion's 475. The irony is that Grok 4.6 is the only chain of the six on the table that obeyed the 80-220-character length spec every single round — and the same 150 characters read as punchy rhythm to the fantasy judges and as poverty to the romance judges. Compliance is a virtue in one genre and a charge in the other.

Questions or ideas? Join our Discord →