Foreverse Research · Fiction Bench
How good is Grok 4.5 at writing fiction?
Not the pick for Chinese fiction under this protocol: #8 on xuanhuan, a unanimous #9 on romance. Its ailment recurs across genres — mid-chain verbatim looping (the same two passages, three times), a romance chain whose plot stalled on the day of the incident for all 20 rounds, and first-person density at half the source's, so the limited POV never holds.
Xuanhuan fantasy
8 / 9
Both packs: p303 第8 · p404 第8
Sensory-barrage run-on sentences, plus a mid window looping the same two passages verbatim three times — found independently by both packs.
Court romance
9 / 9
Both packs: p505 第9 · p606 第9
A unanimous last place; looping and plot-rewind recurred across genres (verbatim late repeats, 20 rounds stalled on the day of the incident) plus first-person density at 5‰ — half the source's — a collapse of POV discipline.
What the 20-round chains actually showed
A unanimous #8 on xuanhuan: sensory-barrage run-ons, and a mid window that looped the same two passages verbatim three times — flagged independently by judges in both mappings. The late window recovered with new text, which is why mid-chain collapse is a fixed-point stall of autoregression on its own output, not a permanent freeze.
A unanimous #9 on romance: the looping-plus-rewind ailment recurred across genres — verbatim late repeats, 20 rounds stalled on the incident day — plus one hard POV number: first-person density of 5‰, half the source's. A first-person limited narrative that gradually loses its “I.”
Bottom of both genres with the same failure shape means this isn't genre mismatch; it's a structural weakness in long-run generation. It may do fine on short single-shot tasks — this board measures 20-round sustained continuation.
Long-run failure mode
Mid-chain collapse
Mid-chain collapse: the xuanhuan mid window looped the same two passages verbatim three times (late recovered with new text); the romance chain repeated verbatim late and rewound its plot, stalling all 20 rounds on the day of the incident — the ailment recurs across genres.
Structural fingerprint
Sensory-barrage run-on sentences; on romance its first-person density of 5‰ was half the source's — limited-POV discipline collapsed.
Cross-genre profile: Bottom of both genres; the looping/rewind ailment recurs across genres.
Test-condition disclosure (hosted models are moving targets)
Model under test: grok-4.5 (released 2026-07-08)
Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.48 list price: in $2/M · out $6/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Grok 4.5.
Continue your book with itHow to cite
Foreverse Research, “How good is Grok 4.5 at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/grok-4-5