Grok 4.5 shipped ten days ago. For fiction, we already had a file on it
Every Grok 4.5 review benchmarks coding. Nobody answers the question people actually type into search: is it any good for stories? As it happens, Grok 4.5 was one of nine models in our long-run continuation benchmark — 20 consecutive rounds on each of two novels, ranked by double-blind review. The file shows a twice-reproduced repetition loop, a dead-last palace-intrigue placement from both reviewers, a first-person density at half the original's — and the scenarios where it is genuinely fine.

Grok 4.5 went into broad release on July 8, 2026, and the coverage since has been a wall of coding numbers. Fair enough — xAI’s own launch post introduces it as a model built to excel at “coding, agentic tasks, and knowledge work”, trained alongside a code editor, priced to run agent loops. Stories are not what it’s being sold for. Meanwhile “grok 4.5 creative writing” is a real thing people type into search bars, and what comes back is mostly air.
We don’t have to improvise an answer, because Grok 4.5 was one of nine models in the two-genre novel-continuation benchmark we published earlier this month. Twenty consecutive rounds per book, every round’s output appended back into a fixed 16k-token context, results anonymized under two random letter mappings and ranked by two reviewers who didn’t know each other existed. Forty rounds of Grok 4.5 writing fiction, judged blind. This post is not a hot take assembled for release week. It’s a records pull.
Where these numbers come from
The two test books are an 8.9M-character Chinese fantasy epic and the palace-intrigue classic Empresses in the Palace — deliberately opposite registers, plain-spoken versus ornate. The full ranking tables live in the nine-model results page, and the taxonomy of long-run failures in the failure-modes study; the complete palace run with per-model reviewer quotes is currently published in Chinese. One cohort note for honesty: the published fantasy ranking table covers the original six models, and Grok 4.5 ran that same fantasy protocol in the follow-up wave — so for fantasy we report its recorded behavior rather than a table placement.
Exhibit one: two paragraphs, three times, word for word
In the fantasy run, Grok 4.5 froze in the middle windows — one stretch repeats the same two paragraphs three times word for word — then partially recovered. That is the “mid-run freeze” pattern from our failure study, and Grok is its signature case: under the same prompt and protocol, most models never did this, and Grok did it in both genres. Kimi K2.6 showed a milder version, copying whole passages from its own earlier rounds; Grok’s loops were harder and came back.
Exhibit two: twenty rounds that never left one afternoon
On the palace novel the freeze was worse. Both blind reviewers independently wrote “plot rewind” in their notes, and across twenty rounds the story clock never left the afternoon of the inciting incident. Final placement: ninth of nine, from both reviewers, under both anonymization mappings.
The file holds one more measurement that matters for anyone writing first-person fiction. The source novel uses “I” about 9 times per thousand characters — the density anchor of its first-person limited narration. Grok’s output: 5. Large stretches drifted into third-person omniscient, which is the one discipline break that ornate first-person fiction cannot absorb.
Why no coding review will catch this
Because the failure is cumulative. A one-shot generation — the shape of nearly every demo and most writing benchmarks — looks fine; the loop needs rounds of the model’s own output feeding back into context before it locks in. That feedback regime is precisely what a reader app does when you keep pressing continue, and it is what coding evals never simulate. Our failure study calls this the autoregressive fixed point: once the window fills with its own text, imitating itself becomes easier than advancing.
The honest page: what Grok 4.5 is actually fine at
This file is not a takedown. Three entries on the positive side. If you already hold an xAI key for development work, short-burst fiction — a brainstorm, a dialogue rewrite — costs you nothing extra and sits below the round count where the loop builds. Its comfort zone is modern, colloquial registers; the place it fell hardest was ornate period prose, so keep the match in mind. And one practical note from our own provider catalog: among the 62 preset BYOK providers in Foreverse, the xAI entry is one of the few with all five modality slots lit — text, vision, image generation, video, voice — per the 2026-07-15 catalog snapshot, so one key covers illustration and narration experiments too.
Specs, checked against xAI’s model page on 2026-07-18: $2 per million input tokens and $6 per million output below 200k prompt tokens, doubling once a request reaches 200k; a 500,000-token context window; text and image input.
The timestamp on this file
Dates matter for model files, so here are ours. Broad release: July 8, 2026. Our benchmark posts shipped July 16, and the runs behind them landed in mid-July — this is the shipped model, tested within days of release, not a preview build. It still ages: endpoints get updated quietly, and every placement above is bound to the version we called. xAI’s next flagship, Grok 5, was still in training as of mid-July 2026 after its Q1 and Q2 windows both slipped, per public reporting. When either model changes, we re-run and update this page; the date at the top is the truth source.
So should you hand it a novel?
Not a whole one. The realistic use is routing: short scenes and modern registers are fine territory, and the moment you see a sentence come back verbatim, switch models — that breaks the loop before it locks in. In Foreverse’s continuation flow the model is a per-segment choice, not a book-level commitment, which is exactly the shape this file argues for. Grok 4.5 can sit in the rotation. It just shouldn’t be the one holding the pen at chapter forty.
FAQ
Is Grok 4.5 good for creative writing?
In short bursts, yes — brainstorms, dialogue rewrites, modern colloquial registers. Over long runs, no: in our 20-round continuation benchmark it locked into verbatim repetition loops in both books we tested. On the palace-intrigue novel, both blind reviewers ranked it last of nine, with twenty rounds of story time never leaving one afternoon. The failure is cumulative, so a quick demo will look fine.
Should I use Grok for roleplay?
Long roleplay is exactly the regime where its documented failure lives: dozens of consecutive turns, each output feeding back into context. In our data Grok 4.5 hit repetition loops in both a fantasy epic and a palace-intrigue novel under that setup. If you run it anyway, keep sessions short, watch for sentences repeating verbatim, and switch models the moment the story clock stops advancing. Short in-character exchanges in modern registers are a much safer fit than long ornate-period campaigns.
Why does Grok keep repeating itself when writing long stories?
It's the classic autoregressive fixed point: as the context fills with the model's own text, imitating itself becomes easier than advancing the plot. Most models resist it better — under the same prompt and protocol, Grok 4.5 was the only system in our field that froze in both test genres. One fantasy stretch repeats the same two paragraphs three times word for word. Prompt tweaks don't cure it; switching models or regenerating the looping segment breaks the cycle.
Did you test the released Grok 4.5 or an older checkpoint?
The released model. Broad release was July 8, 2026; the runs behind our published benchmark posts landed in mid-July 2026, within days of it. Placements are bound to the version we called — xAI iterates, and a future update could change the picture. We re-run the benchmark as models change, and the date at the top of this page is the truth source.
Questions or ideas? Join our Discord →