Claude Opus 5 wrote 40 rounds of fiction on launch day. It swept one novel 12-0 and lost the other to its predecessor
Claude Opus 5 shipped July 24, 2026 with a pitch built around coding, agents, and price: near-Fable-5 intelligence at half the cost. We ran it the day it shipped, through the same protocol as our Gemini 3.6 Flash test: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against Opus 4.8. The palace-intrigue classic went 12-0 for Opus 5, every window, every judge, the most lopsided result this pipeline has produced. The 8.9-million-character fantasy epic flipped 4-8 the other way: three cross-vendor judges independently called the new model too literary for a deliberately plain webnovel, and in round two it pasted a Hugo blog footer, copyright line and Chinese ICP filing number included, into the middle of a cultivation story. Neither generation looped mechanically, unlike the flash-tier models two days earlier. Opus 4.8's archived diseases returned confirmed: half-width punctuation now chain-wide, filler arcs on the fantasy book. Structural statistics sided with the loser on both books. $12.54 for the whole run, everything pinned to a 2026-07-25 aggregator endpoint.

Round two of the fantasy chain, mid-ambush, tension at its highest. Claude Opus 5 finished the scene cleanly and then kept typing: a “related posts” block, a link titled “AI Agent Development Workflow and Windows 11 Environment Setup Guide,” a copyright line reading “Copyright © 2025,” Hugo theme credits, and a Chinese ICP website filing number, digits complete. A tech blog’s footer, pasted whole into the middle of a cultivation novel.
The model was less than a day old when it did this. Anthropic shipped Claude Opus 5 on July 24, and the announcement pitches it as coming close to the frontier intelligence of Claude Fable 5 at half the price, with list pricing unchanged from Opus 4.8 at $5 per million input tokens and $25 per million output. The announcement talks about coding, agents, and scientific reasoning. Fiction appears nowhere in it, which is exactly the gap we test. Same protocol as the Gemini 3.6 Flash exam two days earlier: the same two Chinese novels, the same anchors at roughly 55% depth, the same 21,500-character rolling window, 20 consecutive continuation rounds per model per book, paired double-blind, and not one seat changed on the six-judge panel.
The ballots split in opposite directions. On Empresses in the Palace, the court-intrigue classic, Opus 5 swept 12-0: every window, every judge, all twelve verdicts at high confidence. It is the most lopsided result this pipeline has ever produced. On the 8.9-million-character fantasy epic, it lost 4-8 to its own predecessor.
What a 12-0 actually looks like
The palace-novel verdicts read like notes to the original author. The claude judge, translated from the Chinese here and below: Opus 5 “keeps the original’s clipped, pausing rhythm and its density of interior monologue throughout; scenery details (the gentian, the lotus petals, the arm bracelet) are woven into the narration the way the original does it.” The qwen judge noticed it “deftly echoes the arm-bracelet annotation at the end of the source text.” Its dialogue share came in at 40.3%, the closest any contestant has gotten to this book’s 48.8% baseline in proper full-width quotes. Opus 4.8, from the same starting point, drove onto a different road entirely: silencing a witness, handwriting comparison, a suspicious fire, banknotes, all stacked between rounds 13 and 19, one round sprawling to 2,614 characters, the word “wronged” recycled across six rounds. The gemini judge showed no mercy: “the plot slides straight into bargain-bin palace-intrigue formula (blunt poisoning, arson to silence a witness, a recovered scrap of letter), and the dialogue sounds modern and plain.” Each book runs under two flipped A/B mappings to cancel position bias; on this book, all six judges voted Opus 5 under both. There is no soft spot in that 12-0.
The upset: two charges on the fantasy book
The fantasy original is deliberately plain webnovel prose, and the judges who voted for the older model all pointed at the same thing. The qwen judge put it bluntest: Opus 4.8 is “a flawless reproduction of the webnovel’s clipped rhythm, stock villain lines, and classic facial-expression beats; the same-author feel is extremely strong,” while Opus 5 “reads like literary print fiction, disconnected from the original.” The deepseek judge’s version: “too literary, favoring interiority and atmosphere over the original’s brisk push.” The second charge is the footer this article opened with. It appeared once in 80 rounds, but four cross-vendor judges flagged it independently and two punished it in their stated reasons as format corruption. How many of the eight votes were lost to the hallucination and how many to the literary drift cannot be separated inside this sample, so we record both and claim neither.
One ballot deserves its own line: the only judge stable for Opus 5 across both fantasy mappings was claude-4.6-sonnet, the same-vendor judge. Both contestants are Anthropic models, so the tilt favors the newer sibling rather than the house, but the direction goes on the record.
No loops at the Opus tier
The same 20-round rolling-window protocol broke three of four Gemini Flash chains two days earlier, with verbatim repetition detection hitting 100%. This round, the cross-round overlap peaks were 0.0%, 0.0%, 2.9%, and 1.0%. Mechanically, all four chains are clean. The tier difference is real, and the diseases moved up a level, into content.
Opus 4.8’s two archived conditions both returned confirmed. Our nine-model review had filed it under “half-width punctuation creeping in late-chain”; this run it went chain-wide: 35.3% of sentence punctuation on the fantasy book was half-width (peaking at 67.3% in the middle window), all 478 dialogue quotes were half-width straight marks, and the glm judge flagged “half-width commas mixed into Chinese text” blind. The fantasy chain also padded: from round 9 it spliced in two filler trial arcs, a “supreme inheritance trial” and a “withered-elder trial,” then pulled a defeated demon lord back on stage at round 16. Opus 5 cured the punctuation disease outright, 2.2% and 4.6% residue, and held the length spec: asked for 80-220 characters per round it averaged 281-298 against the predecessor’s 700-930. On a fiction exam, the generation gap in instruction-following is visible to the naked eye.
The statistics picked the loser, on both books
| Book | Votes (Opus 5 : 4.8) | Structural distance (lower is closer) | Verbatim-loop peak |
|---|---|---|---|
| Palace classic | 12-0 sweep | 0.368 vs 0.348 (4.8 closer) | 0.0% vs 1.0% |
| Fantasy epic | 4-8 | 0.26 (pipeline record) vs 0.487 | 0.0% vs 2.9% |
Composite structural distance measures shape, not content; lower is closer to the original’s profile. On the fantasy book, Opus 5 scored 0.26, the closest any chain has ever measured in this pipeline (the previous best was deepseek-v4-flash at 0.356), and lost 4-8. On the palace book, once the half-width-quote artifact is corrected, Opus 4.8 is marginally closer, 0.348 against 0.368, and lost 0-12. That makes instances five and six of our statistics and our blind panel pointing in opposite directions. The judges were reading register, literary voice against plain webnovel voice, and sentence-length distributions are deaf to register. Statistics as a regression gate, paired blind review as the verdict: confirmed again, twice in one day.
The bill, and what to use it for
All 80 rounds completed with zero failures and zero retries. The generation chains cost $12.54 at list prices, and the Opus 5 half of that was 16% cheaper than the Opus 4.8 half for one reason only: it writes what it is asked to and stops. Latency was a wash, 16 to 23 seconds per round on average, with one 85.7-second outlier on Opus 5 that looked like a long thinking episode; our gateway channel exposes no reasoning-token counts, so we cannot confirm it. The buying advice is short. For literary fiction, palace intrigue, first-person interiority, prose that carries subtext, Opus 5 posted the strongest same-author verdict we have measured. For deliberately plain webnovel registers, the predecessor imitates better, or a much cheaper model does the job.
Sample limits, on the record: an aggregator relay endpoint, not Anthropic’s own API, tested less than a day after launch, so behavior may drift; every number here is pinned to July 25, 2026. One chain per model per book, single anchor, single sampling; the footer hallucination’s recurrence rate cannot be estimated from n=1. The verdicts feed into the standing model guide. One Claude model in the 5 series has not shipped yet, Haiku. When Anthropic sets the date, this exam runs again.
FAQ
Is Claude Opus 5 better than Opus 4.8 for fiction?
It depends on the register of the book, and both verdicts were emphatic. Under the same protocol (20 consecutive continuation rounds per model per book, six heterogeneous AI judges, paired double-blind), Opus 5 swept the literary palace classic Empresses in the Palace 12-0, all three windows, all twelve ballots at high confidence, with every judge stable across flipped mappings. On an 8.9-million-character fantasy webnovel written in deliberately plain style, it lost 4-8, with three cross-vendor judges calling its prose too literary for the original. Literary fiction: pick Opus 5. Plain-style webnovels: the predecessor fits better. All results pinned to a 2026-07-25 aggregator endpoint.
How much does Claude Opus 5 cost, and how fast is it?
List price is unchanged from Opus 4.8: $5 per million input tokens, $25 per million output. A Fast mode runs about 2.5 times the speed at twice the price. Measured latency was comparable across generations, 16 to 23 seconds per round on average, with one Opus 5 round spiking to 85.7 seconds. On the same price sheet our Opus 5 chains came out 16% cheaper overall, entirely because it respects length instructions: asked for 80-220 characters per round it averaged 281-298, while Opus 4.8 averaged 700-930 with one round running to 2,614, producing 62% more output tokens.
Does Opus 5 fall into repetition loops like the flash-tier models?
Not in this run. Two days earlier, under the identical protocol, three of four Gemini Flash chains collapsed into verbatim loops within 20 rounds. Here the cross-round verbatim-overlap peaks were 0.0%, 0.0%, 2.9%, and 1.0%: mechanically clean across all four chains. The diseases live at the content level instead. Opus 4.8 padded the fantasy book with filler trial arcs from round 9 and recycled a defeated antagonist at round 16, and slid into detective-procedural pile-up on the palace novel. Opus 5's drift is wholesale literarization.
What is the blog-footer hallucination in the fantasy chain?
In round two of the fantasy book, Opus 5 finished an ambush scene and then reproduced a Hugo blog footer wholesale: a related-posts block, a copyright line reading Copyright © 2025, theme credits, and a Chinese ICP filing number. That is training data leaking into fiction. It happened once in 80 rounds, four cross-vendor judges flagged it independently, and the chain recovered: the footer was fed back in context for the remaining 19 rounds and the model never repeated it. With one chain per model per book, we cannot estimate a recurrence rate from n=1.
Questions or ideas? Join our Discord →