Nine models, one identical brief, one question: whose prose survives a blind test?

Nine LLMs got the same writing brief — imitate a human web-fiction voice — and their output was mixed into real novel excerpts for blind judging. The ranking looks nothing like continuation benchmarks: DeepSeek v4-flash, our continuation champion, came in last; Opus 5 produced the only passage the gatekeeper genuinely believed was human. The winning ingredient across the top three: cheap, genuine flaws — real typos and real verbal habits that can't be faked on instruction.

Warm-paper editorial illustration: a row of fountain pens lined up before a judge's bench, the red-ribboned pen trailing a wobbly smudged ink line while the others draw perfectly uniform strokes

Start with the upset. DeepSeek v4-flash holds two first places in our continuation benchmark. Hand it a different exam — impersonate a human web-fiction writer from scratch — and it finishes dead last out of nine, caught by the gatekeeper at confidence 85, the highest of the field.

Same model, opposite ends of the table. The reason is simple: continuing and impersonating are different skills. Continuation offers an existing voice to lean on. Impersonation removes the crutch, and whatever the model actually is shows through. This post is about the impersonation exam: nine LLMs, whose Chinese web-fiction prose is hardest to unmask?

The exam

Every writer got a byte-identical brief: an anchor excerpt from a real novel (learn this voice), a full anti-AI-flavor discipline sheet, and three narrative voice types to write one short chapter each. Writers were isolated from each other. Output was mixed with real passages from the same books into blind pools, and a gatekeeper judged each passage human-or-machine with confidence and cited evidence.

Why is the gatekeeper a same-generation Claude thinking model and not GPT or Gemini? Because we tested the judges first. Ten mainstream models judged “reads human” in the inverted direction — praising disciplined machine prose, flagging real published novels as AI. The only machine judge whose rationale matched a human reader's went 100% catch with zero false positives across seven rounds. It presides here, with a second judge cross-checking, and neither rules on its own family's output alone.

RankWriterResultOne-line verdict
1Opus 5Only passage the chief judge believed (conf. 72)Zero-payoff timestamps, unevenly clustered verbal tics — diary-grade rough edges
2DeepSeek v4-proOne passage slipped throughA genuine typo, cited as evidence of a lazy human
3Gemini 3.1 ProOne slipped through — with a plagiarism incident on recordHomophone slips read as hand-typed serial fiction
4–5GLM-5.2 / GLM-5One fooled the second judge onlyColloquial seasoning sprinkled too evenly
6Sonnet 5All caughtSentence shapes most human of the field; overall execution gave it away
7Fable 5All caughtExecutes so well even its planted errors come out symmetrical
8–9Gemini 3.6 Flash / DeepSeek v4-flashAll caught, highest confidenceManual-tone exposition, self-contradicting lore

Dossiers

Opus 5 — won by writing useless things

The only machine passage the chief judge sincerely called human. Post-mortem: it wrote timestamps that paid off nothing, and it leaned on one turning-word again and again in one stretch, then forgot the habit elsewhere. That unevenness is what real diaries look like, and it can't be faked to order. In a separate 100k-character long-form test it was also the only model with a rationing instinct — its richest prose went to one character's dialogue while the narration stayed plain. Its recorded weakness: chapters swell about 40% by the end of a long run.

DeepSeek v4-pro — one real typo, worth its weight

Its surviving passage rode on a genuine typo the judge cited as “a typo the author couldn't be bothered to fix.” Not craft — the model's Chinese output naturally carries a little grit, and the grit happens to sit exactly where machines find it hardest to fake. Note the gap to its sibling: pro nearly passes, flash finishes last. The tier difference is much larger on this exam than you'd guess.

Gemini 3.1 Pro — grit earns points, plagiarism burns them

Homophone slips bought it one surviving passage (“hand-typed serial fiction grit — machines almost never do this”). It also owns the experiment's only serious incident: under one calling harness it copied 30+ characters verbatim from the anchor passage in its own brief, book title included. The same model over a different API, same brief, stayed clean. If you use the Gemini family for imitation work, an overlap check against your references has to be a hard gate, not a spot check.

The GLM family — evenness is the tell

GLM-5.2 fooled the second judge once; the chief judge held. The verdict pattern: spoken seasoning sprinkled too evenly, three of the same particle in one paragraph, every paragraph exactly as playful as the last. Real colloquialism arrives in gusts.

Sonnet 5 — statistically most human, holistically exposed

On shape metrics — sentence length, fragment cadence — it was the only writer whose narration landed inside the human band across the board. Caught anyway, on jargon misuse and a compulsion to end every paragraph with a summarizing bow. It also carries a 100k-character rap sheet from our long-form test: one pet cliché written 101 times, density degrading 4.8× in the back half, chapter length collapsing from ~2,600 to ~1,100 characters. Right shape, wrong stamina.

Fable 5 — beaten by its own execution

Our own workhorse writer, and the reason this experiment exists. Its problem in one line: it executes every instruction perfectly, and perfection is the fingerprint. Ask it to plant a typo and it will politely have the narrator correct the typo two lines later, symmetrically. It never slipped a single passage past the gatekeeper in over ten rounds.

Grok 4.5 — a sprinter with a long-distance problem

Unremarkably caught on the short exam. The long-form dossier is the memorable part: its first 100k-character attempt collapsed into synopsis — thirty chapters totalling 27k characters. Re-run with a hard per-chapter word floor, the collapse converted into padding: one “gathering his thoughts” template pasted 18 times, and one passage repeated verbatim three times. A judge's line stings: “the first two chapters are the best prose I've read this round — I'd sign the first two chapters. I would not sign the book.” It fails at self-management, not craft. For opening lines and short pieces, no need to ban it.

DeepSeek v4-flash — the other face of a champion

Manual-tone exposition, lore that contradicts itself, caught at the field's highest confidence. Its reversal between the continuation table and this one is the most useful reminder in the whole exercise: when a leaderboard says “best model for fiction,” ask what task was measured.

The pattern: you can't buy the winning ingredient

Every surviving passage won on flaws the model actually has — real typos, real homophone slips, real verbal habits. All nine briefs were identical, so the spread is entirely the models' own texture. We verified the converse too: instruct a flawless model to write cheap errors, and the errors land neatly in the safest spots. The judge's word for it was distressing — as in furniture. Authentic imperfection doesn't transfer by instruction. Model choice is part of the writing process, and it comes first.

Picking by task

Your taskPickAvoidNote
Prose that must read humanOpus 5, then DeepSeek v4-proDeepSeek v4-flash, Gemini 3.6 FlashModel choice beats prompt stacking
Continuing an existing storyDeepSeek familyDifferent exam — see the continuation benchmark
100k-character long-formOpus 5Sonnet 5, Grok 4.5Watch the back half: cliché density and chapter length decay grow with distance
Short piecesOpus 5 or Grok 4.5Grok's failure mode needs distance to appear
Imitation with GeminiUsable with a hard gate15-character sliding-window overlap check against references, reject on hit

For long runs, three cheap early-warning meters that agreed with human verdicts in our tests: cliché density above ~1.5 per 10k characters; a per-word counter on the model's pet phrases (a healthy run shows a dozen hits per 100k characters, a runaway shows 149); and final-five-chapter average length below 70% of the first five.

Where this ranking ends

Three honest limits. Each writer wrote one piece per voice, so catch confidences are upper bounds — nine writers on one theme handed the judge clustering clues that a single passage in the wild wouldn't. The GPT family sat out entirely on account of account quotas, so this table says nothing about it. And the exam is Chinese web fiction, short form; change the domain and seats will move. We use it as a directional signal, not a podium.

The day the experiment closed, one line in our production config changed: the default writer for pass-as-human jobs went from Fable 5 to Opus 5. Not because Fable writes badly — because it writes too well.

FAQ

Which model should I pick for writing fiction?

Depends on the task. For prose that has to read human-written, Opus 5 first, DeepSeek v4-pro second. For continuing an existing story in its own voice, the DeepSeek family actually tops our separate continuation benchmark. For 100k-character long-form, avoid Sonnet 5 (cliché density degraded 4.8× in the back half of our test) and Grok 4.5 (long-range process collapse) — though Grok remains fine for short pieces; it fails at self-management, not craft.

Why would a strong continuation model rank last at passing as human?

Continuing and impersonating are different skills. Continuation gives the model an existing voice to lean on, and DeepSeek v4-flash leans best of anyone. Impersonation from scratch removes the crutch, and its expository manual-tone and self-contradicting lore were caught at the highest confidence of the field. Before trusting any ‘best model for fiction’ ranking, check which task it measured.

What are “cheap genuine flaws” and why do they win?

The top three passages won on natural errors produced by real capability gaps: DeepSeek v4-pro wrote a genuine typo the judge cited as ‘a typo the author couldn't be bothered to fix’; Opus 5 reused one turning-word like a personal verbal habit, unevenly. These can't be faked. When we instructed a flawless model to write cheap errors on purpose, the errors landed in perfectly safe spots and were called out as distressing — artificial aging — on sight.

How do I stop an LLM from copying my reference text verbatim?

Run a sliding-window overlap check (we use 15 consecutive characters) between the output and every reference you fed it, and reject on any hit. Not theoretical: Gemini 3.1 Pro once copied 30+ characters verbatim from the anchor passage in its brief, book title included, in one harness while the same model over a different API stayed clean. With the Gemini family, make the check a hard gate.

Questions or ideas? Join our Discord →

Nine LLMs Tried to Pass as Human Webnovel Writers. One Nearly Did. · Foreverse · Xinmeng