Ten LLM judges passed our AI prose. A human reader caught it five times out of five.
A day-and-night adversarial experiment: mix LLM-written passages into excerpts from real web novels and see who can tell. Ten mainstream LLM judges scored a 0% catch rate — and called the real books AI. A human reader caught the machine every single round. Here are the six human tells, each with a machine sentence vs. a real-book sentence, plus five disguise tactics that all failed.

On July 25 our own lead read a batch of freshly written character cards and said, “these all reek of GPT.” Nobody argued. Everyone could feel it. The problem was the follow-up question: what exactly is that flavor? Can you measure it? Can you fix it?
So we built an adversarial benchmark we call style-fusion-bench. Take passages written by an LLM, mix them into excerpts from real, human-written Chinese web novels, and run blind tests on two kinds of judges: a panel of mainstream LLMs, and a human reader who lives on web fiction. Every time the writing side lost a round, we patched the writing playbook according to the verdicts and went again. Fifteen rounds between the night of the 25th and the next afternoon; thirty-plus drafts; over a hundred blind pairs.
The short version: against the LLM judges we won embarrassingly hard. Against the human we never won once.
6:30 a.m. — we thought we had it
After seven overnight rounds, the endgame test looked like this: three model judges that had never touched the iteration loop (GLM-5 and two DeepSeek versions) got a pool of ten machine passages and ten real-book passages. Their job: find the machines. To rule out luck, we re-ran the same drafts against fresh excerpt windows a second time.
Catch rate across both rounds: 0%. Fifteen pairing questions: all wrong.
The direction of the errors was the interesting part. The judges weren't confused — they were inverted. GLM-5 rated a machine-written carpenter's shop passage “human, confidence 90” for its natural interplay of gesture and interiority. The same judge rated a passage from Battle Through the Heavens, a novel with millions of readers, “machine, confidence 75,” citing “too much expository text, formulaic scene-setting.” Real books got flagged as AI. Machine drafts got praised for their lived-in warmth.
Ten mainstream models sat on that judge's bench across the experiment: GPT-5, Gemini 3.1 Pro, Qwen, several DeepSeek versions, Doubao, GLM-5, a previous-generation Claude. All inverted, all the same way. Their prior for AI flavor is still clichés, stiffness, templates — so a disciplined machine draft that avoids clichés reads as more human than an actual human. We had met this failure before, when four LLM judges scored 12% against human-consensus gold anchors on character-card copy (that experiment is written up here). This time the sample grew to ten models and the corpus switched to novel prose. Same result, sharper.
Seven hours later, a human sat down
At 7:37 the next morning we mixed the final machine drafts into real-book excerpts and handed the blind sheet to a heavy web-fiction reader. Three machine passages. One glance. Three out of three.
The tell, in the reader's words: the machine passages all pick their details the same way. Our three drafts covered a bus depot, a dye workshop, and a carpenter's shop — three unrelated subjects, one identical rustic oral-history voice in three costumes. The real excerpts were three genuinely different voices: one author rambling at the reader mid-game, one sprinting on dialogue, one storyteller lecturing. The LLM panel had read all night and missed it. The human saw it in the layout before finishing a paragraph.
That reader ran four more blind rounds over the following day. Each time, we patched the playbook per the previous verdict. Each time, caught on sight — and each catch delivered a new criterion. By the end we had six. They are the most valuable output of the whole exercise.
The six human tells
1. One voice wearing ten costumes
Invisible in a single passage, unmistakable across ten. Even our machine gatekeeper eventually admitted it: “one passage alone I might not dare call; put ten together and the same breathing, the same inventory, the same cadence of endings can't hide.”
2. Every sentence flexing
Machines chop sentences short for literary effect and hold beats in suspension. A real author's baseline is mid-length prose, with a fragment thrown maybe three times in thirty lines at an emotional peak; the machine's fragments land like a metronome. The human version of casual is an unceremonious “after a few waves of mobs” — a throwaway connective doing zero rhetorical work, then straight on with the story.
3. Written too well — the incurable one
The gatekeeper's final verdict on one draft: “eight hundred characters without a single miss — and that is the biggest giveaway.” Every detail serving the theme, nothing wasted. Real genre fiction is gloriously wasteful: The King's Avatar spends whole paragraphs on skill points and level caps, information density low and completely unembarrassed about it.
4. Paragraphs welded into slabs
“The machine version has 3 paragraphs; the human version has a dozen.” We measured after the fact: real web fiction averages 25–59 characters per paragraph, over 90% of paragraphs within two sentences, dialogue always on its own line. Machine narration defaults to slabs 4–19× thicker. Worth recording: every machine judge across twelve rounds was blind to this. Models read semantics, not layout. Humans read layout first.
5. Stopping to appraise texture
A machine line: “the cloth shoes of the townsfolk on flagstone are a muffled sound.” The reader called it instantly. It's a connoisseur pausing to grade the world — an “X is Y” appraisal. Real prose hangs the senses on something happening right now: “the courtyard went quiet for a moment; only the firewood crackled in the stove.”
6. Idle hands, busy feelings
In the fifth round the reader pointed at a machine-written player-meets-monster beat and asked: what is this passage even describing? We broke it down: rummage, listen, hold breath, turn, see monster, back off, screenshot, message a friend. Ten actions, zero emotion. The player watches the night's first five-star loot spot get trashed and feels nothing. A real player writes the sting. Real filler always carries something — spite (“well, whose fault is it he got the first kill”), want-but-daren't (“she reached out to grab it, thought again, drew her hand back”), or plain gloating (a toothpick and a well-fed look).
Every disguise we tried, dead on arrival
Knowing the tells doesn't mean you can fake past them. Deliberate typos, up to ten per piece: “errors landing in the safest spots — distressed, not slipped.” Planted contradictions: read as the seams of long-context LLM generation. Sentence-length skeletons copied line-by-line from human passages: the shape transferred, the density didn't. Instructed pointless lifelike details: “the rough edges are configured, not shed.”
And the deepest lesson: whatever repair instruction you give the model, it executes it perfectly — and the perfection is the next round's evidence. Looseness quotas got filled to the decimal. A “can't quite recall” verbal tic, prescribed once, appeared in five passages out of six. This is structural to instruction-driven writing, and no better clause fixes it. The only exit we found is choosing a model that comes with cheap, genuine flaws — real typos, real verbal habits — which is the subject of the companion piece.
What we changed because of this
First, every have-an-LLM-grade-the-AI-flavor scheme is retired here. A model panel's “reads human” now means exactly one thing — good enough for ordinary readers — and is never accepted as evidence of authenticity. Second, acceptance runs on two gates: a same-generation Claude thinking model as the development-time gatekeeper (the only machine judge whose rationale matches the human's; seven rounds, 100% catch, zero false positives), and human spot-reads before anything ships. Third, if the goal is prose that reads human, change the model before you change the playbook — discipline fixes clichés and paragraphing, it does not fix “too well written.”
One last disclosure. Before publishing this article about spotting AI prose, we ran it against our own six tells. The paragraph lengths are deliberately uneven — two sections are one sentence long. An early draft had “the data is cold, the verdicts are hot” in it; textbook appraisal sentence with a parallel-structure bow on top, deleted. Tell #3 can't be self-checked: an author can't see their own too-well-written. So by our own playbook, the final judge of this piece isn't us. It's you.
FAQ
What actually is “AI flavor” in fiction?
After fifteen rounds our answer is: not a word list, a texture. Six concrete tells: a whole batch of texts written in one identical voice; sentences chopped short for literary effect; execution so flawless nothing misses; paragraphs welded into slabs; connoisseur sentences that stop to appraise texture; and idle scenes with zero emotion in them. The first few are fixable. “Too well executed” is not — any repair instruction gets executed too perfectly, and that perfection becomes the next fingerprint.
How do I quickly check whether a passage is AI-written?
Look at the layout before you read a word. Human web fiction averages 25–59 characters per paragraph and dialogue gets its own line; machines default to slabs several times thicker. Then read the idle beats: human filler carries emotion (greed, sulking, gloating, the sting of losing loot); machine filler is an emotionless action log. Finally, zoom out: if ten passages share one voice and not a single sentence misses, suspect the machine.
Can I use an LLM as a judge to detect AI writing?
In our tests the direction was inverted. Ten mainstream models (GPT-5, Gemini 3.1 Pro, Qwen, several DeepSeek versions, Doubao, GLM-5, a previous-generation Claude) caught 0% of disciplined machine passages and flagged real published novels as formulaic, expository, AI-like. Their prior for AI flavor is stuck on clichés and stiffness, so machine prose that avoids clichés reads as human to them. The one exception: same-generation Claude thinking models went 100% catch, zero false positives, seven rounds straight — with rationales that match the human's.
Does adding deliberate typos make AI text pass as human?
No. We pushed it to 7–10 typos per piece and got caught anyway. The verdict that killed it: “the errors land in the safest possible spots — distressed, not slipped.” Real typos are risky; they hurt comprehension and land in load-bearing places. A model choosing where to err errs too tidily, and that tidiness is itself a fingerprint. Planted contradictions, sentence-length mimicry, and instructed “pointless lifelike details” all failed the same way.
Questions or ideas? Join our Discord →