Labs · Human or AI
Four LLM judges scored 12% on this test.
Your move.
One question a day. Two passages continue the same classic novel, read blind — one is what the author actually wrote next, the other is a live model imitation of the same lead-in. Pick the human, or tell two models apart by their fingerprints: slop lexicons, simile density, sentence-rhythm flatness. Then copy your scorecard and go gloat.
Free · no account · record stays in your browser
Where this tool comes from
The bank pairs public-domain prose with live model output. Human side: fourteen pre-1930 novels across three registers — literary (Austen, Wharton, Chopin, Cather, Dickens), genre fiction (Doyle, Wells, Stoker, London, Collins, Orczy) and comedy/dialogue (Jerome, Wodehouse, O. Henry) — anchored mid-book, away from famous openings and endings. AI side: four models from four vendors (Claude 4.6 Sonnet, DeepSeek V4 Pro, Qwen3.7-Max, Gemini 3.1 Pro) continue the exact same lead-in with a matched word target, temperature 0.8, 2026-07; outputs ship untouched. Model-vs-model rounds pair two models on the same anchor. Both passages in every question sit in the same 90–180-word band with the length ratio capped at 1.6, the answer side is randomized, and generations that share an 8-word window with the author's real continuation are rejected as memorization.
The fingerprint lexicon starts from community-documented English AI-slop evidence (Sukino's banned-token list, the Antislop paper's statistics, EQ-Bench's slop score) — then gets calibrated against 1.47 million words of the fourteen source novels themselves: any phrase these authors actually use was dropped or given a quota, so a Victorian habit can't be miscalled as a machine tell. Structural metrics (sentence-length spread, simile density, dialogue share) are computed per passage, and the model dossiers quote numbers measured on this bank's own generations — not vibes. The "12%" anchor comes from our double-blind judge-reliability experiment (4 judges × 7 pairs × both orders, gold from real readers, 2026-07).
What it can't do
Getting one right doesn't make you a detector, and neither are we: a structurally perfect AI passage can pass every rule on this page. Our own gold-calibration work is exactly why we say "no silver bullet for human-ness" — the daily game trains your eye, it doesn't certify it.
The human side is public-domain by necessity, so it's pre-1930 prose: you're judging a model's period imitation, not its contemporary voice. Rounds built on licensed modern human prose are planned and will be labeled when they ship.
The memorization guard rejects verbatim recall (8-word overlap with the real continuation), but a model could in principle slip past it with light paraphrase of a passage it knows — each shipped question also went through a human read, and the reveal always names the source so you can check us.
No global accuracy counter yet — this page has no backend. Your opponent for now is the fixed 12% judge score, which is real, dated and sourced.
FAQ
Where do the questions come from, and how do they rotate?
Two source pools, both with clean display rights. Human passages come from fourteen public-domain novels (Project Gutenberg — Austen, Wharton, Chopin, Cather, Dickens, Doyle, Wells, Stoker, London, Collins, Orczy, Jerome, Wodehouse, O. Henry): each question shows the author's actual next paragraphs after a mid-book anchor. AI passages are real, untouched continuations of the exact same lead-in, generated in 2026-07 by four models from four vendors (Claude 4.6 Sonnet, DeepSeek V4 Pro, Qwen3.7-Max, Gemini 3.1 Pro) with the same instruction, a matched word target and the same temperature. The day number picks the question deterministically, rolling over at midnight UTC+8; past days stay playable as practice.
What's the "LLM judges scored 12%" line about?
A double-blind reliability experiment we ran in 2026-07: four heterogeneous LLM judges rated 7 pairs of character-card prose in both orders, against gold labels from real readers. They agreed with each other 86% of the time — and scored 2/16 = 12% against gold. Unanimously wrong: they read dense, evenly-polished detail as "human" when real readers read it as AI. That number is your opponent here.
Is this fair to the models — and haven't they memorized these books?
The prompt gives each model the book title, the author, the exact preceding passage and the human continuation's word count, and asks it to continue in the author's voice — no sabotage, temperature 0.8, output shipped untouched (we only strip obvious instruction artifacts like a "here is the continuation" lead line, and record it when we do). Memorization is real: some models can reproduce a classic's actual next paragraph nearly verbatim. Any generation sharing an 8-word window with the author's real continuation is rejected from the bank, so what you're judging is genuine imitation, not recall. Every reveal names the model and the generation date.
Where is my record stored? Is there a leaderboard?
In your browser's localStorage only — no account, nothing uploaded. The English and Chinese tracks keep separate records. There's no global accuracy stat yet (this page has no backend); the 12% judge score serves as the fixed benchmark to beat. A crowd accuracy line is on the roadmap.