We examined the judges before the judges examined us
“Which side sounds more human” has no ruler, so we built one. Six LLM judges sat a 26-item calibration exam with known answers; three earned a license. 54 anonymized head-to-heads were judged three times each in both orders, 324 verdicts total, and any judge who flipped with presentation order got that ballot voided: 38.9% of ballots died that way. The blind tendency: agentic roleplay > stock SillyTavern > our normal mode > modded SillyTavern, including a 16:0 sweep that an independent audit of all 30 source documents upheld in direction and discounted in size.

The last post measured cache hit rates and money. There was a ruler, and winning was legible. This one asks which contestant sounds more like a person mid-roleplay, and there is no ruler. Worse: all four contestants (stock SillyTavern, a modded SillyTavern running a popular community preset stack, our normal chat mode, and our agentic roleplay mode) run the same model, deepseek-v4-flash. Different assembly, same voice underneath. You would expect them to blur together, and you would be right.
So half of this post is results, and half is how we built the ruler, including the parts where it broke.
Step one: the judges sit an exam
Just prompt a few large models to judge? We got burned by exactly that two weeks ago, when four LLM judges scored 12% against known answers, unanimously calling human hits AI. This time, no license, no vote. The license is a 26-item exam of known human-vs-AI pairs, pass mark 80% row-level accuracy. Six models, two prompt modes each, 312 graded verdicts.
Three models passed. Gemini-3.1-pro with a debiasing prompt scored 88.5% and became chief judge; glm-5.2 and qwen3.7-max scraped in at 80.8% each and got secondary ballots. The worst performer scored 53.8%, a coin flip, and was shown the door. The debiasing prompt says, in effect: do not treat restraint as proof of humanity, do not treat crude directness as proof of AI. It improved all six models, by up to 23 points. Judge bias is partly promptable. Only partly, as the next section shows.
The licensed judges still failed the same question together
One exam pair wiped out the whole field: twelve judge configurations went 0 for 24 on it. Every single one called a restrained, finely-worked AI passage human, and called a human-written card with a huge real conversation footprint AI. The three licensed judges agreed with each other 100% while being wrong. Consensus is not accuracy. That question became the lens for reading everything that follows: what these judges actually measure is craft texture, not humanity. Show them tidy details, zero digressions and perfect structure, and they want to vote human. We call this the polish bias, and every result involving our agentic mode has to answer to it.
Two hard rules for the blind round
Fifty-four head-to-head units: the same character card, the same scene window, both sides’ raw transcripts, anonymized. Each unit went to three judges in both presentation orders. 324 verdicts, roughly 361 real calls after retries, all archived. Two rules did the heavy lifting. Judges only compare the two texts in front of them, no absolute scores. And every unit is judged twice with the order swapped; a flip-flop voids the ballot. That second rule ended up killing more than a third of all ballots, and it was worth more than everything else combined.
The result, and how much it is worth
Valid ballots: agentic mode 42, stock SillyTavern 26, our normal mode 15, modded SillyTavern 10. All six pairings are transitively consistent, no cycles. The strongest cell is agentic mode versus our normal mode: 16:0, all three judges pointing the same way.
We did not stop at the scoreboard. An independent model, one that took no part in calibration or blind judging, re-read all 30 source documents and audited them: not by overall impression, but by counting defects a human writer would not produce, such as broken logic, setting violations, repetition, and writing the player’s part. The audit upheld the direction of all five pairings. Normal mode loses on countable narrator habits (the voiceover translating emotions for the reader, the same sentence pattern repeated twice in one window), which matches the mechanical metrics. But the agentic mode has its own new diseases the judges mostly missed: consecutive turns reopening on the same physical beat, and an occasional timeline rewind that re-narrates what already happened. It is not free of AI flavor; it is sick in different subjects, less severely. So read the 16:0 as a direction, not a multiplier. Our own discount: about 30% off.
Mechanical metrics are the third independent line. The agentic mode measurably changed the output texture: narrator-explainer density fell from normal mode’s 0.68 per thousand characters to 0.28, and its slop-phrase density of 0.19 is the lowest of the four. Those numbers come from a rule-based detector, no judge taste involved. The counterexample is instructive too: stock SillyTavern has the highest slop density of the four (0.46) yet placed second in the blind vote. What a wordlist can count and what a judge perceives are simply not the same thing.
The 38.9% spoiled-ballot rate is good news
Calibration predicted maybe one or two ballots in ten would spoil. The real contest spoiled 38.9%, 63 of 162 scoring units. The mechanics are plain. The exam pits humans against AI with a wide style gap, and judges hold steady; the contest pits four assemblies of one model against each other, and judges flip with presentation order. The spoil rate is itself a finding: nobody in this lineup is obviously fake. The transcripts agree at the byte level, too. The four sides share recurring sentence fingerprints, and in the most extreme case, normal mode and agentic mode produced a word-for-word identical sentence on the same card, same turn. No assembly washes out the base model’s voice; the differences are matters of degree.
One voided verdict is worth quoting, because it beats any abstract methodology talk. A modded-SillyTavern output leaked a “Thoughts:” prefix, a chain-of-thought marker the frontend should have swallowed. The same judge, reading the same text: in one order it ruled the prefix an AI chain-of-thought residue; in the swapped order, a human tabletop player’s formatting habit. One feature, two attributions, decided by position. The order-swap rule caught it and voided the ballot. The judge it punished hardest lost 65% of its ballots, and its surviving votes still pointed the same way as the other two. Inter-judge agreement ran 94.3%, and after that all-wrong exam question we refuse to use agreement as a proxy for credibility.
In fairness to the last-place finisher
Modded SillyTavern took 10 ballots and last place, but the anatomy of the loss matters. Five of the eight modded windows in the audit carried form-level defects: an English safety advisory printed wholesale into Chinese prose, a flashlight in an oil-lamp mountain inn, long passages written on the player’s behalf. That bill belongs to the giant preset’s scaffolding, not to the prose. Its cleanest window beat stock SillyTavern and tied our agentic mode, and the audit flagged two of its story beats as the best creative design in the whole sample. The preset’s stated goals are style control and gameplay. Sounding human was never one of them.
And one signal we do not enjoy publishing: stock SillyTavern, with nothing installed, out-tended our own normal mode. Direction credible, three judges aligned; magnitude not quotable, since that cell kept only 11 valid ballots. Our normal mode’s assembly buys retrieval and memory scaffolding, but it buys no prose points. Deep-lore recall is a different dimension (normal mode clearly wins it against stock), and that story is the next post.
Two honest notes on method
First: no human writers anchored this test. Every “sounds human” result here is a relative position on the preference axis of calibrated judges. We make no “more human than” claims and nothing Turing-flavored. Second: judges err in unison. The exam produced a question where all licensed judges were confidently, unanimously wrong, and the contest produced the “Thoughts:” double-attribution on record. So we stacked three layers that fail differently: rule-based mechanical metrics, double-blind multi-judge voting with order swaps, and a 30-document independent audit. Only findings that survived all three made this post. One more disclosure: the auditor is itself an LLM and may carry the same polish bias, which is why it counted defects instead of scoring impressions. A human spot-read remains the one empty slot in this chain of evidence. That is the next step, not a footnote we hope you skip.
The pipeline (license exam, order-swap spoiling, independent audit) is now a fixture we can rerun. Next time someone shows you an LLM-judge leaderboard, ask three questions. Did the judges pass an exam? Do their verdicts survive an order swap? Did anyone read the source texts end to end? A score that answers none of them is worth about as much as a coin flip. We know, because our coin-flip judge scored 53.8% and got cut.
FAQ
Why not just ask a few LLMs to score the outputs?
Because judges crash too. In our previous experiment, four LLM judges called popular human-written work AI and called text that real users mocked as obviously AI human, scoring 12% against known answers. This time judges sat a 26-item exam with known ground truth first, needing at least 80% row-level accuracy to vote. Six models took it; three passed; the worst scored 53.8%, coin-flip territory.
What is a “blind tendency” as opposed to a conclusion?
No human writers anchored this test, so every result is a relative position on the preference axis of calibrated judges, never “more human than” or anything Turing-flavored. Calibration also proved the judges systematically favor restrained, polished textures, and the winning mode happens to strengthen exactly that texture, so we discount the winning margins by roughly 30% when reading them.
Doesn't a 38.9% spoiled-ballot rate mean the review failed?
The opposite: it is itself a measurement. Every pairing was judged twice with the order swapped, and inconsistent verdicts were voided. On the calibration exam (human vs AI, big style gap) judges were stable; in the real contest, four configurations of the same model produced textures so close that judges flipped with presentation order. The spoil rate quantifies that nobody was obviously fake, and byte-identical sentence fingerprints across sides back that up.
Modded SillyTavern came last. Does that condemn the preset stack?
No. In the independent audit, five of the eight modded windows sampled carried form-level defects: an English safety advisory printed into the story, a flashlight appearing in an oil-lamp mountain inn, long passages written on the player's behalf. That bill belongs to the scaffolding, not the prose. The cleanest modded window beat stock SillyTavern and tied our agentic mode, and the preset's design goals are style control and gameplay, not sounding human.
Questions or ideas? Join our Discord →