The em dash is not an AI tell in English — we counted
The em dash got branded the “ChatGPT hyphen” and writers started scrubbing it out of their own prose. We went to measure it before adding a dash rule to our English AI-flavor detector — and found that 8 of the 30 top-starred human roleplay cards use em dashes above the density that flags a Chinese text as machine-written, while 2 of 5 tested model families use none at all. Any dash threshold that catches AI also false-flags a quarter of the best human cards. The numbers, the method, and the rule we shipped instead.

Somewhere in 2025 the em dash picked up a nickname — the "ChatGPT hyphen" — and a reputation as the laziest possible way to catch AI text. Rolling Stone covered the discourse; the Washington Post counted em dashes in over half of ChatGPT's replies by summer 2025, up from under one in ten a year earlier; by November, OpenAI was announcing a fix for the model's refusal to stop using them. Writers, meanwhile, started scrubbing a legitimate punctuation mark out of their own prose to avoid the allegation.
We had a concrete stake in whether the accusation holds. We ship a rule-based AI-flavor detector for character-card prose, and in its Chinese calibration the full-width long dash is the single most reliable machine fingerprint we have — our Chinese house rules cap it at roughly 2–3 per thousand characters, and machine text blows through that constantly while human web-fiction prose almost never does. When we built the English version of the detector this July, the obvious move was to port the rule.
We measured first. The rule didn't survive the measurement.
The count
Human side: the first messages of the 30 top-starred SFW character cards on chub.ai — human-written, star-validated, sampled 2026-07. We stripped quoted dialogue (so a character who talks in dashes doesn't skew the narration count), counted em dashes and double-hyphens, and normalized per 1,000 words. Machine side: 9 raw greetings from 5 current model families, generated with a plain prompt and no style instructions.
The human distribution is the interesting one. It's bimodal:
22 of 30 cards contain zero em dashes. The other 8 all sit above 3 per 1,000 words — the density that would flag a Chinese text — topping out at 20.6. There is nobody in between. Humans either don't touch the dash or they lean on it hard, and the leaners include some of the highest-starred cards in the sample. Corpus-wide, that's 3.4 em dashes per 1,000 words of narration in the best human RP prose — comfortably above the density people cite as suspicious.
The machine side refuses to cooperate with the stereotype from the other direction:
| Model family (raw single-shot greeting) | Em dashes per 1,000 words |
|---|---|
| Claude family (2 samples) | 12.3 / 17.3 |
| GPT family (2 samples) | 13.9 / 0 |
| Kimi (1 sample) | 11.0 |
| GLM family (2 samples) | 4.5 / 4.8 |
| Gemini family (2 samples) | 0 / 0 |
Six of nine samples exceed the 3-per-1,000 line, so yes — models as a population love the dash. But one major family produced none at all, in both samples. A detector keyed on dash density waves the dash-free models straight through while convicting the eight heaviest human users in the sample. That's 27% of the best human cards false-flagged, in exchange for a rule some models never trip.
Why the rule works in Chinese and fails in English
The tell was never the punctuation. It was the deviation from the human baseline, and baselines are language-specific. Chinese web-fiction prose essentially doesn't use the double-width dash, so machine overuse stands out like a flare. English narrative prose — and the asterisk-roleplay register in particular, which grew out of dash-happy published fiction — uses it as ordinary furniture. Same mark, different norm, opposite verdict.
This generalizes into a rule we now enforce on ourselves: AI tells do not port across languages without re-measurement. Our English detector nearly shipped with a rule that would have false-flagged a quarter of the very corpus it was calibrated to protect.
What separates, if the dash doesn't
In the same calibration: stock-phrase cluster density (human top-card greetings average 0.07 hits against our 95-phrase lexicon; raw model greetings average 0.22), the "not X, but Y" construction, emotion cocktails, and fishing-line endings. The full list with sources is in the slop lexicon post. And even those separate weakly on frontier models in single-shot mode — word-level detection is a lint, not a verdict, a point our LLM-judge experiment makes at length (the judges scored 12% against human-consensus anchors).
The em dashes in this post — there are a few — were typed on purpose, which is exactly the problem with the accusation: you can't tell. If you're grading essays or moderating a forum by dash density, our numbers say you're wrong about a quarter of your best writers, and you'll never hear about it, because the accused just quietly stop using the mark. We nearly automated that mistake at scale. The count took an afternoon.
FAQ
Is the em dash proof that a text was written by ChatGPT?
No. It's weak evidence at best and it false-flags heavily. In our July 2026 count, 8 of the 30 top-starred human-written English roleplay cards used em dashes above 3 per 1,000 words of narration — the density that reliably flags machine text in Chinese — with one human card hitting 20.6. Meanwhile two of the five current model families we sampled produced zero em dashes in their raw output. A punctuation mark that a quarter of the best human writers lean on and some models never touch cannot carry an accusation.
Why is the long dash an AI tell in Chinese but not in English?
Because the human baselines differ. Chinese web-fiction and roleplay prose almost never uses the double-width dash —— so when a model produces it constantly, the deviation from the human norm is enormous, and in our Chinese detector's calibration it is the single most reliable machine fingerprint. English narrative prose — especially the asterisk-roleplay tradition — inherited heavy dash usage from published fiction. Same punctuation, different baseline, opposite verdict. AI tells are language-specific; porting them across languages without re-measuring produces false accusations.
What should you check instead of em dashes?
Cluster density of stock phrases (in our calibration, top human card greetings average 0.07 slop-lexicon hits; raw model greetings average 0.22), the "not X, but Y" contrast construction, emotion-cocktail formulas ("a mix of anticipation and dread"), and fishing-line endings ("the choice is yours"). None of these is proof either — flagship models in single-shot mode now pass phrase checks clean — but they separate in the right direction, which the em dash does not.
Questions or ideas? Join our Discord →