K3 vs DeepSeek: alignment today, verdict pending

Every K3-vs-DeepSeek comparison online is about code. We have fiction data on both, measured with the same ruler: on the fantasy epic K3's composite distance-to-the-original is 0.388, between DeepSeek V4 Flash's 0.356 and V4 Pro's 0.454; on the palace novel it's 0.396 against Pro's 0.280. Cost is a rout — one K3 continuation buys about twenty-five on Flash. What we refuse to print is a winner: the two sides have never met in the same blind review, and until that paired duel runs, cross-table rankings are protocol noise dressed as a verdict.

Two writing desks facing each other across an ink-dark river, one stacked with dense tidy manuscript pages, the other sparse and scattered — warm-paper hand-drawn ink illustration

Within 48 hours of K3’s launch, the comparison posts arrived in a wave — and every one we found was about code. SWE-bench deltas, agent completion rates, autocomplete latency. Nobody had fiction numbers. We do, and they happen to be measured with the same ruler: K3 just finished two books of 20-round continuation under our benchmark protocol, and DeepSeek’s two tiers already sat in our archive — same books, same starting points, same metric definitions.

The two hardest numbers first. Fantasy epic: K3’s composite distance-to-the-original is 0.388, landing between DeepSeek V4 Flash’s 0.356 (that run’s blind champion) and V4 Pro’s 0.454. Palace novel: K3 scores 0.396 and does not reach Pro’s 0.280. Now the sentence that qualifies both: this is an alignment of two scoreboards, not a duel. K3 and DeepSeek have never sat the same blind exam, and nobody has beaten anybody on paper. If you want the buying advice, skip to the decision tree at the end. If you want to know why we won’t print a winner, the next section is the reason.

Two exams, one ruler

The scoreboards come from different rooms. K3’s opponent was its own predecessor: same two books (an 8.9-million-character fantasy epic and the palace classic Empresses in the Palace), same openings, 20 consecutive rounds each, outputs anonymized and judged pairwise by six heterogeneous AI judges under flipped mappings — a 21:1 sweep over K2.6. DeepSeek’s placements come from our multi-model blind reviews, where anonymized outputs got full rankings from two independent reviewers. Books, openings, window sizes, and metric definitions are identical across all of it, which is why the structural numbers can share a table. The review formats are not, which is why the rankings can’t. Splicing “K3 swept K2.6” onto “Flash won the fantasy run” manufactures a match that never happened.

The alignment also ships with its own warning label. On the palace novel, V4 Flash’s distance score is 0.296 — nearly touching Pro’s 0.280 — yet the blind reviewers ranked Flash 6th–7th and Pro in the top tier. Surface metrics measure texture; they cannot measure “a finer brush that happens to be the original author’s brush.” Metrics sort models into tiers. They do not hand out medals. We measured that limitation on our own instrument, and it caps everything this page can claim.

Table one: textual closeness

The distance score averages eight structural gaps against the original author — sentence length, sentence-length variance, dialogue share, punctuation density, stock-phrase density among them. Lower is closer. One script, four chains, two books:

ModelFantasy epicPalace novel
DeepSeek V4 Flash0.356 (blind champion of that run)0.296 (yet ranked 6th–7th blind)
Kimi K30.3880.396
DeepSeek V4 Pro0.4540.280 (top tier blind)
Kimi K2.60.5040.400

On fantasy, K3 wedges itself between the two DeepSeek tiers — one step behind the champion, ahead of Pro and of GLM 5.2’s 0.565. One of its single metrics is untouched by anything we have tested: sentence-length variation coefficient 0.62 overall, 0.731 in the closing window, nearest of the field to the original’s 0.837. That choppy long-short rhythm is the deepest fingerprint of old-school fantasy serials, and the uniformity disease most models can’t shake. On the palace novel, K3 is simply mid-table: 0.396 against Pro’s 0.280 is not error-bar territory.

One accounting note. About 85% of K3’s dialogue used Western straight quotes, and the dialogue-share metric penalizes that under the raw counting rule (corrected, its dialogue share reads 20.4% on fantasy, 38.9% on palace). The table keeps the unadjusted — conservative — numbers, and GLM 5.2’s 0.565 suffers the same artifact. But the artifact is also just a real defect wearing a lab coat: readers of Chinese fiction see those straight quotes on the first page.

Long-run failure: both sides finish the race

Our 360-round benchmark sorted long-run collapse into three species — restart loops, mid-run freezes, ending rewinds (the taxonomy lives in the failure-mode study). On this axis K3 and both DeepSeek tiers tie: clean. K3’s cross-round verbatim repetition peaked at 6.8% (fantasy, one round) and 1.0% (palace), near zero elsewhere, with 20 rounds of linear plot and no backtracking. Neither DeepSeek tier has ever appeared on a failure list in either review. For contrast: K2.6 once measured 97.6% verbatim self-copying at round 9 of an earlier-protocol fantasy chain, and across the nine-model field the three failure species hit four systems. Finishing 20 rounds intact is the entry ticket both vendors hold — and the reason a shared blind exam is worth running at all.

Table two: cost — this round is a rout

Per million tokens (international list)DeepSeek V4 FlashKimi K3
Input, cache miss$0.14$3.00
Input, cache hit$0.0028$0.30
Output$0.28$15.00

List prices from the DeepSeek pricing page and the Kimi platform docs, both checked 2026-07-22. On top of the sticker gap, K3 carries a thinking tax the table doesn’t show: its reasoning stays on, 81–86% of the output tokens on our chains were invisible thinking billed at the full $15 rate, visible prose ran about 250 Chinese characters a round, and every round waited 37–54 seconds. In money: the whole two-book experiment cost ¥12.21 (about $1.70, roughly four cents per continuation) against Flash’s $0.0016 per segment. One K3 segment buys about twenty-five on Flash. A footnote for completeness: DeepSeek’s domestic platform has announced peak-hour doubling, but the international pricing page still lists a single flat rate as of our check, and the sunset calendar tracks that switch. No suspense in this section either way.

Case files: what each has cured, what each still carries

DeepSeek’s chart is well documented across two reviews. Flash’s strengths are vocative dialogue, onomatopoeia set as standalone paragraphs, and fast plot advance; the reviewers’ line was that it was the only system that read like a continuation chapter of the original. Its disease is genre mismatch: a plain, fast-talking instinct that collapses against ornate court diction, which is exactly where it fell to 6th–7th. Pro is the mirror image (the only model to reliably reproduce the palace novel’s double-layer “words within words, then narrated decoding” structure), and its disease is systematically overgrown description; the fantasy reviewers called it a finer-penned author of the same genre. Two tiers, one vendor, neither able to take the other’s home ground.

K3’s file, cures first. K2.6’s signature metaphor pile-ups are essentially gone: fantasy stock-phrase density fell from 3.68 per thousand characters to 2.02, with the closing stretch near zero; on the palace novel it runs 0.63 against the original author’s own 1.0. The verbatim self-copying didn’t recur. The judges’ praise was specific — faithful short paragraphs, onomatopoeia standing alone, even the original’s “~~~” punctuation habit — a texture that, among the nine models we’d tested before, only GPT-5.6 Terra had matched. Proper nouns from the fantasy canon appeared four to six times, used correctly.

Still on the chart, copied verbatim from the record: straight quotes on ~85% of dialogue (39:9 on fantasy, 52:8 on palace — GLM 5.2’s exact tic); first-person density 1.8× the original on the palace book, the opposite failure direction from Grok 4.5’s sparseness; and it named the Empress’s residence after the TV adaptation’s palace instead of the novel’s — one proper noun betraying which corpus its memory leans on, a probe family we’ve written up in Chinese.

The decision tree

Budget-sensitive, or running a model as a daily reading companion: DeepSeek, tiered by genre (Flash for fast-paced fantasy, Pro for ornate period prose) at pennies per evening. Chasing texture, the choppy-rhythm and standalone-onomatopoeia kind of resemblance: K3 earns an audition, if four cents and a forty-second think per segment read as a fair price; its full sample pages are in the 20-round file, worth reading before topping up. Want both: mix them. Route the pivotal chapters to K3 and the daily mileage to DeepSeek. Continuation in Foreverse switches models per segment, both vendors ride in over BYOK keys, and no chapter forces you to choose a side.

Which leaves the unrun exam. The paired K3-versus-DeepSeek blind review is scheduled, both books, both DeepSeek tiers; when it lands, this page’s alignment becomes a verdict and the date under the title changes. Until then, any “K3 beats DeepSeek” you read — in either direction, including from us — is a scoreline from a match nobody has played.

FAQ

Which writes better fiction, Kimi K3 or DeepSeek?

We can align their scoreboards but not call a winner. Same ruler — same two books, same starting point, 20 consecutive continuation rounds each, identical metric definitions: on the fantasy epic K3's composite distance-to-original is 0.388, between DeepSeek V4 Flash (0.356, that run's blind champion) and V4 Pro (0.454); on the palace novel K3 sits at 0.396, short of Pro's 0.280. The two have never faced the same blind panel, so nobody has beaten anybody on paper.

Why not just run a blind vote and declare a winner?

Because a fair verdict needs both outputs in the same review, under one protocol, judged by reviewers who don't know which is which. K3's paired opponent so far was its own predecessor K2.6 (a 21:1 sweep); DeepSeek's ranks come from separate multi-model blind reviews with a different judging format. The structural metrics were computed by one script on the same two books, so those numbers can share a table — rankings can't. The K3-versus-DeepSeek paired blind run is scheduled, and this page gets rewritten when it lands.

How much more expensive is K3 than DeepSeek?

At international list prices (checked 2026-07-22): K3 charges $3.00 per million input tokens against V4 Flash's $0.14 — about 21× — and $15.00 per million output against $0.28, about 54×. Measured on our runs the gap compounds to roughly 25× per continuation, because K3's reasoning is always on and its thinking tokens bill at full output price: about $0.04 per segment versus $0.0016. Our whole two-book, 40-round K3 experiment cost ¥12.21, roughly $1.70.

What are K3's known quirks for fiction?

Three from the test record, one small and telling. About 85% of its dialogue came out in Western straight quotes on two Chinese novels (fantasy 39:9, palace 52:8 — the same tic GLM 5.2 carries). First-person density on the palace novel ran 1.8× the original author's (16.2 per thousand characters against 9). Reasoning can't be switched off, so every round costs a 37–54 second wait plus the thinking tokens. And it wrote the Empress's palace under the TV adaptation's name rather than the novel's — a one-word tell that its memory leans toward the show's corpus.

Questions or ideas? Join our Discord →

Kimi K3 vs DeepSeek for Fiction: the Numbers We Can Align, and the Duel We Can't Call Yet · Foreverse · Xinmeng