Is Kimi K3 good for fiction? We made it write 40 rounds to find out

Kimi K3 shipped July 16, 2026; within 48 hours we ran it through the same protocol as our nine-model benchmark — two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor K2.6. The vote: 10-0 on the fantasy epic, 11-1 on the palace novel, and the single dissent contradicted itself under mapping flips. K2.6's metaphor pile-ups and verbatim self-copying did not recur. Three honest caveats: ~85% of dialogue in Western straight quotes (a disease K2.6 never had), first-person density 1.8× the original author's, and always-on reasoning that makes every round take 37-54 seconds. Total bill for the whole experiment: ¥12.21.

Two ink-line plants side by side on aged paper: the left one twisting into knots around itself, the right one upright and blooming, with a small stopwatch resting at the right plant's roots

The verdict first. For continuing fiction, Kimi K3 beats its predecessor K2.6 without ambiguity: same protocol, two novels, 20 consecutive rounds each, six heterogeneous AI judges under two flipped blind mappings, combined vote 21 to 1 — and the single dissenting vote contradicted itself when the mapping flipped. The costs are equally unambiguous. Reasoning cannot be turned off, so every round takes 37 to 54 seconds; 81-86% of output tokens are invisible thinking, billed at $15 per million like any other output; and about 85% of its dialogue came out in Western straight quotes — a disease K2.6 never had, freshly acquired in the upgrade.

When K3 shipped on July 16, our GPT-5.6 vs Kimi head-to-head closed with a line: any claim that K3 crushes anything has no exam paper behind it. Forty-eight hours later, here is the paper.

How the exam was set

Protocol copied verbatim from our nine-model benchmark: the same two books (an 8.9-million-character Chinese fantasy epic, and the palace-intrigue classic Empresses in the Palace), the same anchor point at roughly 55% depth, the same context-labeling instruction — the final wording that fixed the Gemini restart loop — the same 16k-token window, each round’s output appended back into context, 20 rounds per chain. The only variable is the model.

The baseline took one extra step, and it matters. The palace-novel K2.6 chain came straight from the benchmark archive, which already ran on the current instruction — directly pairable. The archived fantasy chains, however, all ran on the old wording, so comparing against them would smuggle a prompt variable into a model comparison. We re-ran a fresh K2.6 fantasy baseline on the current wording. That step nearly rewrote the story: under the old wording, K2.6’s fantasy chain once measured 97.6% verbatim cross-round repetition at round 9; under the current wording, the same model peaked at 1.5%. One instruction sentence curing most of a repetition disease is a single-chain anecdote — recorded, not concluded — but using the old chain as the target would have handed K3 a rigged win. Protocol honesty is the entire credibility of every number on this page.

The judges: six AI models across five vendors (claude-4.6-sonnet, gemini-3.1-pro-preview, gemini-3.5-flash, deepseek-v4-pro, qwen3.7-max, glm-5.2), temperature 0. We use AI judges here with a documented precondition: our judge-reliability experiment showed they fail systematically at absolute calls like “which reads more human”, while paired relative judgments on the same setup are within their reach — and every call in this round is paired. The reference excerpt in each blind pack was cut strictly from before the anchor, so no judge ever saw the original’s actual next passage.

21 to 1, and the dissent that dismantled itself

Fantasy: 10-0, plus two invalid ballots with some comedic value — the claude judge twice refused to return a verdict and started continuing the novel instead. Palace: 11-1. Split by early, middle, and late windows, every window votes the same way; there is no “strong start, weak finish” story hiding in the aggregate.

The one dissent deserves its paragraph. In blind pack 2201, the deepseek-v4-pro judge picked option A, which mapped to K2.6. In pack 2203 — same passages, mappings flipped — it picked A again, which now mapped to K3. Two votes that prefer whichever text sits in the first position are a position bias, not a judgment; the other five judges chose K3 consistently across both mappings. Running two mappings per book exists precisely to catch this ballot.

We mechanically verified the judges’ stated reasons against the actual chains. K3 was praised for reproducing the original’s short paragraphing, onomatopoeia standing alone on its own line, and the author’s tilde habit — and its fantasy chain really does contain tilde-marked rumbling set as its own paragraph, a texture that only GPT-5.6 Terra reproduced across the entire nine-model field. Time-skip transitions in the original’s manner and four to six natural uses of the book’s own technique and artifact names all check out. K2.6 was dinged for piled-up rhetoric and long, rushed battle scenes, and several judges independently flagged a rewind feel: its new baseline chain cured the verbatim copying but circled at the plot level — the same villain attacks at round 9, retreats at 11, returns at 17, and is still fighting at 19. K3, same book, same anchor, pushed through to a new arc; the palace chain likewise ran twenty rounds of strictly linear plot with no backtracking.

What the upgrade cured, in numbers

Sentence-length variance (CV) is the first AI fingerprint we measure — machine prose runs on uniformly sized sentences. Both books, original baseline included:

Metric (closer to original is better)Kimi K3Kimi K2.6 (same protocol)Original
Fantasy · sentence-length CV0.620.500.837
Fantasy · cliché density (per 1k chars)2.023.681.0
Palace · sentence-length CV0.4430.4170.495
Palace · cliché density (per 1k chars)0.631.921.0
Palace · first-person density (per 1k chars)16.211.09.0

K3 beats K2.6 on sentence CV on both books, and its fantasy chain climbs to 0.731 in the late windows, approaching the original’s 0.837 — in plain terms, it learned to mix long and short sentences and to let a short one stand alone. The trait that anchored K2.6 to seventh place in the fantasy benchmark, the highest metaphor density in the field, is essentially gone: K3’s palace chain runs at 0.63 clichés per thousand characters, below the original author’s own 1.0, and its fantasy chain opens high (13.1 in round one) then collapses to long stretches of zero.

The long-run failure checkup is equally clean. Against our three failure modes — restart loops, mid-run freezes, ending rewinds — K3’s cross-round verbatim repetition peaked at 6.8% (fantasy, round 10; near zero elsewhere) and 1.0% (palace). No loop, no freeze, no rewind. The self-copying label K2.6 earned in the benchmark did not carry over.

Three honest caveats

One: straight quotes. About 85% of K3’s dialogue used Western straight quotes instead of the full-width marks Chinese prose expects — 39 pairs to 9 on the fantasy book, 52 to 8 on the palace book. This is the same chronic tic we documented on GLM-5.2, and K2.6 never had it; the upgrade brought it in. It even fooled our own detection regex, which only counted full-width quotes and briefly scored both chains at ~4% dialogue; the corrected figures are 20.4% and 38.9%. A quote-format constraint in the prompt is worth trying, but on the GLM precedent, do not expect one sentence to cure a native habit.

Two: too much “I”. On the palace novel — first-person, restrained — K3 wrote the first-person pronoun at 16.2 per thousand characters against the original’s 9.0, nearly 1.8×. The direction is the exact opposite of Grok 4.5’s under-use on the same book. If your book runs on a quiet narrator, watch this one.

Three: slow and expensive. K3 currently ships max reasoning only, no off switch: 81-86% of output tokens are invisible thinking, visible prose runs about 250 characters per round, and the bill counts all of it at output price. Our pessimistic pre-run estimate was one to three hours per chain; actual wall clock came in at 12.4 and 17.9 minutes — better than feared, but the 37-to-54-second wait per press is real inside a reader. Measured cost: roughly four cents per continuation, for work K2.6 does for under two.

A proper-noun Easter egg

K3 used the novel-original name of Ling Rong’s residence, Cunju Tang, correctly fourteen times — that name appears 52 times in the book. But for the Empress’s palace it wrote Jingren Palace, which is the TV adaptation’s name and appears exactly once in the whole novel; the book’s own system is Zhaoyang Hall and Fengyi Palace. Its memory of this story leans slightly toward the show, the same family of evidence as the palace-name probe we published in Chinese: one proper noun can tell you which version of a story a model actually absorbed.

The window that never entered the exam

Expectation management: this entire 21:1 happened inside a 16k-token window. K3’s million-token context never came into play, and our 1M control experiment explains why we did not reach for it — cramming a whole novel into context does not transfer style, and filling K3’s window once costs about $3.15 of input before a single word comes back.

Where K3 stands, and the bill

Careful wording, matched to what the data can carry. On composite structural distance to each original’s profile (lower is closer), K3 scored 0.388 on the fantasy epic against champion DeepSeek V4 Flash’s 0.356, and 0.396 on the palace novel against champion V4 Pro’s 0.280. Inside the top tier’s neighborhood on both books; past the champion on neither. This round paired K3 only against K2.6, and its judging format differs from the nine-model leaderboard (full ranking there, pairwise verdicts here), so we cannot write “K3 ranks Nth” yet. What we can write: it jumped from K2.6’s lower-midfield into the champions’ neighborhood in one generation. A paired K3-versus-DeepSeek run is the next exam.

The bill and the boundaries. K3’s side of the experiment — 41 requests including a connectivity probe and two empty-response retries — cost ¥12.21, about $1.70, with both chains finishing 20 of 20 rounds without interruption. Specs and pricing per the official K3 docs and pricing page (verified 2026-07-18): launched July 16, 2026, 1M-token window at flat pricing, $3.00 per million input tokens on cache miss, $0.30 on hit, $15.00 per million output. Limits, copied from the lab notes: one chain per book, one anchor point, n=1 per condition — the same budget every model got in the benchmark, and still n=1.

Whether to move a book you are mid-way through onto K3 is not a question this page should answer for you. Same opening, three continuations from each model, read them shuffled, flip the cards after — in a reader that switches models per paragraph, the test costs a few cents and the loser costs you nothing to leave behind.

FAQ

Is Kimi K3 good for fiction?

Clearly better than its predecessor, with receipts: under the same protocol, K3 and K2.6 each continued two Chinese novels for 20 consecutive rounds, and six heterogeneous AI judges under two flipped blind mappings voted 10-0 (fantasy) and 11-1 (palace intrigue) for K3 — the single dissent contradicted itself when the mapping flipped. Sentence-length variance moved closer to the original author on both books, and K2.6's signature metaphor pile-ups and self-copying did not recur. Three caveats: ~85% of dialogue in Western straight quotes, first-person density 1.8× the original's on the palace novel, and 37-54 seconds per round from always-on reasoning. No head-to-head against DeepSeek yet, so no field-wide rank.

Should I upgrade from Kimi K2.6 to K3 for creative writing?

The quality direction is unambiguous — 21:1 in paired double-blind — and so is the price of admission. International pricing moves from $0.95 to $3.00 per million input tokens and from $4.00 to $15.00 per million output; our measured cost per continuation went from under two cents to roughly four cents, every round waits half a minute longer, and K3 picked up a straight-quote habit K2.6 never had. If budget or latency matters, staying on K2.6 is defensible. If you are unsure, run the blind test we always recommend: same opening, three continuations from each model, read them shuffled, then flip the cards.

Does Kimi K3's 1M context window help with fiction?

This entire 21:1 result happened inside a 16k-token window; the million-token window never entered the exam. Our five-tier control experiment showed that cramming a whole novel into context with no style instruction still produces boilerplate similes at ten times the original author's density — volume does not transfer style. Filling K3's window once costs about $3.15 of input before it writes a word. The big window earns its keep on one-shot whole-book tasks like summarization and character audits, not on repeated continuation.

How much does Kimi K3 cost per continuation?

Official international pricing (verified 2026-07-18): $3.00 per million input tokens on cache miss, $0.30 on hit, $15.00 per million output, flat across the 1M window, currently max-reasoning only. Measured: our two 20-round chains, 41 requests including a connectivity probe and retries, cost ¥12.21 total (about $1.70) — roughly four cents per continuation. Note that 81-86% of output tokens were invisible reasoning, billed at full output price; visible prose ran about 250 Chinese characters per round.

Is Kimi K3 better than DeepSeek for fiction?

No exam paper yet, so no verdict. Indirect evidence: on composite structural distance to the original author's profile (lower is closer), K3 scored 0.388 on the fantasy epic versus champion DeepSeek V4 Flash's 0.356, and 0.396 on the palace novel versus champion V4 Pro's 0.280 — inside the top tier's neighborhood on both books, ahead of the champion on neither. This round's judging format also differs from the nine-model leaderboard, so cross-table conversion stops there. A paired K3-versus-DeepSeek blind run is our next exam; until it lands, claims in either direction have no paper behind them.

Questions or ideas? Join our Discord →

Kimi K3 Creative Writing Test: 40 Rounds of Novel Continuation, 48 Hours After Launch · Foreverse · Xinmeng