Kimi's 256K window: what it actually buys you when continuing a novel
Kimi K2.6 carries a 256K context window and K3 a full million — and the inference readers draw, 'big window means faithful continuation', fails our tests. In the nine-model blind benchmark K2.6 placed 5th–7th as the only system that improved over 20 rounds, which had nothing to do with window size; a five-tier control showed a whole novel crammed into context still yields clichés at ten times the original's density, billed in full every round. Where the window genuinely wins: one-shot whole-book jobs. Setup steps and July 2026 international prices included.

Kimi’s model page leads with the window: 256K tokens on K2.6, a full million on the new flagship K3. And the inference most people carry into a search for “kimi write fiction” follows straight from the spec sheet — a window that big swallows the whole book, so the continuation must come out faithful. We happen to hold two datasets that test that inference directly: a nine-model blind benchmark on a court novel, and a five-tier context-size control running from 4k tokens up to 200k. This post is the verification report.
The rank: mid-table, with the field’s most unusual trajectory
In the nine-model benchmark — each system continuing an ornate court classic for 20 consecutive rounds, outputs anonymized and ranked by two independent reviewers — Kimi K2.6 placed 5th–7th. The note beside the rank matters more than the rank: the only system whose output improved as the run went on. It opened in a shouty webnovel register and settled, round by round, into period decorum. For calibration, the top tier of that table was DeepSeek V4 Pro (1st–2nd) and GPT-5.6 Terra (1st–3rd), and the fantasy-run champion V4 Flash managed only 6th–7th on this book.
Improvement over a long run is rare in our data. Most models drift the other way: sentence lengths flatten, metaphors pile up; we quantified that downhill slope in the style benchmark. K2.6 is the one system both reviewers independently described as settling in rather than wearing out. The counterweight comes from its fantasy chain: a recorded self-copying habit, whole passages lifted from its own earlier rounds. Run long, and when a paragraph reads like déjà vu, a single reroll — or a one-segment switch to another model — breaks the loop.
Does cramming the whole book into 256K make the output more faithful?
Tested: no. Our five-tier control fed a model between 4k and 200k tokens of original prose (the top tier being an entire 270,000-character novel, crammed in whole) with no style instruction attached. The output still ran on stock similes at ten times the original author’s density. The window solved “does it fit”; it never touched what the material is for, which takes an explicit instruction and a structured feed. In a reader, structured means character sheets plus lorebook entries that enter the request only when the current scene triggers them. Loading the right things beats loading everything — the million-token post runs this argument against every flagship’s spec sheet at once.
The billing side is blunter. Continuation repeats, and a stuffed window re-bills the whole book as input on every segment. At K2.6’s international rate of $0.95 per million input tokens, filling 256K once costs about 24 cents — against roughly one cent for a curated 10,000-token slice. Moonshot’s context caching softens repeat sends, but the discount only holds while the prefix stays identical; slide the window or edit one upstream paragraph and the meter resets. The faithfulness problem stays unsolved, and the bill runs twenty-plus times heavier.
Where the long window earns its spec-sheet billing
One-pass, whole-book jobs. Summarize an entire arc. Audit every scene a character appears in. Draft lorebook entries from the full text. Chase a piece of foreshadowing planted hundreds of chapters back. These tasks are bottlenecked precisely on how much fits in one gulp, and a 256K window — roughly 150,000 to 200,000 English words of fiction by the usual conversion — swallows most standalone novels whole. That capability is real.
Continuation just isn’t on the list. It wants the right small slice each time: recent plot, the characters in scope, the planted threads. That is a selection problem, not a capacity problem, and in a reader the context assembly does the selecting. The pragmatic split: hand the summarize-and-audit work to a long-window model in one pass, and route everyday continuation by genre from the benchmark tables — per-segment model switching makes the two jobs coexist inside one book.
Prices, setup, and two dated warnings
International list prices (checked 2026-07-22): K2.6 at $0.95 per million input tokens and $4.00 per million output, 256K window; the flagship K3 at $3.00 in, $0.30 on cache hit, $15.00 out, 1M window. Prepaid — a drained balance answers with a 402. One segment, on this site’s standard assumption of 10,000 tokens in and 800 out, runs a bit over a cent on K2.6. The input side always dominates, which is why the how-much-context question above matters more to a monthly bill than the sticker price.
Setup runs through the international developer platform platform.kimi.ai: register, create an sk- key on the API Keys page, top up. In Foreverse, Settings → AI models & services → Model providers → Moonshot, paste, save, and run the capability test, a single real call that proves the key works before any book is involved. One host detail that bites: the preset entry ships with the mainland base URL (api.moonshot.cn/v1), so with an international key, change the host to api.moonshot.ai/v1 — key and host must come from the same platform, or the answer is a 401. The provider directory keeps this kind of fine print straight per vendor. Then import the book, long-press a paragraph, continue — new text lands on a branch, the original untouched.
The two dated warnings. First: the legacy moonshot-v1 series sunsets platform-wide on August 31, 2026, so a model name copied from an older tutorial has a shelf life; the sunset calendar tracks it. Second: the flagship question got its verdict — K3 swept K2.6 21:1 in our paired double-blind, with full sample pages in the 20-round file — but that verdict is about prose style. The window never entered the exam, and K3’s million tokens change nothing about this post’s conclusion.
FAQ
Can Kimi continue a novel well?
Yes, with a checkable rank. In our nine-model blind benchmark on an ornate court novel — 20 consecutive rounds per model, two independent reviewers — Kimi K2.6 placed 5th–7th, and the review note is the interesting part: the only system whose output improved as the run went on, opening in a shouty webnovel register and settling into period decorum. Its recorded weakness is a mild self-copying habit on fantasy long runs; when a passage feels like déjà vu, one reroll or a model switch breaks the loop.
Does the 256K context window help fiction writing?
Depends on the job. For one-pass whole-book work — summarizing arcs, auditing a character's appearances, drafting lorebook entries — a window that swallows a full novel is a genuine advantage. For repeated continuation it is not: in our five-tier control experiment, cramming an entire novel into context with no style instruction still produced stock similes at ten times the original author's density, and every segment re-billed the whole book as input. Continuation wants a small, relevant slice, not the shelf.
What does the Kimi API cost, and how do I wire it up?
International list prices, checked 2026-07-22: K2.6 at $0.95 per million input tokens and $4.00 per million output; the flagship K3 at $3.00 in ($0.30 on cache hit) and $15.00 out. The API is prepaid and separate from the free Kimi chatbot. Create a key on the international platform, paste it into a BYOK reader, and set the host to api.moonshot.ai/v1 — Foreverse's preset Moonshot entry ships with the mainland host, and key and host must come from the same platform. Run the capability test, then import a book and continue from a long-press.
The Kimi app is free — doesn't that cover the API?
No. They are two separate account systems. The consumer Kimi assistant on web and mobile offers free chat; the API lives on the developer platform with its own registration, prepaid balance, and per-token billing, and the two balances never mix. Wiring Kimi into a reader to continue a book rides the API. One date to note: the legacy moonshot-v1 model series sunsets platform-wide on August 31, 2026, so don't copy model names from old tutorials.
Questions or ideas? Join our Discord →