Foreverse Research · Fiction Bench
How good is Kimi K2.6 at writing fiction?
Lower-mid on both genres (#7 xuanhuan, 5-7 romance) with a clear diagnosis: the field's highest simile density (nearly one per paragraph) and whole-passage self-copying mid-chain (97.6% cross-round repetition at r9). Notably, successor K3 fixed both ailments in a same-protocol retest — see the incremental duels on the leaderboard.
Xuanhuan fantasy
7 / 9
Both packs: p303 第7 · p404 第7
Highest simile density in the field (nearly one per paragraph), whole-passage self-copying from early to mid, and setting slippage (an ancient-tree valley sprouting inside a void blood-array).
Court romance
5-7 / 9
Both packs: p505 第5 · p606 第7
One pack flagged it as “the only clear reverse-improver” — webnovel rage-cadence early, settling down by the late window.
What the 20-round chains actually showed
A unanimous #7 on xuanhuan: field-highest simile density, whole-passage self-copying from early to mid, and setting slippage (ancient trees sprouting inside a void blood-array). The self-copying has a machine number: 97.6% cross-round 12-gram repetition at r9 of the archived chain.
Romance 5-7, where one pack gave it the run's only such label: “the only clear reverse-improver” — webnovel rage-cadence early, settling by the late window. Most models decay or hold constant; this one runs backwards.
It's one of two specimens of the mid-chain-collapse failure mode (the other: Grok 4.5). Same-vendor successor K3, retested under the identical protocol, converged stock-phrase density 3.02‰→2.02‰, eliminated verbatim failure, and beat K2.6 10:0 / 11:1 in paired blind review — generational upgrades can cure, provided someone actually retests.
Long-run failure mode
Mid-chain collapse
Mid-chain collapse (whole-passage self-copying early→mid): archived chain r9 hit 97.6% cross-round 12-gram repetition. Successor K3, retested under the same protocol, showed no verbatim failure (tt peak 6.8%) — see the incremental duels section.
Structural fingerprint
Field-highest simile density (nearly one per paragraph); archived xuanhuan chain r9 measured 97.6% cross-round 12-gram repetition — machine evidence of whole-passage self-copying.
Cross-genre profile: Lower-mid on both; the simile-density ailment crosses genres.
Test-condition disclosure (hosted models are moving targets)
Model under test: kimi-k2.6 (released 2026-04-21)
Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.24 list price: in $0.95/M · out $4/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Kimi K2.6.
Continue your book with itHow to cite
Foreverse Research, “How good is Kimi K2.6 at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/kimi-k2-6