GPT-5.6 vs Kimi: two personalities on the same exam
GPT-5.6 Terra and Kimi K2.6 each continued two Chinese novels for 20 consecutive rounds in our double-blind benchmarks. Terra is the best prose mimic we have measured and ran the palace-intrigue novel with zero incidents — then rewound the fantasy plot in rounds 18-20, reusing its own round-2 lines. Kimi placed seventh in fantasy with the highest metaphor density in the field, yet was the only system that improved as it went. Honest cutoff: Kimi K3 shipped July 16, 2026 and was never in this exam.

Two moments tell the story. Round 18 of the fantasy run: GPT-5.6 Terra, which had just spent seventeen rounds being the closest prose mimic both blind reviewers had ever scored, rewound the entire plot back to the starting anchor and began replaying chapter one — with sentences recycled from its own round 2. Meanwhile, round 1 of the palace-intrigue run: Kimi K2.6 opened in a shouty webnovel register that had nothing to do with the book’s ornate first-person voice. By the late rounds it was the only system of nine that reviewers flagged as clearly improving.
Short answer: for ornate, dialogue-driven historical fiction, take GPT-5.6 Terra — it held the top tier with a zero-incident run. For fast-paced fantasy, neither of these two is the champion, and each fails in its own signature way. The data comes from two double-blind benchmarks in which both models continued two Chinese novels for 20 consecutive rounds each. What follows reads them as two personalities rather than two rows in a table.
Where does each model place?
The fantasy exam is an 8.9-million-character epic; the palace exam is Empresses in the Palace, a first-person palace-intrigue classic. Nine models ran each book from the same anchor with the same prompt, 20 consecutive rounds, every round’s output appended back into the context. Outputs were anonymized under two random letter mappings and ranked by two reviewers unaware of each other; rank ranges record real reviewer disagreement.
| Exam | GPT-5.6 Terra | Kimi K2.6 |
|---|---|---|
| Fantasy epic (8.9M characters) | 2-4 (placement depends on how hard you punish the crash) | 7 (identical across both mappings) |
| Palace intrigue (Empresses in the Palace) | 1-3 (zero-incident run) | 5-7 (the only system that improved as it went) |
Full nine-model tables, reviewer quotes, and method limits live in the fantasy benchmark and the two-genre rankings page. This page only cross-examines these two.
GPT-5.6 Terra: the best mimic in the field, and the one that forgets where the story is
The ceiling first. In the fantasy run’s early and middle windows, both blind reviewers judged Terra the closest thing to the original author anyone produced. It was the only system of nine that reproduced the author’s micro-habits: tilde-marked onomatopoeia, a trademark colloquial particle substitution, the villain’s exact register. Not word-choice resemblance — typesetting-muscle resemblance.
Then rounds 18 through 20. It rewound the plot to the starting anchor, replayed chapter one, and reused lines from its own round-2 output. The texture stayed perfect; the story folded back on itself. Across all 360 rounds we have run, this failure — we call it the ending rewind — hit exactly one model, and it was the best prose mimic in the field. Surface mimicry and long-range plot coherence are separate capabilities. Terra is the sharpest evidence of that sentence we own.
On the palace novel, the same model produced the only zero-incident run in the sample: no formatting errors, no honorific slips, no continuity accidents in 20 rounds, with evidence-chain plotting (stitch patterns, aged spices, baited traps) reviewers called the closest to the original author. Rank range 1-3, no rewind. Across forty rounds, Terra’s personality is impeccable posture with occasional amnesia: most of the time it is the best-behaved student in the room, but on the fantasy exam it erased its answer sheet just before handing it in and rewrote question one.
Kimi K2.6: loudest opening, steadiest correction
Kimi’s curve runs the other way. On the palace novel it opened in a register the reviewers described as shouty webnovel fury, a full genre away from the book’s restrained first-person voice — then spent the run correcting itself, settling into palace decorum by the late rounds. Rank 5-7, with the note that it was the only system showing clear reverse improvement. Most models loosen over twenty rounds; this one tightened.
The fantasy run exposed the other side. The blind consensus recorded two things: the highest metaphor density in the field — nearly one metaphor per paragraph, against a source text that uses about one per thousand characters — and whole passages in the early-to-middle rounds copied verbatim from its own earlier output. That self-copying is the mild form of the mid-run freeze; the severe form belongs to Grok 4.5, and the full taxonomy is in our three-failure-modes study. Fantasy placement: seventh, identical under both mappings. Not controversial.
So Kimi’s personality is heavy muscle memory with late-run self-correction: it walks into the exam with a webnovel accent it can’t immediately drop, stumbles against an unfamiliar register, then uses the growing pile of corrected context to pull itself back on track. Judge it on the back half of the run, not the first three rounds.
So which one should you write with?
Pick per scene, not per brand. For ornate historical fiction and anything where diction discipline matters, Terra is the only one of the pair that reached a top tier — use it. For fast-paced fantasy, honestly, neither is the champion (that was DeepSeek V4 Flash; the receipts are on the rankings page). If you must choose between these two, Terra places higher — but past a dozen consecutive rounds, scroll back and check whether the plot has quietly rewound; if it has, regenerate that segment and the loop breaks. If you choose Kimi, give it warm-up room. Its strength arrives late.
The more durable conclusion is: don’t lock in. Rewinds and self-copying are cumulative diseases, invisible in any one-shot demo. In a reader that switches models per paragraph, these two opposite personalities become complementary inside the same book, at zero switching cost.
The Kimi you can buy today is not the one that sat this exam
The freshness boundary, stated plainly: we tested Kimi K2.6. On July 16, 2026, Moonshot AI shipped Kimi K3 — 1M-token context, native vision, and per the official pricing docs (verified July 18, 2026), $3.00 per million input tokens, $0.30 on cache hit, $15.00 per million output, flat across the whole window, currently max reasoning effort only. K3 never sat this exam. Every placement, quote, and failure mode above belongs to K2.6 and none of it transfers. Bigger specs do not mean closer prose — in our data, spec sheets and style fidelity have never pointed the same way.
Same disclosure on the other side: the GPT-5.6 family reached general availability on July 9, 2026 in three tiers — Sol, Terra, Luna. We tested Terra, which OpenAI’s announcement prices at $2.50 per million input tokens and $15 per million output. Sol and Luna were never benchmarked, so they get no rank here.
The next re-run puts K3 through the same 20 rounds on both books. This page’s table gets updated then, date at the top as the truth source. Until that happens, any claim that K3 crushes anything at fiction has no exam paper behind it.
FAQ
Is GPT-5.6 good at writing fiction?
Genre-dependent. In our double-blind benchmarks, GPT-5.6 Terra took the top tier (rank range 1-3) on a palace-intrigue classic with a zero-incident run: no formatting, honorific, or continuity errors across 20 rounds. On a fast-paced fantasy epic its early and middle rounds were judged the closest prose mimicry in the field, but in rounds 18-20 it rewound the plot to the starting anchor and reused lines from its own round 2, landing at 2-4. We only tested the Terra tier; Sol and Luna were never benchmarked.
Is Kimi good for fiction?
Depends how long a runway you give it. Kimi K2.6 placed seventh on the fantasy epic (identical across both blind mappings), with the highest metaphor density in the field and whole passages copied from its own earlier rounds. On the palace novel it ranked 5-7 but earned a note no other system got: it opened in a shouty webnovel register and settled into palace decorum as rounds went on — the only system that clearly improved with distance. Don't judge it on its first three rounds.
Will Kimi K3 write fiction better than K2.6?
We don't know, and we won't guess. K3 launched July 16, 2026 with a 1M-token context window, priced at $3 per million input tokens ($0.30 on cache hit) and $15 per million output on the international API. Bigger specs and a higher price tell you nothing about prose fidelity: in our data the fantasy champion was one of the cheapest models in the field. Every placement on this page belongs to K2.6 only. We'll re-run with K3; the date at the top of this page is the truth source.
Can I use both models inside one book?
Yes, and that is the practical hedge. The failures we observed — ending rewinds, self-copying — are cumulative: invisible in a one-shot demo, unmistakable after a dozen consecutive rounds. Switching models mid-book or regenerating one segment resets the loop before it locks in. Foreverse lets continuation switch models per paragraph, and both GPT-5.6 and Kimi plug in with your own API keys.
Questions or ideas? Join our Discord →