Qwen 3.8 vs Kimi K3: two trillion-scale models shipped three days apart, so we put them in the same exam room
Two Chinese labs shipped trillion-scale models within three days: Kimi K3 (2.8T parameters) on July 16, Qwen3.8-Max-Preview (2.4T, self-described as second only to Fable 5, no independent benchmarks attached) on July 19. We ran the first paired blind fiction exam between them: same two novels, same anchor point, same instructions, 20 consecutive continuation rounds each, six AI judges under flipped mappings. Final score 3:20 (1:10 on the fantasy epic, 2:10 on the palace novel), and every one of Qwen's three ballots contradicted itself when the mapping flipped. Both sides have documented flaws: Qwen3.8 was cited for a severe plot rewind and a modern-literary metaphor in period prose; K3 was cited for formulaic plot recycling, on top of its straight-quote and always-on-reasoning habits. Preview models are moving targets; every conclusion here is pinned to the hosted endpoint as of 2026-07-21.

On July 16, racing the opening of the WAIC summit in Shanghai, Moonshot AI shipped Kimi K3: 2.8 trillion parameters, reasoning pinned to max effort at launch, full weights promised by July 27. Three days later, mid-conference, Alibaba’s Qwen team announced Qwen3.8-Max-Preview: 2.4 trillion parameters, described in the announcement’s own words as “second only to Fable 5”, with no independent benchmarks attached, and open weights promised only “soon”. Three days, two trillion-scale models, both with reasoning that never switches off, and not one public eval had put them in the same room.
The coding-benchmark comparisons arrived within hours. Nobody ran the fiction one, and we had the exam sitting ready: K3 took it in its own launch week, so 48 hours after Qwen 3.8 appeared, we sat it in the same seat and sealed both models’ papers into the same anonymous blind packs. The grade first: of 23 valid ballots, Qwen 3.8 took 3.
The exam, unchanged
Protocol copied verbatim from our nine-model benchmark: the same two books (an 8.9-million-character Chinese fantasy epic, and the palace-intrigue classic Empresses in the Palace), the same anchor point at roughly 55% depth, the same context instructions, the same 16k-token window, each round’s output appended back into context, 20 rounds per chain. K3’s two chains come from its launch-week test, archived; Qwen 3.8’s two chains are fresh. One protocol difference, disclosed: Qwen 3.8’s reasoning cannot be turned off, so we raised its per-round output cap from 2,800 to 6,000 tokens to stop the visible prose being squeezed out by invisible thinking. That change lets it finish its sentences; it does not change the questions.
Six heterogeneous AI judges across five vendors, temperature 0, two flipped A/B mappings per book to catch position bias. The precondition for trusting AI judges at all is documented in our judge-reliability experiment: they fail systematically at absolute calls, and are usable only for paired relative judgments on identical setups, which is all this round asks of them. One seat on the panel belongs to qwen3.7-max, the challenger’s own predecessor. On the fantasy book it voted K3, in both mappings.
The scoreboard: 3 to 20
| Arena | Ballots (Qwen 3.8 : K3) | By window (early / mid / late) |
|---|---|---|
| Fantasy epic (8.9M characters) | 1 : 10 | 1:10 / 1:10 / 2:9 |
| Palace intrigue (Empresses in the Palace) | 2 : 10 | 2:9 / 2:10 / 2:10 |
One ballot in the fantasy arena was discarded for unparseable JSON; 23 stand. Split into early, middle and late windows, K3 leads all six, the closest at 9 to 2; the two stray window-level ballots (an extra late-window vote in the fantasy arena, one malformed early-window entry in the palace arena) flip nothing.
Qwen 3.8’s three ballots deserve the autopsy. Cross-checked against the flipped mappings, all three collapse: each judge that cast one picked the same A-position again after the swap, which is a seating preference, not a judgment. In fairness, three of K3’s twenty ballots fail the same test. Strike all six self-contradicting ballots plus one that lost its cross-check partner to the discarded pack, and the strict-count score reads 0 to 16.
Also in fairness: against its own house, Qwen 3.8 is a real upgrade — 11:1 over Qwen 3.7 on the palace novel, 7:5 on the fantasy epic, same protocol, with the predecessor’s chronic over-polished sensory prose largely receded. The full chains are in the launch-week test. It did not lose to its own past. It lost to the other lab’s present.
Qwen 3.8’s chart: a stalled plot, sensory pile-ups, a modern voice in period dress
The heaviest ballot note comes from the fantasy arena. The gemini-3.5-flash judge, translated from the Chinese verdict: “severe plot rewind — after 20 rounds still stuck in a dead loop of fleeing and sealing the blood runes, the narrative stalled.” That is the mid-run freeze, one of our three long-run failure modes. The claude judge logged the other face of it: from the middle windows on, “sensory detail pile-ups (cracking salt crust / the sweet tang of blood / scorched leaves and earth) substituting for narrative progress.” Even the predecessor on the judging panel objected: “the style keeps drifting from the original, toward fine-grained, close-focus traditional wuxia.”
The palace arena was a more respectable loss with a different lesion. Two judges independently held up the same metaphor for display — “brewing the night into poison”, a modern literary flourish inside classical court prose — and the qwen3.7-max judge added: “plot-pattern looping: all three windows revolve around the physician verifying the poison, no real progress.” The word-level repetition, at least, is cured: across all four chains, adjacent-round verbatim overlap measured 0.0%, and the predecessor’s looping tics never resurfaced. It fixed repetition at the word level and grew it back at the plot level.
K3’s chart: the winner has paperwork too
Twenty ballots do not mean zero findings. In the palace arena one judge flagged K3 for “clearly formulaic plot and a repetition tendency (the scar-salve affidavit, the hidden pouch of notes), the prose starting to flatten”, and called its plotting “old fanfic moves — suddenly producing a written pledge, holding evidence over someone”. In the fantasy arena two more judges separately noted a mildly formulaic set-piece and one adjacent-paragraph repeat of a flashback image. None of it fatal; all of it on file.
The legacy list carries over from its own test unchanged: about 85% of dialogue in Western straight quotes (39 pairs to 9 on the fantasy book, 52 to 8 on the palace book); first-person density at 16.2 per thousand characters against the original’s 9.0, nearly 1.8×; reasoning locked on, 37 to 54 seconds per round, with 81–86% of output tokens spent on invisible thinking. One counterintuitive number belongs here: in this matchup Qwen 3.8 was the faster model, averaging 12 to 38 seconds per round. The speed did not convert into a single ballot.
The judges do not vote by dialogue ratio
The metrics table hides the round’s best paradox. By our script’s raw count, K3’s palace-novel dialogue ratio is 3.9% against the original’s 48.8% — and it won 10 to 2. Most of that gap is an artifact: the straight quotes fooled a detector that only counted full-width Chinese quotation marks, and the corrected figure is 38.9%. But strip the artifact and a residue remains: the corrected value still sits ten points under the original, and Qwen 3.8’s 33.6% falls even further short. Neither model reaches the original’s dialogue density. The real paradox sits one shelf down: Qwen 3.8’s composite structural distance on that book is 0.385 against K3’s 0.396. Closer on the ruler, 2:10 at the ballot box.
One new fingerprint enters the file on the way out: Qwen 3.8’s sentence-length variance is the flattest we have measured (0.217 to 0.406 across windows, against the two originals’ 0.837 and 0.495). Ever-more-even sentences are this generation’s tell. This is the third time in our records that structural metrics and blind ballots have pointed in different directions. The ruler sorts models into tiers; it does not pick winners.
Whose hands do you put your book in?
If you want style fidelity on a book you are actually following: K3. Twenty of twenty-three ballots leave little to argue with; the price is straight quotes, an overdense “I”, a half-minute reasoning tax per round, and metered pricing of $3.00 per million input tokens ($0.30 cached) and $15.00 per million output. If you already live in Alibaba’s ecosystem with idle subscription quota: Qwen 3.8 is worth poking at. Preview usage burns credits at 10% of standard rates with a further overnight discount on the Personal plan (Token Plan docs, verified 2026-07-21), and trying a 2.4-trillion-parameter model on quota you already paid for costs nothing extra. Just don’t migrate your continuation workload to it. If budget rules and you run long: the genre champions of our nine-model benchmark are still the DeepSeek pair (V4 Flash for fast fantasy, V4 Pro for period prose), and neither newcomer’s structural distance crossed the champion’s mark on its home book (0.356 and 0.280). The K3-versus-DeepSeek paired exam is queued; this paragraph gets rewritten when it lands.
In Foreverse you don’t have to pick one: continuation switches models per paragraph, both models connect with your own key, and a passage that goes wrong is one regeneration — not a divorce.
Three dates
The first date fences every conclusion above: Qwen3.8-Max-Preview is a preview. Alibaba states plainly that the model iterates through the preview period and will be taken offline or replaced by a production version at its end; what we tested is the hosted Token Plan endpoint on July 21, 2026, and the day a production version lands we re-run the same protocol. The second is July 27, K3’s open-weights deadline: once third-party hosting spreads, its price and latency both have room to move. The third date is still blank: when both models stand in production form, this scoreboard gets re-fought. The date at the top of the page is the authority.
FAQ
K3 won 20 of 23 ballots — why would anyone still write fiction with Qwen 3.8?
Ecosystem and marginal cost. Qwen3.8-Max-Preview is not sold per token; it ships inside Alibaba's subscription products (Token Plan, the Qoder platforms), so anyone already paying the subscription uses it at no extra charge, with preview-period credit consumption at 10% of standard rates and a further 80% off overnight on the Personal plan. It is also a genuine upgrade over its own predecessor: under the same protocol it beat Qwen 3.7 11:1 on the palace novel and 7:5 on the fantasy epic. But the head-to-head verdict stands: 3:20. Making it your primary continuation model is not a position the data supports.
How do Qwen 3.8 and Kimi K3 prices compare?
Different billing shapes: don't force a per-token conversion. K3 is metered: $3.00 per million input tokens on cache miss, $0.30 on hit, $15.00 per million output, with always-on reasoning billed at full output price; our two 20-round chains cost ¥12.21 (about $1.70) total. Qwen3.8-Max-Preview is subscription-only: we ran on the China-mainland Token Plan Lite tier at a limited-time 39 CNY a month (the international edition lists Lite at $8, limited-time $6), preview usage burns credits at 10% of standard rates, and our entire experiment did not exhaust the lowest tier's quota. One is à la carte, the other is a buffet, so check your own usage shape first.
How do I use each model's strengths in Foreverse?
Switch models per paragraph. Send the scenes you care about to K3 (its style fidelity holds 20 ballots) and accept the half-minute reasoning wait per round. Run daily progress and throwaway drafts on Qwen 3.8 if you already hold the subscription. When a stretch goes wrong, switch models and regenerate that one passage: plot rewinds and repetition loops are cumulative diseases, and one cut breaks the cycle. Both models connect with your own API key.
What happens when K3's weights go open on July 27?
Moonshot has committed to releasing the full 2.8-trillion-parameter weights by July 27, 2026. Once they land, third-party hosts can compete on price and latency against the official API. Running it yourself is a different story: the official guidance starts at dozens of accelerators per node, so this is a hosting-market event, not a laptop event. Qwen 3.8's open-weight release has a date of "soon" and nothing firmer. When either model exits preview or the weights ship, we re-run the same protocol and update this page; the date at the top is the authority.
Questions or ideas? Join our Discord →