Foreverse Research · Fiction Bench

How good is Qwen3.8-Max-Preview at writing fiction?

A real upgrade over predecessor Qwen3.7-Max: an 11:1 crush on court romance (all three windows) with 3.7's polished-sensory-stream ailment visibly converged; on xuanhuan only a slim 7:5 — and fixing word-level repetition bought plot-level looping, with judges flagging “still stuck in an escape-and-seal loop after 20 rounds.” Against the same month's Kimi K3 it lost 3:20. Previews are moving targets; conclusions bind to the 2026-07-21 hosted build.

vs Qwen3.7 · 7:5 / 11:1Paired-duel record · not on the full rankingTested 2026-07-21

Duel scorecard (paired double-blind, 6 judges × flipped mappings)

vs Qwen3.7-Max · Same-vendor predecessor · board: 5-6 xuanhuan / #8 romance

Xuanhuan fantasy

7 : 5

Court romance

11 : 1

Tested 2026-07-21

Romance swept all three windows (11:1/11:0); xuanhuan split early 6:6 / mid 7:5 / late 7:5 — a slim win at best.

vs Kimi K3 · Same-month continuation strongman · see its duel page

Xuanhuan fantasy

1 : 10

Court romance

2 : 10

Tested 2026-07-21

Trailed in all six windows; all three of its votes contradicted themselves under mapping flip (position bias) — not one stable judgment.

Where it stands against the board

Predecessor Qwen3.7-Max is the board's textbook row for statistically-closest-yet-temperamentally-furthest: field-best structural metrics, blind-ranked only 5-6 on xuanhuan and a unanimous #8 on romance — scent writing in every window of books that contain none. 3.8 fixed half of that: on romance the sensory stream visibly converged and it swept 11:1; on xuanhuan 7:5 is barely above par.

But fixing word-level repetition bought plot-level looping: judges flagged the xuanhuan chain as “severe plot rewind — still stuck in an escape-and-seal-the-blood-rune loop after 20 rounds, narrative stalled,” and the romance chain as “all three windows orbiting the same poison-test scene, no real progress.” That's the mid-chain-collapse type from the board's taxonomy (the kimi-k2.6 / grok-4.5 family). The mechanical readings are clean (0.0% verbatim overlap between adjacent rounds); the loop lives at plot level, where verbatim detectors can't see it.

Its new signature is ever-more-even sentences: the flattest sentence-length CV in the field (0.318 across the xuanhuan chain, just 0.217 in the late window, against source baselines of 0.837/0.495). It also hosts the third recorded divergence between metrics and blind review: on xuanhuan, 3.7's statistical distance is closer (0.244 vs 0.442) yet the blind vote went to 3.8 — metrics can grade tiers, not human-likeness.

The lateral read in one line: losing 3:20 to the same month's K3 means this upgrade caught up with its own predecessor, not with the current front line.

Test-condition disclosure (hosted models are moving targets)

Model under test: qwen3.8-max-preview (released 2026-07-19)

Evaluated: 2026-07-21 · Access channel: Alibaba Token Plan (OpenAI-compatible endpoint)

Review format: paired double-blind verdicts (two flipped mappings per book against position bias), not the nine-model full ranking

Matches the K3 incremental round item by item (same two books · same 55% anchor · same D2 directive · 20-round chains · temperature 0.7); the sole protocol difference is max_tokens 2800→6000 (qwen3.8 thinks constantly — prevents reasoning tokens from squeezing out the prose).

Honest limits

Qwen3.8-Max-Preview has not entered the nine-model same-protocol full-ranking review, so the board's rank column does not apply to it — this page publishes only ballot-backed paired duels and invents no rank. When it joins the full ranking depends on the next full-board rerun.

Previews are moving targets: every conclusion binds to the 2026-07-21 hosted preview build; the release version may drift.

Its votes against 3.7 can be read alongside 3.7's board ranks, but paired votes and full-ranking positions are different review formats — they don't convert.

Pricing: the preview isn't in the models.dev list-price snapshot, so this page's cost section is honestly absent (the eval ran on Alibaba Token Plan subscription quota).

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Qwen3.8-Max-Preview.

Continue your book with it
Qwen3.8-Max-Preview for Fiction Writing — Paired Blind-Duel Record (Fiction Bench) · Foreverse · Xinmeng