Foreverse Research · Fiction Bench
How good is Qwen3.7-Max at writing fiction?
It's the most instructive row on the board: best-in-field on sentence length, dialogue rate and stock-phrase density, yet blind-ranked only 5-6 (xuanhuan) and 8 (romance). The reason: every window carries smell-and-touch writing (“the scorched reek stung the nose”) in books that contain zero scent description. Statistically alike is not the same as reading alike.
Xuanhuan fantasy
5-6 / 9
Both packs: p303 第5 · p404 第6
A constant “scent fingerprint” — smell/touch description in every window of a source book that contains zero scent writing; consistently passable, consistently unlike.
Court romance
8 / 9
Both packs: p505 第8 · p606 第8
A unanimous #8 in both packs; the “polished sensory stream” reads even more out of place in period prose.
What the 20-round chains actually showed
Field-best structural metrics: sentence length 32.1 (gold 32.1), dialogue 18.1%, lowest simile-stock density at 1.74‰. Blind review still ranked it 5-6 on xuanhuan and a unanimous #8 on romance.
The judges' explanation was consistent and specific: its “sensory stream” — smell, touch and novel similes in every window (“the scorched reek stung the nose”) — is a dimension the source books simply do not have. Sentence-length statistics cannot measure writing what the source never writes.
This row is the direct origin of our methodology rule: structural metrics serve as regression gates only; the likeness verdict needs close-reading judges. Pick by the stats table alone and you'd crown this model #1.
Long-run failure mode
None observed
No long-run failure; its issue is a constant stylistic overlay, not degradation over rounds.
Structural fingerprint
Best structural metrics in the field (sentence length 32.1 vs gold 32, 18.1% dialogue, lowest stock-phrase density 1.74‰) — yet blind-ranked 5-6: statistics can't measure “writing what the source never writes.”
Cross-genre profile: The “polished sensory stream” clashes hardest with period prose; the field's exemplar of statistically-closest, temperamentally-furthest.
Test-condition disclosure (hosted models are moving targets)
Model under test: qwen3.7-max (released 2026-05-21)
Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.60 list price: in $2.5/M · out $7.5/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Qwen3.7-Max.
Continue your book with itHow to cite
Foreverse Research, “How good is Qwen3.7-Max at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/qwen3-7-max