Foreverse Research · Fiction Bench

How good is Qwen3.7-Max at writing fiction?

It's the most instructive row on the board: best-in-field on sentence length, dialogue rate and stock-phrase density, yet blind-ranked only 5-6 (xuanhuan) and 8 (romance). The reason: every window carries smell-and-touch writing (“the scorched reek stung the nose”) in books that contain zero scent description. Statistically alike is not the same as reading alike.

Xuanhuan fantasy · Rank 5-6 / 9Court romance · Rank 8 / 9None observed$0.60 / 10k

Xuanhuan fantasy

5-6 / 9

Both packs: p303 第5 · p404 第6

A constant “scent fingerprint” — smell/touch description in every window of a source book that contains zero scent writing; consistently passable, consistently unlike.

Court romance

8 / 9

Both packs: p505 第8 · p606 第8

A unanimous #8 in both packs; the “polished sensory stream” reads even more out of place in period prose.

What the 20-round chains actually showed

Field-best structural metrics: sentence length 32.1 (gold 32.1), dialogue 18.1%, lowest simile-stock density at 1.74‰. Blind review still ranked it 5-6 on xuanhuan and a unanimous #8 on romance.

The judges' explanation was consistent and specific: its “sensory stream” — smell, touch and novel similes in every window (“the scorched reek stung the nose”) — is a dimension the source books simply do not have. Sentence-length statistics cannot measure writing what the source never writes.

This row is the direct origin of our methodology rule: structural metrics serve as regression gates only; the likeness verdict needs close-reading judges. Pick by the stats table alone and you'd crown this model #1.

Long-run failure mode

None observed

No long-run failure; its issue is a constant stylistic overlay, not degradation over rounds.

Structural fingerprint

Best structural metrics in the field (sentence length 32.1 vs gold 32, 18.1% dialogue, lowest stock-phrase density 1.74‰) — yet blind-ranked 5-6: statistics can't measure “writing what the source never writes.”

Cross-genre profile: The “polished sensory stream” clashes hardest with period prose; the field's exemplar of statistically-closest, temperamentally-furthest.

Test-condition disclosure (hosted models are moving targets)

Model under test: qwen3.7-max (released 2026-05-21)

Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.60 list price: in $2.5/M · out $7.5/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Qwen3.7-Max.

Continue your book with it
Qwen3.7-Max for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng