Foreverse Research · Fiction Bench
How good is DeepSeek V4 Flash at writing fiction?
Depends on the book. On traditional xuanhuan fantasy it ranked #1 of nine in both blind packs — judges called it “the only one that reads like a real next chapter.” On formal court romance it slid to 6-7 of 9. It's also one of the cheapest models on the board.
Xuanhuan fantasy
1 / 9
Both packs: p303 第1 · p404 第1
Won all three windows; judges called it “the only system that gets better as it writes” — opens new arcs in the late window instead of flagging. High-confidence #1.
Court romance
6-7 / 9
Both packs: p505 第7 · p606 第6
The xuanhuan champion slid to lower-mid on court romance — its plain, fast-talking strengths don't transfer to formal period prose.
What the 20-round chains actually showed
Ranked #1 in both xuanhuan blind packs (p303/p404). Judges' sketch: “the only system that reads like a real next chapter” — dialogue-driven, colloquial address, standalone onomatopoeia, fast pacing, first in nearly every window; the nine-model round added “the only system that gets better as it writes.”
On the court-romance track (D2 directive) it slid to 6-7 of 9: the same plain, fast-talking pen stops sounding right in first-person limited, etiquette-heavy period prose. The two genre crowns going to different models is the series' single most important finding.
Structural metrics and blind judgment agree on this one: sentence-length CV recovered 0.56→0.64 late, zero marker echoes in 20 rounds. It's also the cost floor of the board — $0.14/M input at list price; ten thousand characters of new prose costs about three US cents (see cost-column note).
Long-run failure mode
None observed
None of the three long-run failure modes observed across the 20-round chain; zero marker echoes.
Structural fingerprint
Sentence-length CV 0.56→0.64 (gold 0.837); dialogue-driven, standalone onomatopoeia, fast pacing; zero marker echoes in 20 rounds.
Cross-genre profile: King of xuanhuan: plain speech, fast pacing and barked forms of address are its native register; formal literary prose is not.
Test-condition disclosure (hosted models are moving targets)
Model under test: deepseek-v4-flash (released 2026-04-24)
Evaluated: 2026-07-16 · Access channel: DeepSeek official API
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.032 list price: in $0.14/M · out $0.28/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with DeepSeek V4 Flash.
Continue your book with itHow to cite
Foreverse Research, “How good is DeepSeek V4 Flash at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/deepseek-v4-flash