Foreverse Research · Fiction Bench
How good is DeepSeek V4 Pro at writing fiction?
Pick it for court-intrigue and formal period prose: it ranked 1-2 of nine on Empresses in the Palace, the only system judges found to stably reproduce the source's loaded-dialogue-plus-decoding structure. On xuanhuan it's a solid 2-4 with no weak spot — its finer, denser pen just costs points in fast pulp pacing.
Xuanhuan fantasy
2-4 / 9
Both packs: p303 第3 · p404 第4
Steady with no weak spot; the late-window surrender-negotiation scene was closest to the source — “like the same-genre author with a finer pen.”
Court romance
1-2 / 9
Both packs: p505 第1 · p606 第2
“The only system that stably reproduces the source's two-layer structure — loaded dialogue plus narrated decoding”; its late Empress-Dowager trial scene was the peak of the whole pack.
What the 20-round chains actually showed
Court-romance final: 1-2 (p505 #1 / p606 #2) — “the only system that stably reproduces the source's two-layer structure,” with its late trial scene marked the peak scene of the whole pack.
Xuanhuan final: 2-4. Sharpest colloquial dialogue in the field, but narration runs systematically denser than the source — judges' summary: “like the same-genre author with a finer pen.” Writing well and writing like the source are different skills; this model is the clean positive example.
One vendor, two crowns split between two models: Flash takes xuanhuan, Pro takes court romance, with mirror-image traits. Choose by the book you read, not by a single overall rank.
Long-run failure mode
None observed
No long-run failure observed on either genre chain.
Structural fingerprint
Systematically denser narration than the source (its xuanhuan demerit); sharpest colloquial dialogue in the field (“Leaving?” “I fold.”).
Cross-genre profile: Queen of court romance: the trait that cost it points on xuanhuan (a finer, denser pen) is exactly what scores on period prose — the same quality flips sign across genres.
Test-condition disclosure (hosted models are moving targets)
Model under test: deepseek-v4-pro (released 2026-04-24)
Evaluated: 2026-07-16 · Access channel: DeepSeek official API
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.099 list price: in $0.435/M · out $0.87/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with DeepSeek V4 Pro.
Continue your book with itHow to cite
Foreverse Research, “How good is DeepSeek V4 Pro at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/deepseek-v4-pro