Foreverse Research · Fiction Bench
How good is DeepSeek V4 Pro 0813 at writing fiction?
In same-protocol paired blind review, the GA build (0813) lost both books to its own delisted preview: 3:9 on fantasy (stable 0:3) and 1:11 on romance (stable 0:5, the round's strongest signal); a bout against fantasy champion Flash's 0716 snapshot also went 3:9. The diagnosis is a drift into modern literary prose — cosmic-tier fantasy demoted to low-powered wuxia, the palace voice replaced by hurt-core metaphor — while the machinery is among the cleanest measured (12-gram peaks 0%/1.8%, zero half-width quotes across 40 rounds). The systemic finding: the official API hot-swaps weights under the same ID — the romance champion on this board can no longer be called.
Duel scorecard (paired double-blind, 6 judges × flipped mappings)
vs DeepSeek V4 Pro 0716 preview · Its own predecessor · board: romance champion 1-2 / fantasy 2-4 (hot-swapped under the same ID, no dated legacy alias)
Xuanhuan fantasy
3 : 9 (stable 0:3)
Court romance
1 : 11 (stable 0:5)
Tested 2026-08-13
Romance windows 1:11 / 2:8 (2 ties) / 1:11, five judges mapping-stable for the preview and none for GA; fantasy stable 0:3 (six swing ballots voided). Four judges independently named GA's modern-literary drift: hurt-core metaphor pile-ups, over-poeticized diction, and a ~10-round dwell in the punishment-chamber scene.
vs DeepSeek V4 Flash (0716 snapshot) · Reigning fantasy champion · board #1 (also a preview-snapshot rank; hot-swapped under the same ID on 7/31)
Xuanhuan fantasy
3 : 9 (stable 1:4)
Court romance
— (not contested)
Tested 2026-08-13
Losing to the reigning champion has a benign division-of-labor reading (the preview Pro ranked 2-4 on fantasy anyway); losing both books to your own predecessor does not. claude-4.6-sonnet was the sole mapping-stable judge for GA (crediting clipped sentences and restraint) — a taste split, on the record.
Where it stands against the board
This page's headline is not a score but a systemic finding: DeepSeek's official API ships new versions without changing the model ID — the string deepseek-v4-pro served preview weights on 7/16 and serves the 0813 build from 8/13, with no dated legacy alias in the model list. The two builds write like different authors under blind review (0:5 stable on romance), so “pinned model ID = reproducible behavior” does not survive a silent hot swap. For BYOK users this is a direct pain point: the deepseek-v4-pro that suited your book last week may have changed voices this week, with nothing in your settings saying so. All three DeepSeek placements on the main board (Flash's fantasy crown, Pro's romance crown, Pro's fantasy 2-4) were measured on preview-era weights, hot-swapped under the same IDs on 7/31 and 8/13 — the champions on the board can no longer be called.
The pathology of the style regression: one disease, two costumes. On fantasy, three judges independently flagged a power-scale demotion — a cosmos-ruling protagonist spending the late window wading streams, sheltering from rain, lighting the dark with glow-stones, startled by night birds (verified verbatim in r16-19). On romance, four judges independently flagged modern hurt-core prose — simile-family density at 7.36‰, 2.2× the predecessor's, against zero occurrences in the source window. The proper-noun memory probe failed harder than any prior round: the Empress housed in the TV adaptation's “Jingren Palace” six times (zero in the novel proper, whose own system is Zhaoyang Hall ×67), all six from parametric memory. Add a ~10-round dwell in the punishment-chamber scene (r10-19) — invisible to 12-gram detection at 1.8%, caught independently by judges — logged as a “scene dwelling” soft-failure candidate outside the three-mode taxonomy. The Opus 5 contrast is the instructive one: both “went literary”, but Opus 5's classical direction swept romance 12:0 while GA's modern direction lost it 1:11.
The machinery is among the cleanest on record: 40 rounds, zero failed requests, zero retries; 12-gram peaks of 0.0% (all 20 fantasy rounds at zero) and 1.8%; zero half-width-quote artifacts. “Statistically closer ≠ reads closer” recurs on fantasy: GA's composite distance of 0.369 beats the predecessor's 0.454, and it still lost 3:9 — power demotion and literary drift don't show in sentence statistics. The judges also left favorable receipts: claude stably backed GA (“clipped sentences, restraint, scene presence”), and glm quoted its blade-hum line as good writing — “writes well, just not like the book” is GA's portrait across all three bouts.
Three new fingerprints: sentence-length CV of 0.443/0.378, well below the predecessor's 0.592/0.454 — ever-more-uniform sentences, the first AI tell, moving the wrong way; fantasy 的-particle density of 1.08 per 100 characters, 37% of gold (the predecessor overshot at 3.27; GA over-corrected into particle-avoidant polish); romance first-person density of 20.7‰, 2.3× gold, breaking K3's 16.2‰ record. This round's repetition checker also surfaced a new observation on both champion archive chains — “swallow-and-continue” restatement (Flash r13 at a momentary 72%, preview Pro r4 at 53%, neither inside a judged window, no historical ballots affected) — a different animal from verbatim death loops, n=1 each, recorded without conclusions.
Test-condition disclosure (hosted models are moving targets)
Model under test: deepseek-v4-pro (released 2026-08-12)
Evaluated: 2026-08-13 · Access channel: DeepSeek official API (same-ID hot swap; tested 1-2h after GA)
Review format: paired double-blind verdicts (two flipped mappings per book against position bias), not the nine-model full ranking
Same yardstick as the K3 round and every historical DeepSeek chain (same two books · same 55% anchor · same D2 directive · 20-round chains · temperature 0.7 · max_tokens=2800); the romance opponent chain shares the current directive (clean pairing), the two fantasy opponents are old-directive archive chains (ablation-hedged); judge deepseek-v4-pro recused as the contestant, gemini-3.5-flash substituting. The weights behind this model ID were hot-swapped on 8/13: the build under test here is the 0813 GA, while the board's romance champion is the 0716 preview snapshot — the two can no longer both be called.
Honest limits
DeepSeek V4 Pro 0813 has not entered the nine-model same-protocol full-ranking review, so the board's rank column does not apply to it — this page publishes only ballot-backed paired duels and invents no rank. When it joins the full ranking depends on the next full-board rerun.
Paired blind duels ran only against its own two preview snapshots (0716 Pro / 0716 Flash), never against other vendors in the same room; not comparable to main-board ranks (different review formats). Position bias ran high this round — 12 stable ballots out of 36 raw, with kimi-k2.6 swinging on all three bouts (six voided) — so conclusions rest on the tripod of stable votes, independent judge convergence, and mechanical verification.
The two fantasy opponents are old-directive archive chains (hedged by the DeepSeek-family ablation's zero regressions, still strictly a mixed-in variable); the romance opponent chain shares the current wording — a clean pairing.
GA behavior was sampled about 1-2 hours after the announcement (off-peak, early 2026-08-13) and may drift; the TV-palace slips, the scene dwell and the power demotion are single-chain observations — recurrence can't be estimated from n=1.
Pricing: the same-ID hot swap means list-price snapshots can't be tied to a weight version, and the larger repricing announced 8/6 hasn't landed — this page's cost section is honestly absent. Official pricing page (live 2026-08-13): off-peak ¥3/M input on miss, ¥0.025/M on hit, ¥6/M output, doubled in weekday peak hours; the two chains metered ¥2.30, the whole experiment ≈¥5.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with DeepSeek V4 Pro 0813.
Continue your book with itHow to cite
Foreverse Research, “How good is DeepSeek V4 Pro 0813 at writing fiction (Fiction Bench incremental duels),” 2026-07. https://foreverse.app/research/fiction-bench/deepseek-v4-pro-0813