DeepSeek V4 Pro went official. Two hours later, it lost a blind duel to its own delisted preview

DeepSeek V4 Pro's official release (0813) went live on the night of August 12, 2026. Within one to two hours we ran it through the same protocol as our nine-model benchmark: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its own preview predecessor — the 0716 snapshot that holds our romance crown. Palace novel: 1:11 (mapping-stable votes 0:5). Fantasy epic: 3:9 against the predecessor (stable 0:3), and 3:9 against reigning fantasy champion V4 Flash's 0716 snapshot. Diagnosis: a drift into modern literary prose — cosmic-tier fantasy rewritten as low-powered wuxia, the palace novel's voice replaced by contemporary hurt-core metaphors at 2.2× the predecessor's density, and the Empress housed in a TV-adaptation palace six times. The machinery, meanwhile, is spotless: zero verbatim loops, zero half-width quotes across 40 rounds. The finding that outranks any score: the official API hot-swaps weights under the same model ID — the champion on our leaderboard can no longer be called. Total experiment cost: about ¥5.

An embroidered palace fan wearing a champion's rosette faces a bronze mirror on a scholar's table; the mirror reflects the same fan with a fresh sprig of leaves

Late on August 12, DeepSeek's changelog announced V4 Pro's official release. The line that matters: “calling method unchanged — use deepseek-v4-pro to get the latest version.” One to two hours later we had it in the ring. The fight card was an odd one: the challenger is the freshly shipped GA build (the pricing page's version row now reads DeepSeek-V4-Pro-0813); the defending champion is its own preview — the 0716 weights that hold the romance crown on our leaderboard. Three bouts, scored separately: palace novel 1:11. Fantasy epic against the old build, 3:9. A third bout against reigning fantasy champion V4 Flash (also a 0716 snapshot), 3:9. The champion swept its title defense.

And the champion is no longer in the building. GA hot-swapped the weights under the same model ID, leaving no dated alias behind. It won the defense, then vanished from the API.

The rules: archived opponents, an unchanged protocol

The protocol is copied from the pipeline we've used since the nine-model benchmark, through K3, Opus 5 and Gemini 3.6 Flash: the same two books (an 8.9-million-character Chinese fantasy epic, and Empresses in the Palace), the same anchor at roughly 55% depth, the same 16k-token rolling window, each round's output appended back into context, 20 rounds per chain, temperature 0.7. The only variable is the model. The GA side ran on DeepSeek's official API (api.deepseek.com/v1) between 00:36 and 01:20 Beijing time on August 13 — one to two hours after the announcement, entirely off-peak.

Every opponent chain is archive footage — nothing was re-run. The palace-novel opponent chain was generated under the current context-labeling instruction, so that pairing is clean. The two fantasy opponents date from the original benchmark and ran on the older wording, which strictly speaking smuggles a prompt variable into the bout. We accepted it with evidence on file: our ablation showed DeepSeek-family models are insensitive to all three wordings (zero regressions), and the fresh GA chains showed zero continuation-point rewinds in 40 rounds. Recorded here, not buried in a footnote.

Six judges across five vendors (claude-4.6-sonnet, gemini-3.1-pro-preview, gemini-3.5-flash, kimi-k2.6, qwen3.7-max, glm-5.2), two flipped blind mappings per bout. One pool change to disclose: deepseek-v4-pro is normally a judge; this round it was a contestant, so gemini-3.5-flash took its seat per precedent. Both contestants are DeepSeek builds and no DeepSeek judge remained in the pool — no self-review, and vendor loyalty has no side to pick. All 36 verdicts returned without a single failure.

Bout one, the champion's home ground: 1 to 11

No suspense on the palace novel. Raw ballots 1:11; by window, 1:11 / 2:8 (2 ties) / 1:11. Strip out ballots that contradicted themselves under mapping flips and the stable count is 0:5 — five judges consistently for the preview, none for GA. The strongest signal of the night.

Four judges independently converged on the same diagnosis. qwen3.7-max: GA “sinks into modern hurt-core metaphor pile-ups, affected in tone, and Zhen Huan visiting the punishment chamber in person is badly out of character.” gemini-3.1-pro praised the preview for “excellently restoring the author's court voice,” then filed the counterpart: GA's “diction is over-poeticized, the plot jumps, and the original's unhurried, close-woven texture is gone.” The numbers agree: GA's simile density (the Chinese “like/as-if” family) runs 7.36 per thousand characters against the preview's 3.35 — and zero occurrences in the source's continuation window. Lines like “like a dying butterfly” or “like a fist of crushed gauze” would sit comfortably in a modern romance; in this book they read as a different author.

One soft failure sits below the mechanical radar: the punishment-chamber visit stalls from round 10 to round 19 — roughly ten rounds inside one scene, with the model at one point narrating “she repeated in a low voice, again.” Verbatim cross-round repetition peaked at just 1.8%, so string matching sees nothing; the judges flagged it independently. We're logging this “scene dwelling” shape as a candidate outside our three failure modes.

The proper-noun memory probe also failed more systematically than we've seen before. GA housed the Empress in “Jingren Palace” six times — the TV adaptation's setting. The novel's text proper uses it zero times (its single occurrence sits at 99.8% depth, inside a next-book preview appended to the ending); the book's own system is Zhaoyang Hall, 67 occurrences. No palace name for the Empress appears anywhere in the continuation window, so all six came from parametric memory: this build's memory of the story leans toward the show. K3 missed the same probe once; GA missed it six times.

Bouts two and three, fantasy: 3 to 9, twice

Against its predecessor, 3:9 with stable votes 0:3. Against reigning fantasy champion V4 Flash's 0716 snapshot, 3:9 with stable votes 1:4. Losing to the champion has a benign reading — Pro ranked 2-4 on fantasy while Flash owned the crown, a known division of labor. Losing both books to your own previous self does not.

The fantasy ailment is the same disease wearing different clothes. Three judges independently flagged a power-scale demotion: a cosmos-ruling protagonist spends the late window wading streams, sheltering from rain in the woods, lighting the dark with glow-stones, startled by night birds. gemini-3.5-flash's verdict: “sample A drifts severely off-style in the late window, written as the fine-brush scenery prose of wuxia (streams, ancient trees, glow-stones, night birds), badly split from the original's cosmic-sovereign register” — all four images verified verbatim in GA's rounds 16-19. qwen3.7-max put it in five words: “a sovereign, wading a stream.”

The minority report, on the record: claude-4.6-sonnet was the only judge stably backing GA across mappings, crediting its “clipped sentences, action woven with perception, restrained dialogue, and strong scene presence.” A taste split, not position noise. Also on the record: kimi-k2.6 swung on all three bouts — six ballots voided — making this the highest position-bias concentration we've logged, which is why the stable counts above are so much leaner than the raw ones.

The contrast with the Opus 5 round two weeks earlier is the useful one. Opus 5 also “went literary” — in a classical direction, and swept the palace novel 12:0. GA went literary in a modern direction and got killed 1:11 on the same book. Same verb, opposite vector, opposite fate.

The scorecard: statistics and judges part ways again

On composite structural distance to the source profile (lower is closer), GA scores 0.369 on the fantasy epic — better than the preview's 0.454 — and still loses 3:9. Power-scale demotion and literary drift don't show up in sentence-length statistics; the preview's repetition and low dialogue rate wreck its stats without costing it votes. Another entry for the “statistically closer ≠ reads closer” family. On the palace novel the two rulers agree for once: GA loses on both.

Metric (closer to source is better)GA 0813Preview 0716Source
Fantasy · composite distance0.3690.4540
Fantasy · sentence-length CV0.4430.5920.837
Fantasy · 的-particle density (per 100 chars)1.083.272.9
Palace · composite distance0.3860.2800
Palace · simile family (per 1k chars)7.363.350 (window)
Palace · first-person density (per 1k chars)20.711.09.0

GA's three new fingerprints are all in that table. Sentence-length variance dropped well below the preview's on both books — ever-more-uniform sentences, the first AI tell we measure, moving the wrong way. Its 的-particle density runs at 37% of the source's (the preview overshot at 3.27; GA over-corrected into particle-avoidant polish). And on the palace novel it writes “I” at 2.3× the source's rate — 20.7 per thousand characters, breaking the excess record K3 set at 16.2.

The champion's wounds, also on the record

Paired blind review cuts both ways. The preview won and still took hits: glm-5.2 flagged its late window for describing the demon lord's rueful smile and surrender gesture twice over — a rewind-flavored repeat, verified at rounds 17 and 18. qwen3.7-max caught it inventing vocabulary: “forbidden fruit” bursts through five rounds of its chain and zero occurrences of the entire source.

Our repetition checker, which didn't exist when the benchmark archive was built, also dug up old business on both champion chains: Flash's fantasy chain opens round 13 by swallowing all of round 12 verbatim — a momentary 72% cross-round overlap — then continues with new prose; the preview Pro chain does the same at round 4 (53%). Neither round sits inside a judged window, so no historical ballot moves. “Swallow the previous round whole, then keep writing” is a different animal from a verbatim death loop; one specimen each, logged as a fourth mechanical shape, no conclusions drawn.

One judge error became methodology: two judges accused Flash of “introducing a brand new character” late in its chain. We grepped the book — that character appears 27 times in the source. Judges see a 6,000-character style anchor, not the novel; they misread normal cast management as hallucination. From now on, any “new character/setting error” charge gets verified against the full text before we cite it.

The real headline isn't on the scoreboard: same ID, different model

Three sentences carry everything worth remembering. DeepSeek's official API ships new versions without changing the model ID — the string deepseek-v4-pro served preview weights on July 16 and serves the 0813 build today, with no dated alias in the model list. The two builds write like different authors under blind review (0:5 stable on the palace novel). So the assumption “I pinned the model ID, therefore behavior is reproducible” does not survive a silent hot-swap: pinning the ID does not pin the weights.

If you bring your own key to the official endpoint — which is exactly what a BYOK model picker does — this is not a hypothetical: the deepseek-v4-pro that suited your book last week may have changed voices this week, and nothing in your settings will say so. It cuts against our own leaderboard just as hard. All three DeepSeek placements on the board — Flash's fantasy crown, Pro's romance crown, Pro's fantasy 2-4 — were measured on preview-era weights, hot-swapped under the same IDs on July 31 and August 13 respectively. Both crowns are now historical snapshots. We tested Flash's official build separately in the 0731 write-up; and in a neat twist of timing, Grok 4.6, tested the same day, went the opposite direction — sweeping its own predecessor 12:0 and 12:0. August 13 supplied both halves of “a new version is not automatically better at fiction.”

The gate receipts

The bill: 40 GA rounds, zero failures, zero retries; ¥2.11 at list price, ¥2.30 as metered, and about ¥5 for the whole experiment including the judge round — all off-peak. Official pricing (pricing page, verified 2026-08-13): off-peak ¥3 per million input tokens on miss, ¥0.025 on hit, ¥6 per million output, doubled during weekday peak hours; 1M context, 384K max output, thinking on by default. Output tokens are 82-87% invisible reasoning; visible prose runs 250-280 characters per round at 25.8-35.1 seconds each. The larger repricing announced August 6 hasn't landed — the day it does, this section is out of date.

Boundaries, copied from the lab notes: one prompt framework; two books, both Chinese; one chain per book per model, single anchor. The fantasy opponents carry the old-directive condition (ablation-hedged, still a variable). GA behavior was sampled about two hours after release and may drift. The palace stall, the TV-palace slips and the power-scale demotion are all single-chain observations — recurrence rates can't be estimated from n=1. And with only 12 stable ballots surviving from 36 raw ones, every conclusion here leans on the tripod of stable votes, independent judge convergence, and mechanical verification.

Whether to follow the version bump is not a question to outsource: same opening, three continuations from the GA build and three from whatever you use now, read shuffled, flip the cards after. In a reader that switches models per paragraph, the test costs a few cents. And if you'd rather summon the old champion that went 11:1 in its own defense — it is no longer taking challengers.

FAQ

Is DeepSeek V4 Pro's official release good for creative writing?

The ballots say no — relative to its own preview. Under the same protocol as our nine-model benchmark, the GA build (0813) and the preview predecessor (0716 snapshot) each continued two Chinese novels for 20 consecutive rounds, judged by paired double-blind review. The palace-intrigue novel went 1:11 against GA with mapping-stable votes 0:5; the fantasy epic went 3:9 (stable 0:3); a third bout against reigning fantasy champion V4 Flash's 0716 snapshot also went 3:9. Judges independently diagnosed a drift into modern literary prose. Agent work and general benchmarks weren't tested here — but for continuing fiction, the official build is a regression.

Can I still use the romance champion from the leaderboard?

No. Our leaderboard's romance crown belongs to the V4 Pro weights served by the official API on July 16, 2026 — the preview era. On August 12 DeepSeek shipped GA under the same model ID, deepseek-v4-pro, hot-swapping the weights with no dated legacy alias in the model list. Whatever BYOK client you use, selecting deepseek-v4-pro today calls the new build — the one that just lost 1:11 to the old one in blind review. Pinning a model ID does not pin the weights.

Which model should I pick for continuing a novel right now?

By genre. For ornate court romance, the old champion is unreachable; the callable top tier is Kimi K3 (top two of both genres in our July 28 twelve-system re-ranking) and GPT-5.6 Terra, with Claude Opus 5 penciled in. For fast-paced fantasy, note that champion V4 Flash's rank was also measured on a 0716 preview snapshot and the ID was hot-swapped on July 31 — we tested that build separately (peak mechanical repetition fell 72% to 1.2%, style distance grew). Flash remains the budget pick; K3 is the fidelity pick. The evergreen safeguard: same opening, three continuations per candidate, read them shuffled, then flip the cards.

What does DeepSeek V4 Pro cost, and what did this experiment cost?

Official pricing page, verified 2026-08-13: off-peak, ¥3 per million input tokens on cache miss, ¥0.025 on hit, ¥6 per million output; weekday peak hours (9:00-12:00, 14:00-18:00) double everything. Context is 1M tokens, max output 384K, thinking on by default. The larger repricing DeepSeek announced on August 6 has not landed yet — when it does, redo this math. Our bill: the two 20-round GA chains cost ¥2.11 at list, ¥2.30 as metered, and the whole experiment including the judge round stayed within about ¥5, all run off-peak.

Did the GA build win anything?

Yes, and it deserves the record: engineering-wise it is nearly flawless. Forty rounds with zero failed requests and zero retries; cross-round verbatim repetition peaked at 0% on the fantasy chain and 1.8% on the palace chain; full-width quotation marks used correctly throughout — the half-width-quote disease that dogs K3 and GLM-5.2 is entirely absent. The prose itself drew compliments even from judges voting against it (one quoted its line about a blade-hum slicing through blood-light as good writing), and claude-4.6-sonnet backed it as the sole mapping-stable dissenter. It lost on likeness, not on craft — writing well and writing like your book are different skills.

Questions or ideas? Join our Discord →