DeepSeek V4 Flash official API is live: fewer fiction loops, but high and max can still spend 32K on reasoning
Release-day benchmark of DeepSeek V4 Flash official API public testing: Chat and Responses streaming, a two-step tool agent, a matched 20-round fiction run, four reasoning levels, roleplay, and four character cards. The new build cut worst cross-round verbatim overlap from 72.0% to 1.2%, but moved farther from the source style. At an 8K output cap, high spent 8,191 tokens entirely on reasoning and produced no prose. At 32K, high and max both completed 12 fiction rounds, yet length and continuity stayed weak—and max still burned 32,767 reasoning tokens on a complex card with zero visible output.

The most useful result in this benchmark is an empty answer. The request returned HTTP 200 after nearly two minutes, with output_tokens=8191, reasoning_tokens=8191, and zero visible text. The model had worked; it had simply spent the entire output allowance where the reader could not see it.
On July 31, DeepSeek announced in its official changelogthat DeepSeek V4 Flash had entered official API public testing. The stable request ID remainsdeepseek-v4-flash; the date stamp in the changelog is a hosted version label, not a new model name. We first checked the official model list, then ran Chat, Responses, a tool agent, long fiction, roleplay, and card authoring directly against DeepSeek's official endpoints. Every case was sent once, with no automatic retry.
The exam: consecutive continuation, not a one-line prompt
The first layer covered Chat and Responses SSE plus a stateless two-request Responses tool loop. The second reused our existing M69 protocol: the same Chinese novel, the same anchor, the same 21,500-character rolling window, and 20 consecutive rounds, matched against an archived pre-0731 preview chain. The third used only original material: 12-round chains at none, low, high, and max, then two original roleplay cards and one four-genre card brief.
Each original fiction round asked for 450-700 Chinese characters, one concrete new action or fact, and no recap. That constraint matters. “The endpoint returned text” is availability; it is not instruction following or story quality.
Twenty matched rounds: overlap falls from 72.0% to 1.2%, style fit gets worse
| Metric | Pre-0731 preview chain | 0731 release chain |
|---|---|---|
| Completed rounds | 20/20 | 20/20 |
| Visible characters | 6,514 | 8,761 |
| Worst cross-round 12-gram overlap | 72.0% | 1.2% |
| Structural distance from source (lower is closer) | 0.356 | 0.550 |
| Flipped-mapping blind review | 1 vote | 1 vote |
The improvement is concrete. The preview's round 14 overlapped prior output by 72%; the 0731 chain peaked at 1.2% and all 20 rounds were unique. The trade-off is equally concrete: the new chain wrote more, included two rounds over 1,000 characters, and moved farther from the source's structural profile. One blind reviewer preferred the preview's style and factual stability; the other preferred 0731's forward motion and lack of rewinding.
The defensible claim is not “an X% upgrade.” It is that 0731 exchanges a major repetition risk for stronger autonomous continuation—and autonomous continuation sometimes means inventing motives, rules, and longer scenes the source never supplied.
Four levels at 8K: none and low finish; high and max can die before prose
| Effort | 8K long-fiction result | Time to prose | Reasoning and cost |
|---|---|---|---|
| none | 12/12; 0/12 within 450-700 chars | 0.75s median | 0 reasoning; $0.00464 |
| low | 12/12; 0/12 within 450-700 chars | 7.06s median | 6,179 reasoning; $0.00745 |
| high | failed on round 1; zero prose | cut off after 109.5s | 8,191 reasoning; $0.00252 |
| max | three rounds, then zero prose on round 4 | no prose on round 4 | 8,189 reasoning on the failed round |
None and low had a quality problem of their own: both ignored the requested length in every round. Median visible output was 921 characters at none and 1,303 at low. Low was slower and more expensive without a visible gain on this chain.
At 32K, long fiction survives; quality does not turn green
We changed one variable, raising the cap to 32,768, and started fresh 12-round high and max chains. Both completed 12/12 with visible prose, so 32K clearly fixes a practical engineering bottleneck. It did not fix obedience or continuity.
| 32K chain | Completion and length | Reasoning | Latency and cost |
|---|---|---|---|
| high | 12/12; only 3/12 within target | 84,322; per-round max 11,642 | 89.5s median to prose; $0.02752 |
| max | 12/12; only 2/12 within target | 76,972; per-round max 13,232 | 43.5s median to prose; $0.02558 |
In this sample high used more reasoning, took longer, and cost more than max. That does not prove high is generally heavier; it proves that the label is not a monotonic token, latency, or quality control. Max advanced the central mystery more clearly, but was more bloated: 55 uses of the Chinese equivalent of “like/as if” across 12 rounds, versus 23 for high, plus a repeated puzzle machine in which every perfectly fitting object released another clue.
High was terser but contradicted itself more often: time moved backward and then forward; 17 wall marks became seven; one copper piece sat on the workbench and remained inside a notebook at the same time. Max invented its own hard facts too, including a surgical implant reinterpreted as the sister's tuned pickup needle and an unsupported exact date. Neither chain copied sentences mechanically. The failure was semantic and structural repetition, not paste-level duplication.
We then built an anonymous A/B pack from rounds 1-2, 6-7, and 11-12. Two independent reviewers, neither given the effort mapping, both gave B a narrow win at 0.77 and 0.76 confidence; B was max after unblinding. Both credited max with steadier character reactions, causal links, and prose restraint, while also calling high's late reveal—that the ship may never have sailed and that Cen Ye was aboard—the more consequential plot move. This is a small quality win, not enough to make max a default, especially once its zero-output card result is included.
A complex card brief found the edge of 32K
High completed all four genre cards at 32K. All four had the required sections, opening length, and dialogue examples; banned-phrase hits were zero and cross-card 15-gram overlap was 0.08%. But it consumed 30,057 output tokens—27,099 of them reasoning—and did not begin visible text for 236 seconds. The urban and historical cards were the strongest; the rules-horror card contradicted its own floor arithmetic, and the game card slipped out of its established narrative level.
Max is the more important result: all 32,767 tokens were reasoning and visible output was zero, after 293.9 seconds and $0.00918. An earlier independent max run had completed the same four-card shape within 16,617 output tokens. One pass and one failure make the lesson sharper: on a model with a verified large output window, 32K reduces risk; it is not a capability guarantee or a value to send blindly to small-window models.
Roleplay: perfect length, imperfect facts
High at 4K completed eight of eight turns across two original roleplay cards, and all eight landed within the 180-350-character target. Manual reading still found mind-reading of the user, a permission quote invented to resolve an ethical challenge, and a record of an umbrella loan that never happened. Passing a length gate is not passing a roleplay-quality gate.
Responses tools and caching: usable, with a required-mode boundary
Chat and Responses streaming both worked. With tool_choice=auto, the stateless two-step Responses agent returned 200/200, produced the correct final answer, and the second request hit 512 cached tokens. Withthinking=max + tool_choice=required, the API returned 400:Thinking mode does not support this tool_choice. “Supports tools” does not mean every OpenAI-compatible parameter combination is supported.
Cost, cache, and the setting we would ship
The first release-day suite cost about $0.07994 at the prices shown on DeepSeek'sofficial pricing page. The additional 34-request 32K quality batch cost $0.07195. Responses exposed cached input tokens, but every content request leftcache_write_tokens null, so the data proves cache reads, not observable cache writes. Reasoning is already included in output-token billing and must not be charged twice.
For routine fiction continuation, none or low remains the practical default. Raise effort for a specific planning-heavy task. Give high and max at least the 32K envelope, but still detect incomplete/length + zero visible output. A true “thinking off” switch must explicitly send Responses none or Chat thinking=disabled; omitting the field means following the model default. Keep billing pre-authorization, but do not force reasoning and prose to fight inside a legacy global 8K cap.
The sample boundary is real: one book and anchor for the 20-round comparison, and one 12-round chain per 32K effort. This cannot establish a universal win rate. It can establish several engineering facts: 0731 loops less mechanically, effort is not monotonic in actual reasoning, 32K rescues most high-effort long fiction, and it still cannot rescue every complex max task.
For a broader failure taxonomy, see our long-fiction failure study; for a practical model-switching workflow, see continuing fiction with DeepSeek.
FAQ
Is DeepSeek V4 Flash official better at fiction than the preview?
Not across the board. In the matched 20-round run, worst cross-round 12-gram overlap fell from 72.0% on the preview chain to 1.2% on the official build, so the mechanical looping problem improved dramatically. But structural distance from the source moved from 0.356 to 0.550, and length and factual discipline did not improve with it. Two flipped-mapping blind reviews split 1-1: one preferred the preview's style and stability, the other preferred the official build's forward motion and lower repetition.
Why did DeepSeek use 8,191 reasoning tokens and return no answer?
In the Responses API, max_output_tokens is shared by hidden reasoning and visible output. With a cap of 8,192, the high run reported 8,191 output tokens, all 8,191 of them reasoning, then ended incomplete/length before visible prose began. This is an output-budget exhaustion, not an input-context or network failure.
Does setting max_output_tokens to 32,768 solve the problem?
It helps substantially but does not guarantee completion. At 32K, both the high and max long-fiction chains completed all 12 rounds with visible prose. In an independent four-card run, however, max still spent 32,767 tokens entirely on reasoning and returned no card text. Treat 32K as a practical default only for models verified to have a large enough context and output window—not as a promise that max cannot exhaust it.
Do tool calls work in DeepSeek V4 Flash official?
Yes with a specific boundary. A stateless two-request Responses tool loop completed 200/200 with tool_choice=auto, and the second request hit 512 cached input tokens. The combination thinking=max plus tool_choice=required returned HTTP 400 with 'Thinking mode does not support this tool_choice.' An adapter should not translate every mandatory-tool intent into required.
Which reasoning effort should I use for fiction writing?
For routine continuation, none or low were more practical in this sample. Both completed 12/12 under 8K; none was faster and cheaper. High and max did not produce a stable quality upgrade even at 32K, while greatly increasing time to first prose and still introducing factual contradictions. Raise effort for a specific planning-heavy task rather than making high or max the global default.
Questions or ideas? Join our Discord →