Field notes

Notes from the reading tavern

We are building a pocket reader and a pocket tavern. These are the potholes, trade-offs, and field tests we hit along the way — written by the people doing the work.

An embroidered palace fan wearing a champion's rosette faces a bronze mirror on a scholar's table; the mirror reflects the same fan with a fresh sprig of leaves
Model evalsAug 13, 20269 min

DeepSeek V4 Pro went official. Two hours later, it lost a blind duel to its own delisted preview

DeepSeek V4 Pro's official release (0813) went live on the night of August 12, 2026. Within one to two hours we ran it through the same protocol as our nine-model benchmark: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its own preview predecessor — the 0716 snapshot that holds our romance crown. Palace novel: 1:11 (mapping-stable votes 0:5). Fantasy epic: 3:9 against the predecessor (stable 0:3), and 3:9 against reigning fantasy champion V4 Flash's 0716 snapshot. Diagnosis: a drift into modern literary prose — cosmic-tier fantasy rewritten as low-powered wuxia, the palace novel's voice replaced by contemporary hurt-core metaphors at 2.2× the predecessor's density, and the Empress housed in a TV-adaptation palace six times. The machinery, meanwhile, is spotless: zero verbatim loops, zero half-width quotes across 40 rounds. The finding that outranks any score: the official API hot-swaps weights under the same model ID — the champion on our leaderboard can no longer be called. Total experiment cost: about ¥5.

Read the post →
Ink-and-cinnabar illustration of a phone showing a chat app behind a half-raised shutter, with a browser window off to the side holding the key-shaped toggle that raises it
JanitorAIAug 12, 20267 min

The JanitorAI app is not broken. It boots in restricted mode, and the switch is not where you think.

JanitorAI's official mobile app (in beta since early 2026) starts every install in a restricted safe mode: high-score characters blocked, definitions hidden, explicit terms filtered. The unlock toggle deliberately lives on the website, not in the app. The exact steps, the sync failure modes, what stays filtered even after unlocking — and why cards-as-files sidestep the whole class of problem.

Read the post →
An old-paper illustration of two writing desks divided by a river of ink, with a neat manuscript on one desk and loose continuation pages on the other
Model benchmarkJul 31, 20269 min

DeepSeek V4 Flash official API is live: fewer fiction loops, but high and max can still spend 32K on reasoning

Release-day benchmark of DeepSeek V4 Flash official API public testing: Chat and Responses streaming, a two-step tool agent, a matched 20-round fiction run, four reasoning levels, roleplay, and four character cards. The new build cut worst cross-round verbatim overlap from 72.0% to 1.2%, but moved farther from the source style. At an 8K output cap, high spent 8,191 tokens entirely on reasoning and produced no prose. At 32K, high and max both completed 12 fiction rounds, yet length and continuity stayed weak—and max still burned 32,767 reasoning tokens on a complex card with zero visible output.

Read the post →
Ink-line illustration of a small figure on a ladder pulling a glowing book from a tall shelf, with pages and index cards drifting toward a desk below
Model benchmarkJul 31, 20268 min

On DeepSeek V4 Flash official API day, we let the in-app Agent choose and rewrite scenes inside a 385-chapter novel

On the day DeepSeek V4 Flash entered official API public testing, we re-ran the same 385-chapter urban novel and user prompt inside the Foreverse App Agent: read relevant chapters, judge which scenes deserve expansion, then rewrite. The book survived, write failures hit zero, chapter-tool calls numbered about 174, and whole-file edits stayed at zero. Thirteen chapters grew by ≥300 characters; the run took about 19 minutes. Versus a July 21 baseline on the same model ID, tool discipline improved; throughput did not.

Read the post →
Six calligraphy brushes copying the same passage on one scroll, with the cinnabar-red brush making the steadiest strokes
Model evalsJul 31, 202610 min

Every Luna reasoning effort vs Terra Medium in a 20-round fiction benchmark

A controlled 32K-output fiction benchmark of GPT-5.6 Luna none/low/medium/high/xhigh/max and Terra medium: 560 planned rounds, 548 completed prose responses, five incomplete streams, two Chinese stories, and two 20-round replicates per story. Two isolated model-review sessions—not humans—ranked Terra medium first at 4.725/5 and Luna high as the best Luna setting. Includes quality, normalized cost, time to first prose, chain completion, and 450–700-character compliance.

Read the post →
Three writing exam papers labeled DeepSeek V4 Flash, GPT-5.6 Luna, and GPT-5.6 Terra laid side by side
Model comparisonJul 31, 202610 min

DeepSeek V4 Flash, GPT-5.6 Luna, or Terra for fiction? Put the two test suites on the right footing first

A data-backed comparison of DeepSeek V4 Flash 0731 and GPT-5.6 Luna/Terra for fiction: official input, cache, and output rates; 32K long-form completion; time to first prose; 450-700-character compliance; reasoning exhaustion; agent tools; and cache reads. The two suites were not one head-to-head blind test: Terra medium won the Luna/Terra matrix, while DeepSeek none is the cheapest practical tier.

Read the post →
Ink-and-cinnabar illustration of an unplugged cable between a phone chat window and a server rack, with a wall calendar page reading July 24 drifting to the floor
Time-sensitiveJul 28, 20267 min

JanitorAI proxy stopped working? Two different breakages, two different fixes

If your JanitorAI DeepSeek proxy died this week, the cause has a timestamp: July 24, 2026, 15:59 UTC, when DeepSeek retired the deepseek-chat and deepseek-reasoner model names. If your configs came up blank instead, that's a separate April incident with an official Recover button. The one-line fix for each, plus the full triage table: 401s, 402s, the save-then-refresh ritual, and region age checks.

Read the post →

Earlier posts (89)

Jul 2026