Foreverse Research · Fiction Bench

How good is Gemini 3.1 Pro at writing fiction?

One thing first: its last-place xuanhuan rank was mostly our own instruction-wording bug — the old phrasing made it treat AI-continued passages as “ignorable reference,” so it rewrote the opening for 20 straight rounds; with corrected wording, 0/20 rewinds. Under the fixed condition it placed a mid-field 4-6 on romance — but its simile-stock density stayed the field's highest. The looping was our bug; the purple prose is its nature.

Xuanhuan fantasy · Rank 9 / 9Court romance · Rank 4-6 / 9Restart loop$0.56 / 10k

Xuanhuan fantasy

9 / 9

Both packs: p303 第9 · p404 第9

Under the old directive it rewound to the continuation point 20/20 rounds (rewriting the opening every time, zero progress) — ablation proved this was a bug in our instruction wording; 0/20 after the fix.

Court romance

4-6 / 9

Both packs: p505 第6 · p606 第4

With the fixed directive it climbed from disqualified to mid-field; its simile-stock density of 2.50‰ stayed the field's highest — the purple prose is its nature, the looping was our bug.

What the 20-round chains actually showed

Its xuanhuan chain (old directive) replayed the same opening in all three windows — 20 rounds, zero progress, a unanimous #9 in both packs. A three-way ablation later dissected the disqualification: old wording (“for plot-transition reference only”) → 20/20 rewinds; explanation removed → ~6/8; corrected wording (“established story canon”) → 0/20. Faced with an unexplained source-marker glyph, its default reading was “annotated text isn't canon.”

The fix principle graduated into the product: the explanation sentence declares status only (established canon, not ignorable reference) and never commands actions. DeepSeek-family models were insensitive to all three wordings; Gemini was the easiest to mislead — so instruction robustness gets designed against it.

On the romance track (fixed directive) it competed normally at 4-6 — and its 2.50‰ flavor density stayed the field's highest. What remains after our bug was fixed is the model's own temperament: a constant, non-drifting lean toward ornate simile.

Long-run failure mode

Restart loop

Restart loop (triggered by our old directive): every round returned to the source's last line and rewrote the opening — 20 rounds, zero progress. Three-way ablation: old wording 20/20 rewinds, no explanation ~6/8, fixed wording 0/20. Its default reading of an unexplained marker was “annotated text isn't canon”; the explanation sentence is load-bearing.

Structural fingerprint

Highest simile-stock density in the field (2.47‰ xuanhuan post-fix / 2.50‰ romance vs the 1.0‰ source baseline); constant purple-prose lean.

Cross-genre profile: The directive-wording fix was its watershed: from disqualified to normal competition.

Test-condition disclosure (hosted models are moving targets)

Model under test: gemini-3.1-pro (released 2026-02-19)

Evaluated: 2026-07-16 · Access channel: yunwu aggregator gateway

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.56 list price: in $2/M · out $12/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Gemini 3.1 Pro.

Continue your book with it

How to cite

Foreverse Research, “How good is Gemini 3.1 Pro at writing fiction (Fiction Bench),” 2026-07. https://foreverse.app/research/fiction-bench/gemini-3-1-pro

Keep going

← Back to the leaderboard

Gemini 3.1 Pro for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng