Every Luna reasoning effort vs Terra Medium in a 20-round fiction benchmark

A controlled 32K-output fiction benchmark of GPT-5.6 Luna none/low/medium/high/xhigh/max and Terra medium: 560 planned rounds, 548 completed prose responses, five incomplete streams, two Chinese stories, and two 20-round replicates per story. Two isolated model-review sessions—not humans—ranked Terra medium first at 4.725/5 and Luna high as the best Luna setting. Includes quality, normalized cost, time to first prose, chain completion, and 450–700-character compliance.

Six calligraphy brushes copying the same passage on one scroll, with the cinnabar-red brush making the steadiest strokes

The short answer is unusually clean. Under the same 32K output envelope, on the same two Chinese fiction fixtures, across the same 20 consecutive rounds, Terra medium placed first overall and Luna high placed first within Luna. Luna xhigh and max did not buy another quality step; they bought more waiting, more reasoning, or worse chain completion. For routine passage-by-passage continuation, none and low are the practical Luna defaults.

This was not a seven-opening beauty contest. Each setting continued two original stories, with two replicates per story and up to 20 linked rounds per chain. Every response became context for the next round, so the model had to preserve the characters, objects, and promises it had introduced itself. The plan contained 560 slots. The runner sent 553 requests, received 548 complete responses with prose, and recorded five incomplete_stream outcomes.

The exam: seven settings, two genres, 20 rounds per chain

The field was Luna at none, low, medium, high,xhigh, and max, plus Terra at medium. The fixtures were an original historical-supernatural investigation, Long Night Lamp, and an original contemporary emotional mystery, The Seventh Call. Every setting ran two replicates of each fixture, 20 rounds per chain:7 × 2 × 2 × 20 = 560 planned slots.

Every formal sample used the same Requesty Responses route with a 32,768-token maximum output envelope. “32K” here means the output ceiling for this run. It does not mean each call generated 32K tokens, and it is not a claim that the input context was 32K. Reasoning and visible prose share the output allowance. An earlier 8K screen had produced a response where reasoning exhausted the allowance before prose; all 548 completed calls in this matrix contained prose. Raising the envelope removes that particular cause of empty output. It does not prove long-term API stability.

The blind judges were models, not humans

“Double blind” can imply a human panel, so the identity belongs near the top: both judges were isolated model-review sessions, not people. They were also not two models from different vendors. Each session received a different anonymous label mapping, knew neither the Luna/Terra identity nor the effort, and had no access to the other session’s judgment. Both score sets were frozen on July 31, 2026 before the mappings were revealed.

Every chain was sampled in early, middle, and late windows and scored from 1 to 5 on continuation coherence, character/fact stability, style fit, plot advancement, and control of self-repetition. The primary score is an equal-weight mean across two judges, four planned chains, three windows, and five dimensions. Completion and length compliance do not add quality points. Across the 28 chains, judge agreement was Pearson 0.868 andSpearman 0.853; the mean absolute difference between chain means was 0.336. Both judges agreed on the top two. Fine ordering lower down deserves restraint.

Overall ranking: Terra Medium first, Luna High second

RankSettingFive-dimension meanJudge AJudge BReading
1Terra medium4.725 / 54.8174.633Stable across both genres and all four chains; clear winner
2Luna high4.175 / 54.4173.933Best Luna quality, with a large latency tradeoff
3Luna xhigh3.975 / 54.2173.733Effectively tied with max; highest latency and reasoning cost within Luna
4Luna max3.950 / 54.0003.900Completed every chain but did not beat high
5Luna none3.658 / 53.7673.550Slightly better low-tier quality; good routine default
6Luna low3.600 / 53.7333.467Slightly faster and cheaper; good routine default
7Luna medium3.592 / 53.6173.567No gain over none / low in this run

Terra medium was not carried by one friendly genre. It scored 4.717 on the historical mystery and 4.733 on the contemporary mystery, with all four chains landing in the field’s top four. Its quality moved from 4.850 early to 4.550 late, a decline of 0.300. Luna high moved from 4.625 to 3.850, down 0.775; the three low efforts ended at only 2.975–3.100 late. What separated long runs was not sentence polish at the opening. It was whether the model stopped adding another door, another hidden identity, and another mechanism, then resolved the conflict already on the page.

High’s position inside Luna is also clear. It beat xhigh by 0.200 and max by 0.225. Xhigh and max differed by only 0.025—far too little to market as a meaningful tier win. More reasoning effort was not monotonically better fiction. High found the best balance in this run; above it, most of the gain was waiting and hidden reasoning rather than visible quality.

Cost, speed, completion, and length compliance belong beside—not inside—the quality score

“Median first prose” below uses completed calls only and averages the two central values for even-sized groups. A “complete 20-round chain” reached round 20 for one setting, fixture, and replicate. Length compliance means only that a completed response landed inside the requested 450–700 Chinese-character band; it is not a style or plot score. Dollar figures reprice each completed call with Foreverse’s normalized user rates. They are not Requesty’s final invoice.

SettingComplete prose / sentComplete 20-round chains450–700 charsMedian first proseNormalized repricing
Luna none80 / 804 / 421 / 80 · 26.25%2.256s$0.4806
Luna low80 / 804 / 417 / 80 · 21.25%2.132s$0.4723
Luna medium80 / 804 / 416 / 80 · 20.00%3.255s$0.4919
Luna high77 / 783 / 432 / 77 · 41.56%16.420s$0.6521
Luna xhigh73 / 761 / 424 / 73 · 32.88%56.049s$1.1782
Luna max80 / 804 / 427 / 80 · 33.75%14.652s$0.6811
Terra medium78 / 793 / 434 / 78 · 43.59%3.429s$4.4941
Total548 / 55323 / 28171 / 548 · 31.20%4.217s$8.4503

Terra medium and Luna medium form the clean same-effort comparison: Terra scored 1.133 points higher and reached first prose only about 0.174 seconds later, but its normalized cost was about 9.14 times as high. You buy quality, not value. Luna none and low completed all four chains, reached prose in roughly 2.1–2.3 seconds, and each cost less than $0.49 for the entire matrix slice. Their quality difference was only 0.058, so routine continuation does not need brand-like loyalty to one of them.

High makes sense as an on-demand “this passage matters” upgrade: it was Luna’s quality winner, but median first prose rose to 16.420 seconds. Xhigh used 421,227 reasoning tokens, took 56.049 seconds to first prose, and completed only one of four chains. Max used 120,270 reasoning tokens and completed 4/4 chains, yet still scored below high, which used 116,066. That is why xhigh and max are poor routine long-form defaults on this evidence—not a claim that they are useless for every task.

The fixed summary omitted a stage-name relationship

The blind material itself had one disclosure-worthy defect. In The Seventh Call, the mother’s real name is Shen Yunqiu and Ye Hang is her broadcasting stage name. The fixed-facts summary said only “mother Ye Hang” and omitted that the two names belong to one person. Judge B therefore treated several correct uses of Shen Yunqiu as name drift, depressing the absolute character/fact-stability scores for the contemporary fixture.

The scores were already frozen; reading the answer key is not permission to rewrite the exam. We kept the five-dimension ranking and ran one sensitivity check: remove the facts dimension from every sample, then recompute the remaining four dimensions.

SettingOriginal five-dimension rank / scoreWithout facts: rank / score
Terra medium1 · 4.7251 · 4.792
Luna high2 · 4.1752 · 4.177
Luna xhigh3 · 3.9754 · 3.969
Luna max4 · 3.9503 · 4.052
Luna none5 · 3.6585 · 3.719
Luna low6 · 3.6006 · 3.677
Luna medium7 · 3.5927 · 3.646

The core result survives: Terra medium remains first and Luna high remains the best Luna setting. Only the effectively tied xhigh and max swap places. A rerun should state “mother Shen Yunqiu (broadcasting stage name Ye Hang)” in the fixed facts. It should not rescore this batch with a repaired summary and pretend the new score came from the original blind review.

What to use

If you continue one ordinary passage at a time and want low cost, quick prose, and four fully completed chains, Luna none or low is the default answer. None was slightly better on quality; low was slightly faster and cheaper. Luna medium showed no clear benefit over either in this run.

If one chapter deserves an extra wait, Luna high is Luna’s quality setting. If budget permits and you want the strongest cross-genre long-form prose in this matrix, Terra medium is first. Xhigh and max did not return enough extra quality to recommend keeping them on for routine fiction.

Keep the boundary with the recommendation: this was one Requesty Responses matrix with two original Chinese genres and two replicates. Model-judge agreement is not human preference, five incomplete streams are not an API SLA, and a 450–700-character hit is not the same thing as good fiction. To place this suite beside DeepSeek V4 Flash 0731 without fabricating a shared leaderboard, read the DeepSeek, Luna, and Terra comparison. Readers of Chinese can inspect the unedited round 1 and round 20 samples before seeing which configuration wrote each one. For protocol, caching, and normalized-pricing evidence on these model IDs, read the Luna / Terra relay API field test. For why long-form systems need consecutive-round tests, see our long-run continuation guide.

FAQ

Which is better for fiction, GPT-5.6 Luna or Terra?

In this 32K-output long-form matrix, Terra medium ranked first with a two-judge five-dimension mean of 4.725/5. Luna high was the best Luna setting at 4.175/5. Choose Terra medium when prose quality outranks budget, Luna none or low for routine passage-by-passage continuation, and Luna high for an occasional quality-focused step up. These findings cover two Chinese genres, not every language or writing task.

Were the blind reviewers human?

No. They were two isolated model-review sessions, not people and not models from two different vendors. Each session received a different anonymous mapping and froze its scores before identities were revealed. Across the 28 continuation chains, judge agreement was Pearson 0.868 and Spearman 0.853. That supports this run's relative ranking, but it does not replace reader research.

Does Luna xhigh or max write better fiction than high?

Not in this run. Luna high scored 4.175, xhigh 3.975, and max 3.950. Xhigh took a median 56.049 seconds to first prose and completed only one of four 20-round chains. Max completed all four chains, but did not beat high on quality and used more reasoning and normalized cost. Neither is a good routine long-form default on this evidence.

Was $8.4503 the benchmark's actual API bill?

No. It is the sum obtained by repricing the 548 completed calls, one by one, with Foreverse's normalized user rates. The five incomplete streams had no terminal usage, so their possible upstream charges are unknown. They cannot be added to $8.4503 or declared free.

Do 548 completed responses and five incomplete streams measure API stability?

Not as a long-term stability rate or SLA. The matrix planned 560 slots and sent 553 requests. Of those, 548 reached a terminal event with non-empty prose; five HTTP 200 streams ended before the terminal event and stopped seven later slots from being sent. This is completion behavior from one controlled Requesty Responses run, and it cannot isolate model, network, or relay responsibility.

Questions or ideas? Join our Discord →