The validation record

What is the headline result?

Across 4 frontier model families the regulated arm was calmer in 15 of 16 generator×judge cells by sign (11/12 off-diagonal), and 13 of 16 reached p<.05. One cell returned a flat null at −0.04, and homeostatic recovery replicated in 4/4 families.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

What is the null cell?

One cell of the sixteen, GPT generating and Gemini judging, came back at −0.04, which is flat rather than a difference in either direction. GPT×Gemini. The three sub-threshold cells all sit on the near-saturated GPT host. The null is reported beside the headline everywhere the headline appears.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

Why is the result 15 of 16 and not every cell?

One cell of sixteen returned a null at −0.04 and three fell short of the p<.05 threshold, all three sitting on the near-saturated GPT host. Rounding the figure up to every cell would misstate the record, so the published claim stays at 15 of 16 by sign with 13 of 16 significant.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

Which model families were tested?

Four frontier families were tested as both generator and judge in a symmetric design: GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash and Grok 4.3. The fourth family, from xAI, is lineage-independent of the other three, which is a deliberate control against shared judge lineage or house style.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

Why did every family judge every other?

A single judge model can flatter its own lineage, so every family scored every generator including itself, producing a symmetric four by four matrix. The off-diagonal cells, where judge and generator share no lineage, carry the weight of the result: 11/12 of them are calmer by sign.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

How many judged units were scored?

Each generator contributed 202 judged units, scored by panels drawn from all four families. The design is pre-registered, so the number of units, the scoring protocol and the analysis were fixed before any data was collected.

Source: Pre-registered cross-family evaluation. Paper: arXiv 2607.24339, DOI 10.5281/zenodo.21303518.

What is homeostatic recovery, and did it replicate?

Homeostatic recovery is the controller returning to baseline after a provocation instead of staying elevated. Recovery replicated in 4/4 frontier families, and the signature is arousal 0 → 0.273 → 0.469 → 0.338 → 0.152 → 0.068 → 0.031 → 0.022, DEFAULT by turn 6.

Source: Pre-registered ablation, 2026-07-21, gpt-5.6-sol, N=3.

The spend receipts

What does the receipts benchmark measure?

The receipts benchmark measures upstream spend on identical work. Both arms make the same number of attempts against the same battery, so the difference in the bill is the whole measurement rather than an artefact of one arm attempting less.

Source: Pre-registered receipts benchmark, 2026-07-19. Scored output of harness/report.py.

How much did the governed arm save?

The governed arm paid 4.1% to 20.2% of the ungoverned bill across 7 model families on the verbatim-loop battery, a reduction of 79.8% to 95.9%. Best measured case was GPT-5.6 Sol at $0.1669 ungoverned against $0.0068 governed.

Source: Pre-registered receipts benchmark, 2026-07-19. Scored output of harness/report.py.

The full 12-row table

What qualifiers travel with the spend figures?

Three qualifiers are mandatory. OpenRouter rows are the upstream's own metered usage cost rather than our token math, the Gemini free-tier rows are token deltas that carry no dollar claim, and every figure is a scoring-script output rather than an estimate.

Source: Pre-registered receipts benchmark, 2026-07-19. Scored output of harness/report.py.

Is the governor just a system prompt in disguise?

A pre-registered ablation compared three arms on one host model. The governed arm paid 23% to 63% of baseline spend, while a strong static calming system prompt paid 117% to 192% of baseline. The prompt did not contain the runaway and cost more than doing nothing.

Source: Pre-registered ablation, 2026-07-21, gpt-5.6-sol, N=3.

Guide: the measured comparison

Scope and limits

Which latency figure is the real one?

Three latency figures exist and none substitutes for another. One controller tick is p99 2.8 µs as an in-process microbenchmark over 100,000 iterations. End-to-end governed proxy overhead is p50 1.2 ms and p99 2.4 ms. The WebAssembly build runs at roughly ~125 ns per tick at the edge.

Source: Bench run 2026-07-18, latency_bench.py, Windows 11 AMD64, Python 3.14.5.

What does Gubernaut deliberately not claim?

Several figures stay dark because they were not measured cleanly. Absolute dollars for Llama from token math are withheld because the catalog price sits 24% off the metered cost, Gemini free-tier rows carry no dollar claim at all, and long-history overhead against a real upstream came back inconclusive at n=5 and is reported that way.

Source: Pre-registered receipts benchmark, 2026-07-19. Scored output of harness/report.py.

What are the documented failure modes?

Five failure modes are documented as F1 to F5: recovery failure under intensity-only drive, scar tissue in the arbiter, a de-escalation false positive, a harness gap caught by the practice gate, and the null cell. Each fix was one bounded change, declared before it was written and re-tested against frozen criteria.

The full record

Last reviewed 2026-08-02 · paper arXiv 2607.24339, DOI 10.5281/zenodo.21303518