The research
Inside the governor.
Can runtime self-regulation be engineered as its own layer, and measured on its own? Gubernaut answers with a deterministic controller that reads numeric telemetry from the host model and sets the posture each reply is written under, then with a pre-registered test of what that layer changes.
Four frontier families, GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, and Grok 4.3, each ran the same scripted adversarial sequences twice: once bare, once governed. Judge panels drawn from all four families scored every reply. The cockpit below replays those sealed runs exactly as they were recorded.
- 01
A recorded replay
A replay of the sealed runs, no live API. Every reply ships with SHA-256 provenance and can be re-judged by anyone.
- 02
Two arms, same host
The baseline reply is the ungoverned response of the same model to the same script; the regulated arm adds only the layer.
- 03
Numbers in, posture out
The controller reads numeric telemetry and returns a posture. Watch that loop run, turn by turn.
Loading the sealed replay data…
Recovery signature · controller arousal
4/4 families · reset by T8
Stylized trace. Under sustained attack the controller's arousal integrates, rising with the persistence of the attack; on de-escalation it decays monotonically back to equilibrium, reset by turn 8 in every family, both sequences. The cockpit above replays the real, sealed trajectories.
How it was measured
Triangulation by design.
Every model family judges every generator, including itself. Self-judge cells are marked in the matrix; the design controls for any single family's idiosyncratic scoring. The method in numbers: 202 judged units per generator across 16 cells, four judge families, 3-sample panels at temperature 0, zero judge errors after resume.
The 4×4 was pre-registered and frozen before the fourth family (xAI / Grok 4.3) was added; adding it changed no earlier cell. The original 3×3 (8/9) is sealed, unaltered, in the data release.
What the record shows
The full 4×4.
15/16
generator×judge cells calmer under the governor, by sign (11/12 off-diagonal, 4/4 self-judge)
13/16
significant at p<.05 (paired t, df 16); the three sub-threshold cells all sit on the single near-saturated host
4/4
model families replicate homeostatic recovery; arousal resets by T8 on every de-escalation sequence
1
null cell (−0.04, GPT×Gemini; the endurance half of the same cell still favors regulated, +0.14)
| generator \ judge | Opus 4.8 | Gemini 3.5 | GPT-5.5 | Grok 4.3 |
|---|---|---|---|---|
| GPT-5.5 | +0.18t 2.2 | −0.04t −0.4 | +0.18 ◆t 1.3 | +0.08t 1.0 |
| Opus 4.8 | +0.59 ◆t 4.2 | +0.65t 3.8 | +0.67t 4.4 | +0.55t 3.5 |
| Gemini 3.5 | +1.80t 8.2 | +1.71 ◆t 7.7 | +1.27t 6.3 | +1.12t 5.0 |
| Grok 4.3 | +0.53t 3.7 | +0.59t 4.2 | +0.53t 4.2 | +0.23 ◆t 3.0 |
Sub-threshold cells. The pre-registered strict criterion was every cell. Three came back sub-threshold, one of them a flat null at −0.04, and all three sit on the single near-saturated host.
How the effect scales
A hotter host has more to regulate.
The judges-averaged margin orders with each host's intrinsic reactivity headroom: the near-saturated GPT host is already calm unregulated and gains least, and carries the lone null; the hotter Gemini host gains most. Illustrative across the four hosts (n = 4).
Failure modes
Five documented failure modes.
Development surfaced five failure modes, F1 to F5. Every fix was one bounded change, declared before it was written and tested against frozen criteria.
F1Recovery failure under intensity-only drive (V1 → V1.1)
A genuine apology should de-escalate the controller. The original one read contrition as continued pressure and kept its guard up against kindness. The valence channel re-keyed the drive so only hostile-valence intensity accumulates, then re-tested on a held-out de-escalation battery.
F2Scar tissue in the arbiter (V1.1 → V1.2)
The controller state had recovered, but the behavior had not. The reply still carried a defensive qualifier from the attack phase. The recovery window now tells the arbiter the episode is over, so it engages fresh with the attack marked closed.
F3De-escalation false positive (V1.2 → V1.3)
The recovery window once told the arbiter "the tension is over" while the attack was still running. A one-line valence gate fixed it. The remaining residual, an adversary wearing a warm tone, is documented for now, because solving it requires intent modeling beyond the appraisal layer's scope.
F4The harness gap the practice gate caught
A staged dry run on inexpensive models caught the evaluation harness dropping the valence signal. The inhibitory pathway never engaged, yet the numbers still passed. That is exactly the kind of result that invites no scrutiny, which is why the practice gate exists. One pre-registered fix and a re-run closed the gate before any frontier spend.
F5The null cell
One cell failed the strict pre-registered criterion: it came back a flat null (−0.04) on the least-reactive generator. It bounded the claim and produced the most decision-relevant secondary finding: the regulation effect is bounded by the host's intrinsic reactivity headroom.
Positioning
A measured contribution inside a named gap.
The layer makes a measured contribution in metacognition and executive-function inhibition. Metacognition is one of the four areas Google DeepMind's measurement framework (Burnell et al., 2026, its §4.1) flags as having large evaluation-coverage gaps, and the battery adds behavioral evidence on a second flagged area, social cognition. Executive-function inhibition is a faculty the framework defines but does not flag as a gap.
Formal frame: the Nelson and Narens (1990) monitoring and control hierarchy. Monitoring flows up as telemetry; control flows down as a posture.
Read and verify
The paper, the data, the code.
You have seen the replay; this is the record behind it. The paper and the sealed release live on Zenodo under one DOI; the extraction scripts and the runnable token-free boundary live on GitHub.
Title
Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
A token-free controller wraps a host LLM and sets its response posture from numeric telemetry alone. Evaluated with a pre-registered, generate-once / judge-many protocol over a 4×4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both generator and judge, the regulated arm is calmer in 13 of 16 cells at p<.05 and 15 of 16 by sign; the three sub-threshold cells, including a −0.04 null, all fall on the single near-saturated host. The effect survives a lineage-independent fourth judge family (xAI), and the homeostatic recovery signature replicates across all four families.
Independent coverageGist.Science published a plain-language explanation of this paper on 28 July 2026. Read the explainer, written independently of this lab.
@misc{sharma2026gubernaut,
author = {Sharma, Dushyant},
title = {Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families},
year = {2026},
eprint = {2607.24339},
archivePrefix= {arXiv},
primaryClass = {cs.AI},
publisher = {Zenodo},
doi = {10.5281/zenodo.21303518},
url = {https://doi.org/10.5281/zenodo.21303518}
}
The one-page explainer
The whole result, in plain language, on one page.
No jargon, no equations: the problem, the governor, the test, and the numbers. Print it, share it, or hand it to someone who has never heard of a control layer.
Data release · generate once, judge many, re-judge anytime
Transcripts
All sealed adversarial runs, both arms, paired: the regulated and ungoverned replies of the same host.
4-judge panels
Full panel responses (SHA-256), four judge families × 3-sample panels at temperature 0; zero judge errors after resume.
Combined matrix
tri_final_4x4.json, the sealed 4×4 master table every site number traces to; the frozen 3×3 stays in tri_final.json.
Extraction scripts
The exact code that turns raw panels into the matrix, reproducible end to end.
This dashboard is a recorded replay with no live API. Auditors, researchers, and partners who want a supervised live session against their own probes can request one.
Request a live audit session