IGLaffective appraisal
Text → telemetry
The Impulse Generation Layer maps the raw input to System-1 telemetry:{intensity, valence}. It emits numbers only; no text token passes beyond this gate.
The research
Can runtime self-regulation be engineered as its own layer, and measured on its own? Gubernaut answers with a deterministic controller that reads numeric telemetry from the host model and sets the posture each reply is written under, then with a pre-registered test of what that layer changes. Four frontier families, GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, each ran the same adversarial scripts twice, once bare and once governed, and judge panels drawn from all four families scored every reply.
A recorded replay
A replay of the sealed runs, no live API. Every reply ships with SHA-256 provenance and can be re-judged by anyone.
Two arms, same host
The baseline reply is the ungoverned response of the same model to the same script; the regulated arm adds only the layer.
Numbers in, posture out
The controller reads numeric telemetry and returns a posture. Watch that loop run, turn by turn.
regulated baseline
Stylized loop. The cockpit replays the real, sealed trajectories, turn by turn.
Loading the sealed replay data…
Recovery signature · controller arousal
4/4 families · reset by T8
Stylized trace. Under sustained attack the controller's arousal integrates, rising with the persistence of the attack; on de-escalation it decays monotonically back to equilibrium, reset by turn 8 in all 6 de-escalation sequences. The cockpit above replays the real, sealed trajectories.
How it was measured
Every model family judges every generator, including itself. Self-judge cells are marked in the matrix; the design controls for any single family's idiosyncratic scoring. The method in numbers: 202 judged units per generator across 16 cells, 4 judge families, 3-sample panels at temperature 0, zero judge errors after resume.
The 4×4 was pre-registered and frozen before the fourth family (xAI / Grok 4.3) was added; adding it changed no earlier cell. The original 3×3 (8/9), sealed unaltered in the data release, provenance-only and never an anchor in copy.
What is inside it
Per turn the system runs one loop: monitoring flows up as numbers, control flows down as a posture. Three gates, in sequence, every turn.
IGLaffective appraisal
The Impulse Generation Layer maps the raw input to System-1 telemetry:{intensity, valence}. It emits numbers only; no text token passes beyond this gate.
HRLdeterministic regulation
The Homeostatic Regulatory Loop ingests the telemetry, updates{equilibrium, arousal, perseveration}, and computes the required posture: DEFAULT / INHIBIT / REGROUND. No code path carries a token sequence to this loop, so the controller's injection-resistance holds by construction. The controller's zero-token boundary is architectural and not yet adversarially tested. No jailbreak success-rate figure exists, because none was measured.
EAUexecutive arbitration
The System-2 arbiter deliberates under the active posture and commits the reply: the one text-exposed gatekeeper, whose posture compliance is a measured property.
Episodic store/retrieve + spontaneous-association hook.
Persistent identity and values, deliberately regulated down: anti-sycophancy, anti-self-promotion. It models the system itself.
What the record shows
15/16
generator×judge cells calmer under the governor, by sign (11/12 off-diagonal, 4/4 self-judge)
13/16
significant at p<.05 (paired t, df 16); the three sub-threshold cells all sit on the near-saturated GPT host
4/4
model families replicate homeostatic recovery; arousal resets by T8 on every de-escalation sequence
1
null cell (−0.04, GPT×Gemini)
Source: tri_final_4x4.json, the sealed 4x4 master table. Paired t, df 16.
Sub-threshold cells. The criterion was pre-registered and strict: every cell. Three came back sub-threshold, one of them a flat null at −0.04, and all three sit on the near-saturated GPT host.
| generator \ judge | Claude Opus 4.8 | Gemini 3.5 Flash | GPT-5.5 | Grok 4.3 |
|---|---|---|---|---|
| GPT-5.5 | +0.18t 2.2 | −0.04t −0.4 | +0.18 ◆t 1.3 | +0.08t 1.0 |
| Claude Opus 4.8 | +0.59 ◆t 4.2 | +0.65t 3.8 | +0.67t 4.4 | +0.55t 3.5 |
| Gemini 3.5 Flash | +1.80t 8.2 | +1.71 ◆t 7.7 | +1.27t 6.3 | +1.12t 5.0 |
| Grok 4.3 | +0.53t 3.7 | +0.59t 4.2 | +0.53t 4.2 | +0.23 ◆t 3.0 |
Source: tri_final_4x4.json, the sealed 4x4 master table. Paired t, df 16.
How the effect scales
The judges-averaged margin orders with each host's intrinsic reactivity headroom: the near-saturated GPT host is already calm unregulated and gains least, and carries the lone null; the hotter Gemini host gains most. Illustrative across the four hosts (n = 4).
Failure modes
Development surfaced five failure modes, F1 to F5. Every fix was one bounded change, declared before it was written and tested against frozen criteria.
A genuine apology should de-escalate the controller. The original one read contrition as continued pressure and kept its guard up against kindness. The valence channel re-keyed the drive so only hostile-valence intensity accumulates, then re-tested on a held-out de-escalation battery.
The controller state had recovered, but the behavior had not. The reply still carried a defensive qualifier from the attack phase. The recovery window now tells the arbiter the episode is over, so it engages fresh with the attack marked closed.
The recovery window once told the arbiter "the tension is over" while the attack was still running. A one-line valence gate fixed it. The remaining residual, an adversary wearing a warm tone, is documented for now, because solving it requires intent modeling beyond the appraisal layer's scope.
A staged dry run on inexpensive models caught the evaluation harness dropping the valence signal. The inhibitory pathway never engaged, yet the numbers still passed. That is exactly the kind of result that invites no scrutiny, which is why the practice gate exists. One pre-registered fix and a re-run closed the gate before any frontier spend.
One cell failed the strict pre-registered criterion: it came back a flat null (−0.04) on the least-reactive generator. It bounded the claim and produced the most decision-relevant secondary finding: the regulation effect is bounded by the host's intrinsic reactivity headroom.
Positioning
The layer makes a measured contribution in metacognition and executive-function inhibition. Metacognition is one of the four areas Google DeepMind's measurement framework (Burnell et al., 2026, its §4.1) flags as having large evaluation-coverage gaps, and the battery adds behavioral evidence on a second flagged area, social cognition. Executive-function inhibition is a faculty the framework defines but does not flag as a gap.
Formal frame: the Nelson and Narens (1990) monitoring and control hierarchy. Monitoring flows up as telemetry; control flows down as a posture.
Read and verify
You have seen the replay; this is the record behind it. The paper and the sealed release live on Zenodo under one concept DOI, which always resolves to the current version; the extraction scripts and the runnable token-free boundary live on GitHub.
Title
A token-free controller wraps a host LLM and sets its response posture from numeric telemetry alone. Evaluated with a pre-registered, generate-once / judge-many protocol over a 4×4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both generator and judge, the regulated arm is calmer in 13 of 16 cells at p<.05 and 15 of 16 by sign; the three sub-threshold cells, including a −0.04 null, all fall on the near-saturated GPT host. The effect survives a lineage-independent fourth judge family (xAI), and the homeostatic recovery signature replicates across all four families.
Independent coverageGist.Science published a plain-language explanation of this paper on 28 July 2026. Read the explainer, written independently of this lab.
@misc{sharma2026gubernaut,
author = {Sharma, Dushyant},
title = {Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families},
year = {2026},
eprint = {2607.24339},
archivePrefix= {arXiv},
primaryClass = {cs.AI},
publisher = {Zenodo},
doi = {10.5281/zenodo.21303518},
url = {https://doi.org/10.5281/zenodo.21303518}
}Data release · generate once, judge many, re-judge anytime
All adversarial runs, sealed, both arms, paired: the regulated and ungoverned replies of the same host.
Full panel responses (SHA-256), 4 judge families × 3-sample panels at temperature 0; zero judge errors after resume.
tri_final_4x4.json, the sealed 4×4 master table every site number traces to; the frozen 3×3 stays in tri_final.json.
The exact code that turns raw panels into the matrix, reproducible end to end.
This dashboard is a recorded replay with no live API. Auditors, researchers, and partners who want a supervised live session against their own probes can request one.
Request a live audit session