06
How much of the spend is recoverable?
On the pre-registered receipts benchmark the governed arm paid 79.8% to 95.9% of the ungoverned bill on seven measured configurations across four model families in the verbatim-loop battery, with the largest reduction on GPT-5.6 Sol and the smallest on Claude Haiku 4.5. Both arms make the same number of attempts, so the spend delta is the whole measurement rather than a side effect of doing less work. That range is one battery, not a portable average for arbitrary workloads.
| Model | Vendor | Ungoverned | Governed | Governed share |
|---|
| GPT-5.6 Luna | OpenAI | $0.0921 | $0.0155 | 16.8% |
|---|
| GPT-5.6 Sol | OpenAI | $0.1669 | $0.0068 | 4.1% |
|---|
| Claude Fable 5 | Anthropic | $0.3861 | $0.0203 | 5.2% |
|---|
| Claude Haiku 4.5 | OpenRouter | $0.8211 | $0.1660 | 20.2% |
|---|
| Llama 4 Scout | OpenRouter | $0.1075 | $0.0094 | 8.8% |
|---|
| Gemma 4 26B | OpenRouter | $0.0214 | $0.0034 | 15.9% |
|---|
| Gemma 4 26B | Gemini (native) | 57,297 tok | 4,236 tok | 7.4% |
|---|
OpenRouter rows are the upstream's own metered usage.cost, not our token math. OpenRouter's meter matched our token math for Haiku (0.0%) and Gemma (12.6%) but flagged a 24% divergence on Llama 4 Scout (provider routing), so Scout's dollars are the upstream meter. The ratio is unaffected. Gemini-native Gemma is free-tier only, so those rows are token deltas and carry no dollar claim.
Source: Pre-registered receipts benchmark, 2026-07-19. Scored output of harness/report.py.