Certificates: the bands, the witness, the replay
Cody Lee Walker · s2ar · calibration seen 2026-10-10 · a living page: it follows the landing’s calibration line
The bands on this page are the ones the live landing serves (
data/calibration/live.json, line sha2c115bd3cae8…); the tower measures them and this page never carries an older set than the landing. A certificate is a signed Assay record that an image, compressed, still reads the same to a vision model, within a band measured per reader. MEASURED — not argued.
17 models measured; the landing shows 4 demo items at scales ×1.0, ×0.85, ×0.7, ×0.55, ×0.4, ×0.3, ×0.2, ×0.12, ×0.08, ×0.05; the certify walkthrough's record `f164152c1dd55564…` verifies free at `/v1/verify/<sha>`.
1. What a band is
certify compresses an image and has a witness measure what a vision model would still read from it. The witness is deterministic
(recon: gradient-energy over the reconstruction, with RMSE for coarse images and OCR for text); the band is how far the witness’s
score may drop before that model’s answers start to move, measured per model on the calibration items. Tolerance is per
reader: how far the band must tighten so the witness never outruns the model. A strict tolerance is a model that loses detail early.
2. The bands, per model
Strictest tolerance first; typical and max are vision-token savings at the certified scale.
| model | provider | tolerance | typical savings | max | n |
|---|---|---|---|---|---|
| Gemini 3 6 Flash | 0.35 | −62% | −99% | 9 | |
| Qwen3 7 Flash | OpenAI | 0.5 | −96% | −99% | 26 |
| Qwen3 7 Plus | OpenAI | 0.5 | −91% | −99% | 91 |
| Llama 4 Scout | Meta | 0.5 | −89% | −99% | 83 |
| Llama 3.2 11B | Meta | 0.5 | −88% | −99% | 77 |
| Grok 4 5 | OpenAI | 0.5 | −87% | −99% | 91 |
| Claude Sonnet 4.6 | Anthropic | 0.6 | −91% | −99% | 88 |
| GPT-4o mini | OpenAI | 0.6 | −89% | −99% | 76 |
| Claude Fable 5 | Anthropic | 0.6 | −88% | −99% | 28 |
| Gemini 3.5 Flash | 0.6 | −88% | −99% | 77 | |
| GPT-5.4 mini | OpenAI | 0.6 | −86% | −99% | 87 |
| GPT 5 6 Luna | OpenAI | 0.65 | −96% | −99% | 26 |
| GPT 5 6 Terra Pro | OpenAI | 0.65 | −96% | −99% | 26 |
| Gemini 3 5 Flash Lite | 0.65 | −94% | −99% | 25 | |
| Claude Sonnet 5 | Anthropic | 0.65 | −89% | −99% | 88 |
| GPT 5 6 Luna Pro | OpenAI | 0.7 | −97% | −99% | 26 |
| Gemini 3 1 Flash Lite | 0.75 | −96% | −99% | 81 |
3. Where the witness over-certifies
The chips on the landing, as a table: the calibration items where recon would pass a scale the model fails. The certificate pulls
back on those, so an over-certified item is a tighter band, not a wrong certificate.
| model | witness | tolerance | over-certified items | which |
|---|---|---|---|---|
| Llama 4 Scout | recon |
0.5 | 7 | tiny_code, which_color, dense_config, small_label_number |
| Grok 4 5 | recon |
0.5 | 8 | count_circles, tiny_code, dense_config, small_label_number |
| Qwen3 7 Plus | recon |
0.5 | 5 | tiny_code, dense_config, small_label_number, subscript_number |
| Llama 3.2 11B | recon |
0.5 | 6 | tiny_code, which_color, dense_config, small_label_number |
| Claude Sonnet 4.6 | recon |
0.6 | 6 | tiny_code, which_color, dense_config, small_label_number |
| Claude Fable 5 | recon |
0.6 | 7 | count_circles, tiny_code, which_color, legend_lookup |
| Claude Sonnet 5 | recon |
0.65 | 7 | count_circles, tiny_code, bar_value, dense_config |
| GPT-5.4 mini | recon |
0.7 | 7 | tiny_code, legend_lookup, dense_config, small_label_number |
| Gemini 3.5 Flash | recon |
0.75 | 6 | count_circles, tiny_code, dense_config, small_label_number |
| GPT-4o mini | recon |
0.75 | 3 | tiny_code, dense_config, small_label_number |
| Gemini 3 1 Flash Lite | recon |
0.8 | 4 | tiny_code, dense_config, small_label_number, subscript_number |
4. The witness bake-off
On the render corpus the witnesses were raced for safety (band-ratio geomean; 1.0 = never over-certifies): RMSE over-certifies its blind
spots about five times as often as the gradient-energy witness, which is why recon is primary and RMSE is kept for coarse images only.
The full table, per item and per model, is in the calibration reports (the public corpus report).
5. How a band is replayed
A record keeps the two images’ hashes and dimensions, the witness, its score, the model’s band and the verdict, never the pixels.
Anyone holding the two images can run the witness again and compare; anyone at all can check the seal and the signature at
/v1/verify/<record_sha256> (the certify walkthrough does both). The record this site’s walkthrough
issued: f164152c1dd555641a29b7fb2a297fc54c3d805114b426c3f3ea40983fd7bc1a.