Blog
Writing from s2ar.dev — reward integrity (starscry) and certified image compression (starlens): one measurement discipline, every claim replayable.
The Refusal Nobody Measures
The industry counts how often models comply with harmful requests and barely counts how often they refuse legitimate ones. Over-refusal is a real alignment failure — and its headline metric is a self-graded scalar that misranks its own subjects. The case for measuring it like a calibration.
Read →Why a Certificate, Not a Compressor
Images bill by dimensions, not bytes — and every existing answer to 'how far can I shrink this?' is an aggregate. The case for a per-request, per-model, sealed fidelity certificate, priced as a share of the savings it proves.
Read →Your JPEG Quality Knob Saves 0%
We requantized to quality 25 and measured the token bill: exactly 0.0% saved. The knob the industry ships optimizes a number vision-LLMs don't bill on — here is what they bill on instead, per provider.
Read →The Average Is Not a Guarantee
A pixel-error metric calls a 32-pixel-wide image 'identical' while every model we tested stops reading at six times that size. On held-out content the naive policy silently loses 42–50% of answers — while reporting the best savings number on the board.
Read →Certificates Name Their Model
A compression that one model reads perfectly loses another 29% of its answers at the same rung. Cross-model transfer is one-directional — which is why a fidelity guarantee that doesn't name its reader isn't one.
Read →Don't Trust Me — Re-run the Seed
A public RL environment whose reward a seeded search games in one simulation — and why the fix is to construct reward-hacks deterministically and ship the apparatus with the claim, not to detect them with a bigger model.
Read →Run the Checker: Replayable Audits, Difficulty Certificates, and the Capability Floor
The platform plus a season of pointing the instruments at my own hypotheses: three refutations with named mechanisms, and two results — a capability floor and a solver-effort certificate — that survived the same guillotine.
Read →The Reward Integrity Index, in brief
The accessible on-ramp: what reward-hacking is, the six deterministic instruments, and what the live board of 32 audited environments found — every number re-runnable from its seed.
Read the introduction →Reward Integrity: deterministic, replayable audits of RL-environment rewards
The rigorous treatment — the epiplexity frame, the difficulty-certification results (−0.924 / −0.870 vs pass-rate, +0.872 vs oracle), the 1372× decomposition collapse, the Reward Integrity Score, and an open research agenda. Heavily cited, annotated in the margins.
Read the paper →Action-space decomposition — a live explorer
The flagship result, interactive: a flat search explodes 7 → 23 → 469 → 1372 simulations on a rule-assembly puzzle while an LM-proposed macro holds 1 — play the live Rust→WASM engine in your browser.
Open the explorer →The paper, reproduced as runnable code
Fourteen hermetic Rust prototypes — thirteen reproduce the paper end-to-end (every instrument, the Index, the ExiT story, the epiplexity theory), the fourteenth extends it — each asserting its claim.
See the prototypes →More posts soon. RSS.