Blog

Writing from s2ar.dev — reward integrity (starscry) and certified image compression (starlens): one measurement discipline, every claim replayable.

starscry · Essay

The Refusal Nobody Measures

The industry counts how often models comply with harmful requests and barely counts how often they refuse legitimate ones. Over-refusal is a real alignment failure — and its headline metric is a self-graded scalar that misranks its own subjects. The case for measuring it like a calibration.

Read →
starlens · Essay

Why a Certificate, Not a Compressor

Images bill by dimensions, not bytes — and every existing answer to 'how far can I shrink this?' is an aggregate. The case for a per-request, per-model, sealed fidelity certificate, priced as a share of the savings it proves.

Read →
starlens · Measurement note

Your JPEG Quality Knob Saves 0%

We requantized to quality 25 and measured the token bill: exactly 0.0% saved. The knob the industry ships optimizes a number vision-LLMs don't bill on — here is what they bill on instead, per provider.

Read →
starlens · Measurement note

The Average Is Not a Guarantee

A pixel-error metric calls a 32-pixel-wide image 'identical' while every model we tested stops reading at six times that size. On held-out content the naive policy silently loses 42–50% of answers — while reporting the best savings number on the board.

Read →
starlens · Measurement note

Certificates Name Their Model

A compression that one model reads perfectly loses another 29% of its answers at the same rung. Cross-model transfer is one-directional — which is why a fidelity guarantee that doesn't name its reader isn't one.

Read →
starscry · Essay

Don't Trust Me — Re-run the Seed

A public RL environment whose reward a seeded search games in one simulation — and why the fix is to construct reward-hacks deterministically and ship the apparatus with the claim, not to detect them with a bigger model.

Read →
starscry · Research note

Run the Checker: Replayable Audits, Difficulty Certificates, and the Capability Floor

The platform plus a season of pointing the instruments at my own hypotheses: three refutations with named mechanisms, and two results — a capability floor and a solver-effort certificate — that survived the same guillotine.

Read →
Introduction

The Reward Integrity Index, in brief

The accessible on-ramp: what reward-hacking is, the six deterministic instruments, and what the live board of 32 audited environments found — every number re-runnable from its seed.

Read the introduction →
Technical paper

Reward Integrity: deterministic, replayable audits of RL-environment rewards

The rigorous treatment — the epiplexity frame, the difficulty-certification results (−0.924 / −0.870 vs pass-rate, +0.872 vs oracle), the 1372× decomposition collapse, the Reward Integrity Score, and an open research agenda. Heavily cited, annotated in the margins.

Read the paper →
Live demo

Action-space decomposition — a live explorer

The flagship result, interactive: a flat search explodes 7 → 23 → 469 → 1372 simulations on a rule-assembly puzzle while an LM-proposed macro holds 1 — play the live Rust→WASM engine in your browser.

Open the explorer →
Reproductions

The paper, reproduced as runnable code

Fourteen hermetic Rust prototypes — thirteen reproduce the paper end-to-end (every instrument, the Index, the ExiT story, the epiplexity theory), the fourteenth extends it — each asserting its claim.

See the prototypes →

More posts soon. RSS.