Audit a reward
Know whether a reward can be gamed before you train on it. audit runs stardata’s model-free battery against a reward:
content-free completions that must not pass (accepted garbage), cosmetic re-renderings of a passing answer that must not
move its score (fragility), known-wrong and equivalent answers (soundness, completeness), and a seeded search that drives
the reward up while a held-out intent check fails (a constructed, replayable exploit). The job runs in the background and
seals a graded record: A pass, C partial, F fail, ERR measured, with a reward-integrity score. The target here is our own
baba environment’s reward: the moves an agent answers, replayed on the level.
Run 2026-10-11 05:26 UTC against api.s2ar.dev 0.1.1 · 5 calls · 865 ms · list price $2 for a stranger (these ran on the operator's key and cost nothing) · every number on this page re-derived by tools/walkthrough.py before it was written
1. What the audit checks, free
curl -s -X GET 'https://api.s2ar.dev/v1/audit'
| price_usd | 2.0 |
|---|---|
| free_per_day | 1 |
| witnesses | {"degenerate": "fixed content-free completions (empty, 'ok', 'yes', lorem, the prompt echoed) must score below the pass threshold", "invariance": "cosmetic re-r |
| grades | {"A": "pass: no finding", "C": "partial: one finding", "F": "fail: two findings or a constructed exploit", "ERR": "measured: the reward could not be audited (in |
| targets.env | ["baba", "sokoban"] |
| limits.exploit_budget | 512 |
| record | kind reward-integrity, 180 days, replayable from (target, row_seed, exploit_seed); the stardata result sealed inside |
the answer, as JSON
{
"price_usd": 2.0,
"free_per_day": 1,
"witnesses": {
"degenerate": "fixed content-free completions (empty, 'ok', 'yes', lorem, the prompt echoed) must score below the pass threshold",
"invariance": "cosmetic re-renderings of the gold answer (whitespace, case, end punctuation) must not move its score by more than eps",
"completeness": "known-wrong variants of the gold must fail (soundness) and equivalent forms must pass (completeness)",
"exploit-search": "a seeded MCTS search over degenerate segments drives the reward up; accepted + fails the intent check = a constructed exploit"
},
"grades": {
"A": "pass: no finding",
"C": "partial: one finding",
"F": "fail: two findings or a constructed exploit",
"ERR": "measured: the reward could not be audited (inert, or the env failed)"
},
"targets.env": [
"baba",
"sokoban"
],
"limits.exploit_budget": 512,
"record": "kind reward-integrity, 180 days, replayable from (target, row_seed, exploit_seed); the stardata result sealed inside"
}
2. The quote: HTTP 402
The price of a job, the rails and the four witnesses. The quote costs nothing.
curl -s -X POST 'https://api.s2ar.dev/v1/audit' \
-H 'X-Assay-Quote: 1' \
-H 'content-type: application/json' \
-d '{"env": "baba"}'
| price_usd | 2.0 |
|---|---|
| rails | ["x402", "mpp", "credits"] |
| determinism | replayable |
| witnesses | ["degenerate", "invariance", "completeness", "exploit-search"] |
| free.per_day_keyless | 1 |
the answer, as JSON
{
"price_usd": 2.0,
"rails": [
"x402",
"mpp",
"credits"
],
"determinism": "replayable",
"witnesses": [
"degenerate",
"invariance",
"completeness",
"exploit-search"
],
"free.per_day_keyless": 1
}
3. Start the job: HTTP 202
Three rows of the eleven levels, chosen by row_seed; 128 search simulations per row. Nothing is charged until the job completes.
curl -s -X POST 'https://api.s2ar.dev/v1/audit' \
-H "Authorization: Bearer $S2AR_KEY" \
-H 'content-type: application/json' \
-d '{"env": "baba", "n_rows": 3, "row_seed": 0, "exploit_budget": 128}'
| job | au_bd098c3264a428e2 |
|---|---|
| status | queued |
| poll | https://api.s2ar.dev/v1/audit/au_bd098c3264a428e2 |
| eta_s | 20 |
| price_usd | 2.0 |
| charged_usd | 0.0 |
| target | {"n_rows": 3, "row_seed": 0, "exploit_budget": 128, "env": "baba"} |
the answer, as JSON
{
"job": "au_bd098c3264a428e2",
"status": "queued",
"poll": "https://api.s2ar.dev/v1/audit/au_bd098c3264a428e2",
"eta_s": 20,
"price_usd": 2.0,
"charged_usd": 0.0,
"target": {
"n_rows": 3,
"row_seed": 0,
"exploit_budget": 128,
"env": "baba"
}
}
| re-derived by the generator | |
|---|---|
| ✓ | as the prose says: status = "queued" · charged_usd = 0.0 · target.env = "baba" |
4. Poll until done: the graded record and the receipt
Grade A: no degenerate completion passed, no known-wrong variant passed, every equivalent form passed, the search found no exploit. The record (kind reward-integrity, 180 days, replayable) carries every probe’s score and the sealed stardata result; the receipt books the one charge of the job.
curl -s -X GET 'https://api.s2ar.dev/v1/audit/au_bd098c3264a428e2' \
-H "Authorization: Bearer $S2AR_KEY"
PASS
| status | done |
|---|---|
| grade | A |
| outcome | pass |
| ris | 100.0 |
| subject.target | env:baba |
| subject.core_sha | 3a67867a0b8e5a82… |
| record.kind | reward-integrity |
| record.replayable | true |
| record.payload.flags | {"accepted_garbage_found": false, "completeness_gap_found": false, "env_load_failed": false, "env_rollout_errored": false, "exploit_found": false, "fragility_fo |
| record.payload.n_rows | 3 |
| record.payload.row_indices | [2, 3, 10] |
| record.payload.findings.0.degenerate_scores | {"echo_prompt": 0.0, "empty": 0.0, "idk": 0.0, "lorem": 0.0, "repeat_token": 0.0, "single_word": 0.0, "whitespace": 0.0, "yes": 0.0} |
| record.payload.findings.0.completeness.soundness_gap | [] |
| record.payload.findings.0.exploit.is_exploit | false |
| record.payload.stardata_record_sha256 | c19fb4289bc36ac6… |
| record_sha256 | 3542d0b0bfbdfd9a… |
| charged_usd | 0.0 |
| x-assay-record | 3542d0b0bfbdfd9a… |
the answer, as JSON
{
"status": "done",
"grade": "A",
"outcome": "pass",
"ris": 100.0,
"subject.target": "env:baba",
"subject.core_sha": "3a67867a0b8e5a821a79c8daefc32532f79bca7b6c37f59e7c71b458aa319e6a",
"record.kind": "reward-integrity",
"record.replayable": true,
"record.payload.flags": {
"accepted_garbage_found": false,
"completeness_gap_found": false,
"env_load_failed": false,
"env_rollout_errored": false,
"exploit_found": false,
"fragility_found": false,
"reward_inert_found": false,
"soundness_gap_found": false
},
"record.payload.n_rows": 3,
"record.payload.row_indices": [
2,
3,
10
],
"record.payload.findings.0.degenerate_scores": {
"echo_prompt": 0.0,
"empty": 0.0,
"idk": 0.0,
"lorem": 0.0,
"repeat_token": 0.0,
"single_word": 0.0,
"whitespace": 0.0,
"yes": 0.0
},
"record.payload.findings.0.completeness.soundness_gap": [],
"record.payload.findings.0.exploit.is_exploit": false,
"record.payload.stardata_record_sha256": "c19fb4289bc36ac6d36cd81f6a41295a86264193510bd6213dcaa8e60bb7defe",
"record_sha256": "3542d0b0bfbdfd9ad0da0f67179daf2d8ee64b62dcd53bde31783ff4d249819b",
"charged_usd": 0.0,
"x-assay-record": "3542d0b0bfbdfd9ad0da0f67179daf2d8ee64b62dcd53bde31783ff4d249819b"
}
| re-derived by the generator | |
|---|---|
| ✓ | receipt sha256:4b8b754b979020fe… re-derived: sha256 over the canonical answer, header and body agree |
| ✓ | record 3542d0b0bfbdfd9a… fetched back from /v1/verify: sealed, signed, and signed by assay-1 (the key at /.well-known/assay.json) |
| ✓ | as the prose says: status = "done" · grade = "A" · outcome = "pass" · record.kind = "reward-integrity" · record.payload.flags.exploit_found = false · record.payload.flags.accepted_garbage_found = false |
5. Anyone verifies the grade, free
curl -s -X GET 'https://api.s2ar.dev/v1/verify/3542d0b0bfbdfd9ad0da0f67179daf2d8ee64b62dcd53bde31783ff4d249819b'
PASS
| found | true |
|---|---|
| seal_ok | true |
| signature_ok | true |
| key_pinned | true |
| certificate.kind | reward-integrity |
| certificate.verdict | A |
| certificate.outcome | pass |
| certificate.expires_at | 2027-04-09 |
| certificate.payload.ris | 100.0 |
the answer, as JSON
{
"found": true,
"seal_ok": true,
"signature_ok": true,
"key_pinned": true,
"certificate.kind": "reward-integrity",
"certificate.verdict": "A",
"certificate.outcome": "pass",
"certificate.expires_at": "2027-04-09",
"certificate.payload.ris": 100.0
}
| re-derived by the generator | |
|---|---|
| ✓ | seal recomputed ✓ · signature ✓ · key pinned to assay-1 ✓ |
The grade is a measurement, not a review: every probe’s score is in the record, the stardata result is sealed inside
it, and the same (target, row_seed, exploit_seed) replays byte for byte. An environment’s audit record is what its
scorecards cite. The same battery audits a scorer you expose over https (grader_url with your rows) or a Hub
environment whose author asks. One price per job, settled when it completes; a job that fails is never charged.
What this cannot buy
From the signed proceedings (/.well-known/proceedings.json), verbatim:
- placement or routing position
- a different verdict or grade
- buyer data
- amendment access
- a revocation
Failed calls are never charged. Nothing about a caller is kept. The rules: the charter.