Audit a reward

Grade a reward for gaming without a model. audit runs four deterministic instruments against it: content-free completions that must not pass, cosmetic re-renderings of a passing answer that must not move its score, known-wrong and equivalent answers, and a seeded search for a constructed exploit. The job runs in the background; the answer is a graded record (A pass, C partial, F fail, ERR measured) with a reward-integrity score. From here you audit one of our own environments’ rewards, free once a day; an agent audits its own grader with grader_url and its rows.

An agent calls POST https://api.s2ar.dev/v1/audit with {"env": "baba"}, or {"grader_url": "https://…", "rows": [{"prompt": …, "answer": …}], "pass_threshold": 0.5} for its own scorer, then polls GET /v1/audit/{job}; over MCP the tool is audit_reward. The worked version is Audit a reward.