Token economics: what a call costs against what a session costs
Cody Lee Walker · s2ar · 2026-10-10 · a living page: a point per bench stamp, the prices as published
Every number here is a living number from a bench result stored beside this page (
data/bench/, the ◆ names the file by hash) or from the published OpenAPI. A bench is paired: the same tasks, the same model, one arm with the tools and one without, the headline a median of per-task deltas. MEASURED — not argued.
The market's median price per paid call is 0.01 USD ±0 ◆; ours run $0.005–$0.01. A sym-bench task costs 0.068 USD ◆ without the tools against 0.045 USD ◆ with sym (-32.8% ◆ paired); asserting an answer before returning it costs 79.2% ◆ more per task and lifts the pass rate from 100% ◆ to 100% ◆.
1. What a call costs
| route | price per call | what |
|---|---|---|
/v1/assert |
$0.005 | Assert an agent’s output |
/v1/audit |
$2.000 | Audit a reward for gaming: degenerate probes, invariance, soundness, a seeded exploit search; a graded signed record. POST |
/v1/certify |
$0.010 | Compress an image and verify a vision model still reads it |
/v1/diff |
$0.005 | Diff two screenshots |
/v1/env/reset |
$0.020 | Open an episode of a deterministic environment (baba, sokoban); step it free; the sealed, signed episode. GET ?env=baba |
/v1/extract_text |
$0.020 | Extract the text from an image |
/v1/score |
$0.005 | Grade a text against a form |
/v1/sym/find |
$0.005 | Code search by name in a GitHub repo |
/v1/sym/ls |
$0.005 | File skeleton from a GitHub repo |
/v1/sym/map |
$0.010 | GitHub repo map |
/v1/sym/read |
$0.005 | Function source from a GitHub repo |
/v1/sym/where |
$0.010 | Code search by meaning in a GitHub repo |
/v1/x402/inspect |
$0.005 | Inspect an x402 endpoint |
The market around these prices: median 0.01 USD ±0 ◆, p90 0.05 USD ±0 ◆, 79.2% ▼1.5 ◆ of listings at or under a cent (the agent markets, measured). A call is one HTTP request; an agent that pays with x402 adds one signature and no account.
2. What a session costs
A session is the unit an operator pays for: a model, its context, its tool calls, its turns. The benches run the same task list through Claude Code headless, one arm per configuration, and read each session’s cost from the transcript.
| bench | stamp | model | n per arm | plain median | with the tools | paired median Δ | note |
|---|---|---|---|---|---|---|---|
| sym | 2026-10-08T2020Z | claude-sonnet-5 | 36 | $0.0676 | mod $0.0454 · plugin $0.0585 | -32.8% (mod) | Clean run (plugin not installed in the bench scope) after the read fix: sym read accepts |
| s2ar plugin | 2026-10-10T1447Z | claude-sonnet-5 | 16 | $0.0416 (pass 88%) | $0.0600 (pass 94%) | +86.8% | asserted in session 10/16 |
| s2ar plugin | 2026-10-10T2309Z | claude-sonnet-5 | 16 | $0.0419 (pass 100%) | $0.0513 (pass 100%) | +79.2% | asserted in session 10/16 |
sym plain 0.068 · sym with the mod 0.045 · plugin plain 0.042 · plugin with the tools 0.051 · plugin pass rate 100
3. Reading the two together
A paid call is two orders of magnitude below a session: a cent against a nickel to a dime of model time. So the question is never whether a call is cheap, but whether it changes what the session does. The sym bench says a repository read through the tools cuts the session’s cost (fewer whole-file reads); the plugin bench says asserting an answer before returning it costs more per task, because the agent now does the check it used to skip, and it is right more often for it. Both are n of 16–36 per arm on one model; the paired median is the honest headline because the per-task spread is wide.
4. Method and limits
Paired medians over identical task lists; the plugin arm’s extra cost includes the tool calls’ tokens, never the calls’ prices (they
ran on the operator’s key). One model per stamp (claude-sonnet-5 on every stamp so far); a second model is a second series. The
bench harnesses: sym/bench/run.py and sym/bench-s2ar/run.py on GitHub (cody-walker/sym), with their task lists and raw results.