tidbit
Benchmarks

Results backed by committed raw evidence.

This page includes only values the repository can recalculate today. Each value names the workload, axis, evidence record, and fingerprint used to produce it.

Evidence policy

Retained results

Outcome values recalculated by `npm run bench:public`.
WorkloadSavedBasisQuality evidence
Large tool-heavy traceone 197,521-token tool-heavy requestEvidence large-trace-lossless, SHA-256 b2cbde3fb63b14%billed input tokens, with the answer check passing. one live lossless run.recorded answer sentinel passed in both compared calls
Repeated stable contextthe same stable context sent four timesEvidence cache-reuse-4, SHA-256 9eb27fc01b1561%cumulative billed cost after one initial send and three reuses. one live four-send run; conditional on reuse.input and output token counts matched across all 4 measured turns
Long repeated sessionone pre-warmed 24-turn window with repeated stable contextEvidence cache-session-24, SHA-256 83f492784fe988%cumulative billed cost over the measured pre-warmed window. one live pre-warmed 24-turn window; conditional on reuse.input and output token counts matched across all 24 measured turns
Long tool-heavy sessionturns 8 through 14 of one growing tool-heavy sessionEvidence long-session-8-14, SHA-256 d8efefcfe60971–80%cumulative billed input-token reduction. one live run; older context changed; off by default.fresh-answer key checks passed on all 14 turns; old-context recall was not tested

Provenance

No numeric value is emitted without an evidence id, SHA-256 fingerprint, and checked-in derivation.

Savings depend on workload. Each retained value is recalculated from a committed raw result file and carries its own evidence fingerprint. Reuse-dependent values do not apply to one-off traffic.

Questions about a retained result or its applicability to your traffic: get in touch.