How savings are measured
3 min readUpdated July 2026
Each retained public benchmark value is recalculated from a committed raw result and carries an evidence ID, source fingerprint, and checked-in derivation.
What we compare against#
- One 197,521-token tool-heavy request used 14% fewer billed input tokens in one live lossless-profile run.
- The same stable context sent four times reduced cumulative billed cost by 61% in one live four-send run.
- One pre-warmed 24-turn window with repeated stable context reduced cumulative billed cost by 88%.
- Turns 8 through 14 of one growing tool-heavy session reduced cumulative billed input tokens by 71% to 80% in an off-by-default profile.
The lossless-profile fidelity statements are limited to checks recorded in their source runs. The off-by-default long-session result passed fresh-answer key checks, but old-context recall was not tested. Reuse-dependent values do not describe one-off traffic.
If a workload is not saving you anything, the console shows an honest zero rather than a flattering estimate.