Measured workloads, fewer billed tokens
Tidbit publishes four retained results from committed raw evidence. They cover two outcome families: billed input-token reduction on tool-heavy workloads and billed-cost reduction when stable context is reused. They are examples from measured runs, not estimates for a different call or session.
What the lossless profile measured
One 197,521-token tool-heavy request used 14% fewer billed input tokens. The recorded answer sentinel passed in both compared calls.
Sending the same stable context four times reduced cumulative billed cost by 61% after one initial send and three reuses. Input and output token counts matched across all four measured turns.
A pre-warmed 24-turn window with repeated stable context reduced cumulative billed cost by 88%. Input and output token counts matched across all 24 measured turns. This result depends on reuse and does not describe one-off traffic.
What those checks do not prove
The lossless-profile checks describe those source runs only. A passing sentinel or matching token counts in one workload is not a general quality or latency guarantee.
A separate off-by-default profile reduced cumulative billed input tokens by 71% to 80% at turns 8 through 14 of one growing tool-heavy session. Fresh-answer key checks passed on all 14 turns, but old-context recall was not tested.
Read the evidence by workload
Each public value links to an evidence ID, a SHA-256 fingerprint, and a checked-in derivation. Compare your workload with the measured shape before using a result as a planning input. Actual savings depend on request size, session length, and repeated stable context.