tidbit
← All posts

Measured workloads, fewer billed tokens

Tidbit publishes four retained results from committed raw evidence. They cover two outcome families: billed input-token reduction on tool-heavy workloads and billed-cost reduction when stable context is reused. They are examples from measured runs, not estimates for a different call or session.

What the lossless profile measured

One 197,521-token tool-heavy request used 14% fewer billed input tokens. The recorded answer sentinel passed in both compared calls.

Sending the same stable context four times reduced cumulative billed cost by 61% after one initial send and three reuses. Input and output token counts matched across all four measured turns.

A pre-warmed 24-turn window with repeated stable context reduced cumulative billed cost by 88%. Input and output token counts matched across all 24 measured turns. This result depends on reuse and does not describe one-off traffic.

What those checks do not prove

The lossless-profile checks describe those source runs only. A passing sentinel or matching token counts in one workload is not a general quality or latency guarantee.

A separate off-by-default profile reduced cumulative billed input tokens by 71% to 80% at turns 8 through 14 of one growing tool-heavy session. Fresh-answer key checks passed on all 14 turns, but old-context recall was not tested.

Read the evidence by workload

Each public value links to an evidence ID, a SHA-256 fingerprint, and a checked-in derivation. Compare your workload with the measured shape before using a result as a planning input. Actual savings depend on request size, session length, and repeated stable context.