The blog
Crumbs & notes.
Short reads on measured workload savings and evidence limits.
- August 2026
Batching and prompt engineering, in that order
A batch endpoint halves the rate on work nobody is waiting for, and it is a config change. Prompt work pays too, but it is smaller and easier to fool yourself about.
Read → - August 2026
Where your LLM bill actually goes
Most of an LLM bill is spent re-reading text the model has already seen. Seven straightforward ways to cut it, most of them arithmetic rather than engineering.
Read → - July 2026
Measured workloads, fewer billed tokens
Four tracked runs show workload-specific reductions in billed input tokens or billed cost. The fidelity statements are limited to checks recorded by the lossless profile.
Read →