Chip Huyen explains how to cut inference costs without new hardware
Inference, not training, is where the money goes: Chip Huyen puts the training-to-inference compute ratio at 1:10 to 1:100, and higher for reasoning models. The cheapest win is prompt caching, where Claude Code reports 90-97% hit rates, then quantization, continuous batching and splitting prefill from decode. When comparing providers, check quality impact alongside cost and latency, and track goodput rather than raw throughput.