ALEXKRI.NET
Resume Contact
← Back

Chip Huyen explains how to cut inference costs without new hardware

Inference, not training, is where the money goes: Chip Huyen puts the training-to-inference compute ratio at 1:10 to 1:100, and higher for reasoning models. The cheapest win is prompt caching, where Claude Code reports 90-97% hit rates, then quantization, continuous batching and splitting prefill from decode. When comparing providers, check quality impact alongside cost and latency, and track goodput rather than raw throughput.

Read the source ↗