ALEXKRI.NET
Resume Contact
← Back

GKE Pod Snapshots Cut Model Load Times, and Move the Work to Snapshot Lifecycle Management

Cold starts are the tax nobody budgets for on GPU workloads. GKE Pod Snapshots checkpoint a running pod - GPU memory, threads, filesystem - and restore it without re-running init: a 70B model back in 37 seconds, an 8B in 15, and Codeway went from ~60s to 8s. Up to 89% off startup latency. The honest caveat in the writeup: the hard part isn’t capture, it’s snapshot invalidation and upgrades.

Read the source ↗