GKE Pod Snapshots Cut Model Load Times, and Move the Work to Snapshot Lifecycle Management
Cold starts are the tax nobody budgets for on GPU workloads. GKE Pod Snapshots checkpoint a running pod - GPU memory, threads, filesystem - and restore it without re-running init: a 70B model back in 37 seconds, an 8B in 15, and Codeway went from ~60s to 8s. Up to 89% off startup latency. The honest caveat in the writeup: the hard part isn’t capture, it’s snapshot invalidation and upgrades.