From Memory-Hungry HNSW to Quantized SPANN: The Technical Evolution of Pinterest's Manas Platform
Good reminder that “keep the whole index in RAM” is a default, not a law. Pinterest moved Manas from in-memory HNSW to SSD-backed SPANN and saved over 40% of CPU time on production queries, with roughly 3x the QPS of DiskANN at a third of the latency for a 5% recall drop. Scalar quantization shrank HNSW indexes 59% while holding recall above 90%. Across 80 clusters and 5 billion embeddings: 20-30% off serving cost.