Mooncake: Turning KV Cache into a Distributed LLM Serving Layer
Mooncake reframes KV cache as a first-class distributed resource. This post walks through its architecture, the KV store design, and why prefilling and decoding benefit so differently from cache disaggregation.