Diagram of disaggregated LLM serving with a shared KV cache pool across prefill and decode nodes.

Mooncake: Turning KV Cache into a Distributed LLM Serving Layer

Mooncake reframes KV cache as a first-class distributed resource. This post walks through its architecture, the KV store design, and why prefilling and decoding benefit so differently from cache disaggregation.

September 5, 2026 · 13 min · 2716 words · martinuke0
Feedback