ProjectsInference
How LLMs work, the KV cache visualizer
An animated, open-source explainer of LLM inference and the KV cache.
- Year
- 2026
- Area
- Inference
- Built with
- Open source
I built this while studying inference infrastructure: KV-cache offloading, PagedAttention, and disaggregated prefill and decode.
Reading about the KV cache only got me so far. Animating it step by step is what made it click, so I turned it into something other people can learn from too.