Naveen Vivek

ProjectsInference

How LLMs work, the KV cache visualizer

An animated, open-source explainer of LLM inference and the KV cache.

Year
2026
Area
Inference
Built with
Open source
Code
github.com/naveenvivek/howllmworksanimation

I built this while studying inference infrastructure: KV-cache offloading, PagedAttention, and disaggregated prefill and decode.

Reading about the KV cache only got me so far. Animating it step by step is what made it click, so I turned it into something other people can learn from too.