🤖 AI & Machine Learning💻 Tech & Development🎨 Design & UX
In this blog, we will learn about Prefill-Decode Disaggregation, a way of running a large language model where the reading of the prompt and the writing of the answer happen on separate machines. We will also see how an LLM answers a request in two phases, what the KV Cache is, why the two phases ne...
Sep 22, 2026