teaching_llm_applications

Paged attention and vLLM (virtual LLM)

Break up the KV cache into small pieces

image

image