Map logical KV blocks onto non-contiguous GPU pages to reduce fragmentation, batch requests efficiently, and share cached prefixes.
vLLM · Kwon et al. 2023