From the previous lesson
Inference Optimization · Lesson 06
PagedAttention
How can KV caches grow without one large contiguous allocation?
Why this lesson comes next
How can variable KV caches grow without large contiguous regions?
Map logical cache blocks onto non-contiguous physical pages to reduce reservation waste, fragmentation, and expensive movement as sequences grow.Terms, if you need themLogical block · Physical page · Fragmentation+
- Logical block
- A consecutive group of tokens as the request sees them.
- Physical page
- A fixed-size slot in GPU memory that can hold one KV block.
- Fragmentation
- Memory lost inside oversized reservations or split into gaps that are hard to reuse.
Keep in mind: PagedAttention improves KV allocation, not attention math; block tables and partially filled final pages still have some overhead.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.Unused slots inside Request A are internal fragmentation. Free slots split into small gaps are external fragmentation.
Keep token order logical while letting the KV cache occupy any free physical pages.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Fragment
Reservation creates two kinds of fragmentation
Unused slots inside Request A are internal fragmentation. Free slots split into small gaps are external fragmentation.Live diagram
From fragmented reservations to block-at-a-time allocation
Teaching diagram · illustrative values unless marked as measuredA owns all six slots, but only three contain KV data.
The GPU may have enough free memory in total, but not enough adjacent memory for the next request.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
PagedAttention turns reclaimed KV memory into throughput
- Model
- OPT-13B, OPT-66B, OPT-175B
- Hardware
- 1×, 4×, and 8× NVIDIA A100 GPUs
- Benchmark
- ShareGPT and Alpaca request-length traces
What this result shows · The vLLM paper reports 2–4× higher throughput at the same latency, with larger gains on longer sequences and larger models.
What it does not prove · The range is the paper’s end-to-end serving result, including vLLM’s PagedAttention-based memory manager.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · PagedAttention paper, SOSP 202304 · Quiz
Can you reason through the concept?
What we can now explain
PagedAttention, in one sentence
Near-demand allocation supports larger batches and safe block sharing.Next lesson