Skip to PagedAttention lesson
LLM Inference Visualizer

Inference Optimization · Lesson 06

PagedAttention

How can KV caches grow without one large contiguous allocation?

From the previous lesson

Carry forward: Continuous Batching

Admit and retire sequences at iteration boundaries so requests with different arrival times and output lengths can continuously share the GPU.

Why this lesson comes next

How can variable KV caches grow without large contiguous regions?

Map logical cache blocks onto non-contiguous physical pages to reduce reservation waste, fragmentation, and expensive movement as sequences grow.
Terms, if you need themLogical block · Physical page · Fragmentation
Logical block
A consecutive group of tokens as the request sees them.
Physical page
A fixed-size slot in GPU memory that can hold one KV block.
Fragmentation
Memory lost inside oversized reservations or split into gaps that are hard to reuse.

Keep in mind: PagedAttention improves KV allocation, not attention math; block tables and partially filled final pages still have some overhead.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementReservation creates two kinds of fragmentation

Unused slots inside Request A are internal fragmentation. Free slots split into small gaps are external fragmentation.

BottleneckKV fragmentation

Keep token order logical while letting the KV cache occupy any free physical pages.

Metric to watchCapacity + batching
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Fragment

Reservation creates two kinds of fragmentation

Unused slots inside Request A are internal fragmentation. Free slots split into small gaps are external fragmentation.

Live diagram

From fragmented reservations to block-at-a-time allocation

Teaching diagram · illustrative values unless marked as measured
Lesson 06 · 14 total

The GPU may have enough free memory in total, but not enough adjacent memory for the next request.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
AllocationOn demandBlock by block
WasteLowerMostly final blocks
Batch sizeLargerMore KV fits

Actual benchmark

PagedAttention turns reclaimed KV memory into throughput

Named model · measured workload
Relative throughput at the same latency× prior systems
PagedAttention turns reclaimed KV memory into throughputFasterTransformer / Orca: 1×; vLLM with PagedAttention: 2–4×FasterTransformer / Orca1×vLLM with PagedAttention2–4×
Model
OPT-13B, OPT-66B, OPT-175B
Hardware
1×, 4×, and 8× NVIDIA A100 GPUs
Benchmark
ShareGPT and Alpaca request-length traces

What this result shows · The vLLM paper reports 2–4× higher throughput at the same latency, with larger gains on longer sequences and larger models.

What it does not prove · The range is the paper’s end-to-end serving result, including vLLM’s PagedAttention-based memory manager.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · PagedAttention paper, SOSP 2023

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01What does the block table separate?
02What happens when one generated token starts a new logical block?
03A request ends after storing 9 tokens with four-token KV blocks. What waste remains even with paging?
Answer all three, then check your reasoning.

What we can now explain

PagedAttention, in one sentence

Near-demand allocation supports larger batches and safe block sharing.

Next lesson

A new bottleneck now comes into view

Can identical prompt blocks be reused?Continue to Prefix Caching
Back to Continuous Batching