Skip to Prefix Caching lesson
LLM Inference Visualizer

Inference Optimization · Lesson 07

Prefix Caching

Why recompute the same system prompt for every user?

From the previous lesson

Carry forward: PagedAttention

Map logical cache blocks onto non-contiguous physical pages to reduce reservation waste, fragmentation, and expensive movement as sequences grow.

Why this lesson comes next

Why recompute the same system prompt for every user?

Reuse content-addressed KV blocks for exact shared prefixes, then examine cache hits, eviction, tenant boundaries, and cache-aware routing.
Terms, if you need themPrefix · Cache hit · Eviction
Prefix
The sequence of tokens at the beginning of a request.
Cache hit
A request finds compatible stored KV blocks for part of its exact prefix.
Eviction
Removing cached blocks to free memory for more useful or newer state.

Keep in mind: Prefix caching helps only when exact prefixes repeat and reach a worker that still holds them; unique prompts remain full prefills.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementMany requests begin with the same tokens

System prompts, documents, and few-shot examples often repeat exactly.

BottleneckRepeated prefill

If two requests begin with exactly the same tokens, reuse the completed prefill state for their shared blocks.

Metric to watchTTFT + throughput
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Repeat

Many requests begin with the same tokens

System prompts, documents, and few-shot examples often repeat exactly.

Live diagram

Reuse exact prompt blocks across requests

Teaching diagram · illustrative values unless marked as measured
Lesson 07 · 14 total

Without reuse, every request performs the same prefill work.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
Cache hitSkip prefillFor matched blocks
TTFTLowerMore prefix reused
ThroughputHigherLess repeated work

Actual benchmark

Prefix reuse can preserve throughput beyond HBM capacity

Named model · measured workload
Relative throughput when the working set exceeds HBM× SGLang
Prefix reuse can preserve throughput beyond HBM capacitySGLang baseline: 1×; SuffixReplay prefix caching: 2.3–4.3×SGLang baseline1×SuffixReplay prefix caching2.3–4.3×
Model
OLMo-Hybrid-7B, Qwen3.5-4B, Qwen3.6-27B-FP8
Hardware
NVIDIA H100 GPUs
Benchmark
LongBench + RULER; branching traffic beyond HBM

What this result shows · On branching hybrid-LLM workloads beyond HBM capacity, SuffixReplay reports 2.3–4.3× SGLang throughput and 15–70% lower median TTFT.

What it does not prove · This is a recent hybrid-LLM prefix-caching result; gains depend on repeated prefixes, cache pressure, and routing locality.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · SuffixReplay paper

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why is semantic similarity insufficient for a prefix-cache hit?
02What latency component improves most when a long prefix is reused?
03Worker A has the exact cached system prompt but a slightly longer queue; Worker B is idle with no matching prefix. Which policy can improve TTFT without blindly choosing either?
Answer all three, then check your reasoning.

What we can now explain

Prefix Caching, in one sentence

Cache hits reduce TTFT and free prefill capacity for new work.

Next lesson

A new bottleneck now comes into view

Can one target pass verify several tokens?Continue to Speculative Decoding
Back to PagedAttention