From the previous lesson
Inference Optimization · Lesson 07
Prefix Caching
Why recompute the same system prompt for every user?
Why this lesson comes next
Why recompute the same system prompt for every user?
Reuse content-addressed KV blocks for exact shared prefixes, then examine cache hits, eviction, tenant boundaries, and cache-aware routing.Terms, if you need themPrefix · Cache hit · Eviction+
- Prefix
- The sequence of tokens at the beginning of a request.
- Cache hit
- A request finds compatible stored KV blocks for part of its exact prefix.
- Eviction
- Removing cached blocks to free memory for more useful or newer state.
Keep in mind: Prefix caching helps only when exact prefixes repeat and reach a worker that still holds them; unique prompts remain full prefills.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.System prompts, documents, and few-shot examples often repeat exactly.
If two requests begin with exactly the same tokens, reuse the completed prefill state for their shared blocks.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Repeat
Many requests begin with the same tokens
System prompts, documents, and few-shot examples often repeat exactly.Live diagram
Reuse exact prompt blocks across requests
Teaching diagram · illustrative values unless marked as measuredThe fingerprint includes the exact token block, the prior-prefix fingerprint, and compatible model/cache settings.
Without reuse, every request performs the same prefill work.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Prefix reuse can preserve throughput beyond HBM capacity
- Model
- OLMo-Hybrid-7B, Qwen3.5-4B, Qwen3.6-27B-FP8
- Hardware
- NVIDIA H100 GPUs
- Benchmark
- LongBench + RULER; branching traffic beyond HBM
What this result shows · On branching hybrid-LLM workloads beyond HBM capacity, SuffixReplay reports 2.3–4.3× SGLang throughput and 15–70% lower median TTFT.
What it does not prove · This is a recent hybrid-LLM prefix-caching result; gains depend on repeated prefixes, cache pressure, and routing locality.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · SuffixReplay paper04 · Quiz
Can you reason through the concept?
What we can now explain
Prefix Caching, in one sentence
Cache hits reduce TTFT and free prefill capacity for new work.Next lesson