From the previous chapter
Memory Model · Lesson 01
KV Cache
What are Keys and Values, and why does saving them make generation faster?
Why this lesson comes next
Why does faster decoding create a growing memory bill?
Reuse past keys and values instead of recomputing the entire prefix, then watch cache memory grow across layers, tokens, and active requests.Terms, if you need themKey (K) · Value (V) · ITL+
- Key (K)
- A learned label used to decide whether a token is relevant to the current Query.
- Value (V)
- The learned content mixed into the attention result when its Key receives weight.
- ITL
- Inter-token latency: the delay between consecutive streamed output tokens.
Keep in mind: KV caching removes repeated K/V projection work, but attention still reads the growing history and the cache consumes more memory for every token and request.
Current setup & bottleneck
See why decoding repeats work
First understand the attention values that the uncached path rebuilds.Reuse past keys and values instead of recomputing the entire prefix, then watch cache memory grow across layers, tokens, and active requests.
Save the unchanging past so each decode step computes only the newest token.
Before the cache
Meet K and V inside self-attention
Every token creates three learned vectors. Think of them as a question, a label, and the content behind that label.
Turn the current token into three vectors
Compare Q₄ with every Key
Q₄Kᵀ = [0.2, 3.1, 1.4, 2.7]softmax[0.04, 0.52, 0.10, 0.34]attention weightsUse the scores to mix the Values
.04V₁ + .52V₂ + .10V₃ + .34V₄Context vectorThe numbers are a small teaching example. In a real model, Q, K, and V are much longer vectors computed independently in every attention layer.
Improvement visualization
Save the past, then append one new row
Compare the uncached and cached decode paths side by side.Live comparison
One decode step, two execution paths
Each token sees itself and the past
Rows are Queries; columns are Keys. The diagonal is self-attention, and future tokens are masked.
Replay the prefix
All prompt K/V projections are required.
Reuse, then append
Prefill writes all prompt K/V rows once.
How the numbers usually move
Trade repeated projection work for growing memory
Use the calculator and published comparison to connect context length with serving cost.Actual benchmark
Reusing stored KV avoided rebuilding a real Llama context
- Model
- Llama 3 70B
- Hardware
- NVIDIA H100 with x86 host over PCIe
- Benchmark
- Multiturn reuse · 1,024–7,168 input tokens
What this result shows · NVIDIA measured 5–14× faster TTFT when a saved Llama 3 context was reloaded instead of recomputed from scratch.
What it does not prove · The lesson first shows reuse inside one response. This benchmark measures the same stored K/V state reused across turns, where transfer cost is also present.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Llama 3 KV-cache benchmarkWhat changed?
Past keys and values become persistent state
- Past K and V are created once, then saved.
- Each decode step adds only the newest token.
What did not change?
The model still attends over the full cached context
- Prefill still processes the whole prompt.
- The Query still checks all visible Keys.
The boundary
Saved compute becomes a capacity problem
- Longer contexts need more cache memory.
- More simultaneous users multiply that cost.
04 · Quiz
Can you reason through the concept?
What we can now explain
KV Cache, in one sentence
Reuse the immutable past to avoid repeated K/V projections, while remembering that cache memory and attention reads still grow with every token and request.Next lesson