Skip to KV Cache lesson
LLM Inference Visualizer

Memory Model · Lesson 01

KV Cache

What are Keys and Values, and why does saving them make generation faster?

From the previous chapter

Carry forward: Why Local Models + SLOs

Choose between hosted, self-hosted, and on-device inference, then define TTFT, token pace, end-to-end latency, goodput, reliability, and quality targets for that deployment.

Why this lesson comes next

Why does faster decoding create a growing memory bill?

Reuse past keys and values instead of recomputing the entire prefix, then watch cache memory grow across layers, tokens, and active requests.
Terms, if you need themKey (K) · Value (V) · ITL
Key (K)
A learned label used to decide whether a token is relevant to the current Query.
Value (V)
The learned content mixed into the attention result when its Key receives weight.
ITL
Inter-token latency: the delay between consecutive streamed output tokens.

Keep in mind: KV caching removes repeated K/V projection work, but attention still reads the growing history and the cache consumes more memory for every token and request.

01

Current setup & bottleneck

See why decoding repeats work

First understand the attention values that the uncached path rebuilds.
Before the improvementWhy does faster decoding create a growing memory bill?

Reuse past keys and values instead of recomputing the entire prefix, then watch cache memory grow across layers, tokens, and active requests.

BottleneckRepeated decode compute

Save the unchanging past so each decode step computes only the newest token.

Metric to watchITL + memory

Meet K and V inside self-attention

Every token creates three learned vectors. Think of them as a question, a label, and the content behind that label.

Step 1

Turn the current token into three vectors

Layer inputVector for “fox”x₄
q_projQ₄ = x₄WQ
k_projK₄ = x₄WK
v_projV₄ = x₄WV
Q₄What do I want to find?
K₄What kind of information do I contain?
V₄What information do I carry?
Step 2

Compare Q₄ with every Key

The QueryQ₄[0.8, 0.2, 0.6]
K · Keys
K₁K₂K₃K₄
one dot product per token
Dot productsQ₄Kᵀ = [0.2, 3.1, 1.4, 2.7]softmax[0.04, 0.52, 0.10, 0.34]attention weights
Step 3

Use the scores to mix the Values

V · Values
V₁V₂V₃V₄
.04.52.10.34
Weighted sum.04V₁ + .52V₂ + .10V₃ + .34V₄Context vector
o_projNew vector for “fox”

The numbers are a small teaching example. In a real model, Q, K, and V are much longer vectors computed independently in every attention layer.

02

Improvement visualization

Save the past, then append one new row

Compare the uncached and cached decode paths side by side.

Live comparison

One decode step, two execution paths

Request timelinePrefill
PromptGeneratedNext token
Causal self-attention

Each token sees itself and the past

Rows are Queries; columns are Keys. The diagonal is self-attention, and future tokens are masked.

Without cache

Replay the prefix

Repeated work

All prompt K/V projections are required.

Create Keys (K)4 token rows
Create Values (V)4 token rows
K/V rows created this step8
With KV cache

Reuse, then append

Saved state

Prefill writes all prompt K/V rows once.

Saved Keys (K)4 resident rows
Saved Values (V)4 resident rows
New K/V rows this step8
During prefill, both paths compute K and V for every prompt token. The cache earns its benefit on later decode steps.
03

How the numbers usually move

Trade repeated projection work for growing memory

Use the calculator and published comparison to connect context length with serving cost.
TTFTMostly unchangedPrefill still runs
ITLLowerPast K/V is reused
ThroughputHigherLess repeated projection
MemoryHigherLinear cache growth
QualityUnchangedExact stored values

Actual benchmark

Reusing stored KV avoided rebuilding a real Llama context

Named model · measured workload
Time-to-first-token speedup× recomputing the prompt KV
Reusing stored KV avoided rebuilding a real Llama contextRecompute KV: 1×; Reload cached KV: 5–14×Recompute KV1×Reload cached KV5–14×
Model
Llama 3 70B
Hardware
NVIDIA H100 with x86 host over PCIe
Benchmark
Multiturn reuse · 1,024–7,168 input tokens

What this result shows · NVIDIA measured 5–14× faster TTFT when a saved Llama 3 context was reloaded instead of recomputed from scratch.

What it does not prove · The lesson first shows reuse inside one response. This benchmark measures the same stored K/V state reused across turns, where transfer cost is also present.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Llama 3 KV-cache benchmark

Past keys and values become persistent state

  • Past K and V are created once, then saved.
  • Each decode step adds only the newest token.

The model still attends over the full cached context

  • Prefill still processes the whole prompt.
  • The Query still checks all visible Keys.

Saved compute becomes a capacity problem

  • Longer contexts need more cache memory.
  • More simultaneous users multiply that cost.

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why can earlier Keys and Values be reused during decoding?
02Which cost still grows even when KV projections are cached?
03A server holds 20 equal-length requests and its KV region is nearly full. Traffic rises to 40 active requests with the same model and context length. What should you expect?
Answer all three, then check your reasoning.

What we can now explain

KV Cache, in one sentence

Reuse the immutable past to avoid repeated K/V projections, while remembering that cache memory and attention reads still grow with every token and request.

Next lesson

A new bottleneck now comes into view

Where do weights, KV cache, activations, and workspaces fit?Continue to GPU Memory
Back to Why Local Models + SLOs