From the previous lesson
Memory Model · Lesson 02
GPU Memory
Why can a model fit on paper and still run out of memory?
Why this lesson comes next
Why can a model fit numerically and still fail or run slowly?
Build the complete device-memory budget, then separate capacity limits from the bandwidth costs that often dominate token-by-token decoding.Terms, if you need themHBM / VRAM · SRAM · Headroom+
- HBM / VRAM
- Large, high-bandwidth GPU memory that holds weights, KV cache, and active tensors.
- SRAM
- Small on-chip memory beside compute units, used for the tiles needed right now.
- Headroom
- Memory deliberately left unused so temporary allocations do not trigger an out-of-memory failure.
Keep in mind: More host DRAM can extend capacity through offload, but it cannot match HBM bandwidth; frequent host transfers can slow every token.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.At startup the host loads the checkpoint, then copies the serving weights into GPU memory.
Capacity decides what can stay resident; bandwidth decides how fast the token loop can read it.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Start
Weights begin in storage and CPU DRAM
At startup the host loads the checkpoint, then copies the serving weights into GPU memory.Live diagram
One token step through one GPU layer
Teaching diagram · illustrative values unless marked as measuredOnly the tiles needed right now stay beside the compute units.
Hot-path rule: keeping active weights and KV in HBM avoids crossing the much slower host link for every layer and generated token.
- 01Ready
- 02Fetch tiles
- 03Project QKV
- 04Attend
- 05Run MLP
- 06Write back
The token activation, this layer's weights, and the existing KV cache are resident in HBM.
CPU DRAM can hold the checkpoint, but hot layer weights should remain in HBM during generation.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Smaller KV state raised long-context serving throughput
- Model
- Llama-3.1-8B
- Hardware
- 1× NVIDIA H100
- Benchmark
- 150 requests · concurrency 8 · ~20K input + ~2K output tokens
What this result shows · In vLLM’s measured run, halving KV-cache storage raised output throughput 14.9% and reduced median inter-token latency from 15.18 ms to 12.93 ms.
What it does not prove · This is a decode-heavy long-context case where KV traffic matters; short prompts or unsupported FP8 kernels can move differently.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM FP8 KV-cache benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
GPU Memory, in one sentence
Predict safe concurrency before the server reaches an out-of-memory cliff.Next chapter