Skip to GPU Memory lesson
LLM Inference Visualizer

Memory Model · Lesson 02

GPU Memory

Why can a model fit on paper and still run out of memory?

From the previous lesson

Carry forward: KV Cache

Reuse past keys and values instead of recomputing the entire prefix, then watch cache memory grow across layers, tokens, and active requests.

Why this lesson comes next

Why can a model fit numerically and still fail or run slowly?

Build the complete device-memory budget, then separate capacity limits from the bandwidth costs that often dominate token-by-token decoding.
Terms, if you need themHBM / VRAM · SRAM · Headroom
HBM / VRAM
Large, high-bandwidth GPU memory that holds weights, KV cache, and active tensors.
SRAM
Small on-chip memory beside compute units, used for the tiles needed right now.
Headroom
Memory deliberately left unused so temporary allocations do not trigger an out-of-memory failure.

Keep in mind: More host DRAM can extend capacity through offload, but it cannot match HBM bandwidth; frequent host transfers can slow every token.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementWeights begin in storage and CPU DRAM

At startup the host loads the checkpoint, then copies the serving weights into GPU memory.

BottleneckCapacity + bandwidth

Capacity decides what can stay resident; bandwidth decides how fast the token loop can read it.

Metric to watchFit + utilization
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Start

Weights begin in storage and CPU DRAM

At startup the host loads the checkpoint, then copies the serving weights into GPU memory.

Live diagram

One token step through one GPU layer

Teaching diagram · illustrative values unless marked as measured
Lesson 02 · 14 total
Part AWhere model state livesHBM/VRAM holds the active serving state; SRAM holds only the tiles being computed right now.
One decode token · layer 17 of 32Watch one transformer layer use GPU memory
Fastest · smallestGPU SRAMRegisters + shared memory
W tileKV tilexttemp

Only the tiles needed right now stay beside the compute units.

Matrix engineTensor coresOne layer, one token step
Query · key · value projectionsAttention + o_projMLP projections
Large · high bandwidthGPU memory · HBM / VRAMPersistent state for the running server
Layer weightsWQKV · WO · WMLPresident
KV cacheK1:t−1 · V1:t−1read past rows
Activationxt(17)layer input
Workspacetemporary buffersavailable
Largest · slower host tierCPU DRAM
Usually outside the per-layer decode hot path
Model checkpointweights · scales · configsource for startup or reloadCommon
Offloaded model statecold weights · old KV blocksextends capacity, adds transfer latencyOptional
Pinned transfer buffersHost→device / device→hostpage-locked buffers feed GPU copiesCommon
Request + runtime statetoken IDs · queues · samplingscheduler, networking, output buffersCommon

Hot-path rule: keeping active weights and KV in HBM avoids crossing the much slower host link for every layer and generated token.

Part BWhat one decode step does in one layerThe highlight moves through reads, matrix work, attention, and write-back.
  1. 01Ready
  2. 02Fetch tiles
  3. 03Project QKV
  4. 04Attend
  5. 05Run MLP
  6. 06Write back

The token activation, this layer's weights, and the existing KV cache are resident in HBM.

CPU DRAM can hold the checkpoint, but hot layer weights should remain in HBM during generation.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
CapacityMust fitAll resident tensors
BandwidthSets paceBytes moved per token
HeadroomKeep someAvoid runtime OOM

Actual benchmark

Smaller KV state raised long-context serving throughput

Named model · measured workload
Output throughputtokens/s · 150 requests
Smaller KV state raised long-context serving throughputBF16 KV cache: 450.3 tok/s; FP8 KV cache: 517.5 tok/sBF16 KV cache450.3 tok/sFP8 KV cache517.5 tok/s
Model
Llama-3.1-8B
Hardware
1× NVIDIA H100
Benchmark
150 requests · concurrency 8 · ~20K input + ~2K output tokens

What this result shows · In vLLM’s measured run, halving KV-cache storage raised output throughput 14.9% and reduced median inter-token latency from 15.18 ms to 12.93 ms.

What it does not prove · This is a decode-heavy long-context case where KV traffic matters; short prompts or unsupported FP8 kernels can move differently.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM FP8 KV-cache benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why is model size alone insufficient for deciding whether a deployment fits?
02Which memory tier should hold active weights and KV state for the token loop?
03An 8B model fits with 5 GiB free, yet decode remains slow while tensor cores wait on data. Which change targets the likely bottleneck?
Answer all three, then check your reasoning.

What we can now explain

GPU Memory, in one sentence

Predict safe concurrency before the server reaches an out-of-memory cliff.

Next chapter

A new bottleneck now comes into view

Can the same model use fewer bytes?Continue to Quantization
Back to KV Cache