From the previous chapter
Model Optimization · Lesson 03
Quantization
How can the same model values occupy fewer bytes?
Why this lesson comes next
How can the same model values occupy fewer bytes?
Map weights and activations into lower-precision formats, inspect scaling and error, and connect reduced storage to memory traffic and supported GPU kernels.Terms, if you need themScale · Weight-only · W8A8+
- Scale
- A value that maps a group of real numbers onto a smaller set of quantized levels.
- Weight-only
- Weights use fewer bits while intermediate activations remain in higher precision.
- W8A8
- Both weights and the inputs to eligible linear layers use 8-bit formats during matrix multiplication.
Keep in mind: Quantization reduces storage and traffic, but it does not remove the KV cache, fix scheduling, or guarantee acceptable accuracy on every model.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.BF16 uses 16 bits for each weight and provides a wide numerical range.
Store and move model numbers with fewer bits, then check whether kernels and quality still meet the goal.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Baseline
High precision stores more detail
BF16 uses 16 bits for each weight and provides a wide numerical range.Live diagram
What quantizes, and how much memory it saves
Teaching diagram · illustrative values unless marked as measuredINT and FP formats can use the same bit width but encode numbers differently. Actual GPU count is higher when KV cache, workspace, activations, and serving headroom are included.
More bits preserve detail, but increase model size and memory traffic.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
W8A8 nearly halved measured OPT-175B memory
- Model
- OPT-175B
- Hardware
- NVIDIA A100-80GB GPUs
- Benchmark
- FasterTransformer decode benchmark · batch 1 · sequence length 128
What this result shows · SmoothQuant reduced measured memory from 369 GB to 182 GB while keeping W8A8 accuracy close to FP16 in the paper’s evaluation.
What it does not prove · This is W8A8 on a supported INT8 path. Different bit-widths, kernels, batch sizes, and models can change both latency and quality.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · SmoothQuant, ICML 202304 · Quiz
Can you reason through the concept?
What we can now explain
Quantization, in one sentence
Smaller models can fit on fewer GPUs and move less data per token.Next lesson