Skip to Quantization lesson
LLM Inference Visualizer

Model Optimization · Lesson 03

Quantization

How can the same model values occupy fewer bytes?

From the previous chapter

Carry forward: GPU Memory

Build the complete device-memory budget, then separate capacity limits from the bandwidth costs that often dominate token-by-token decoding.

Why this lesson comes next

How can the same model values occupy fewer bytes?

Map weights and activations into lower-precision formats, inspect scaling and error, and connect reduced storage to memory traffic and supported GPU kernels.
Terms, if you need themScale · Weight-only · W8A8
Scale
A value that maps a group of real numbers onto a smaller set of quantized levels.
Weight-only
Weights use fewer bits while intermediate activations remain in higher precision.
W8A8
Both weights and the inputs to eligible linear layers use 8-bit formats during matrix multiplication.

Keep in mind: Quantization reduces storage and traffic, but it does not remove the KV cache, fix scheduling, or guarantee acceptable accuracy on every model.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementHigh precision stores more detail

BF16 uses 16 bits for each weight and provides a wide numerical range.

BottleneckWeight storage + traffic

Store and move model numbers with fewer bits, then check whether kernels and quality still meet the goal.

Metric to watchMemory + throughput
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Baseline

High precision stores more detail

BF16 uses 16 bits for each weight and provides a wide numerical range.

Live diagram

What quantizes, and how much memory it saves

Teaching diagram · illustrative values unless marked as measured
Lesson 03 · 14 total
OPT-175B parameter arithmetic175B parameters
Weights only · benchmark below reports measured memory
Model parameter RAM
175B × 2 bytes
≈ 350 GBBaseline
BF16 arithmetic350 GB
Selected · BF16350 GB
Minimum 80 GB GPUs for weight bytes
GPU 1GPU 2GPU 3GPU 4GPU 5
5 × 80 GB

INT and FP formats can use the same bit width but encode numbers differently. Actual GPU count is higher when KV cache, workspace, activations, and serving headroom are included.

Inside one transformer blockWhat is actually quantized?
Input activationxBF16
Self-attention linear layers
q_projk_projv_projo_proj
Weights: WQ, WK, WV, WO
Feed-forward linear layers
gate_projup_projdown_proj
Weights: Wgate, Wup, Wdown
Linear inputBF16 activation
Stored matrixBF16 weight
KernelBF16 matmul
Output / accumulateBF16 or FP32
Quantized weightsQ, K, V, O projectionsGate, up, down projectionsEmbedding and LM head depend on the recipe.
Quantized activations when enabledInputs to each quantized linear layerCalibrated per tensor or per tokenCurrently kept in BF16.
Usually kept higher precisionLayerNorm · Softmax · residual addsMatmul accumulation and outputsKV-cache precision is a separate serving option.

More bits preserve detail, but increase model size and memory traffic.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
BF162 bytesPer weight
INT81 byteAbout 50% smaller
INT40.5 byteAbout 75% smaller

Actual benchmark

W8A8 nearly halved measured OPT-175B memory

Named model · measured workload
Measured inference memoryGB · sequence length 128
W8A8 nearly halved measured OPT-175B memoryFP16: 369 GB; SmoothQuant W8A8: 182 GBFP16369 GBSmoothQuant W8A8182 GB
Model
OPT-175B
Hardware
NVIDIA A100-80GB GPUs
Benchmark
FasterTransformer decode benchmark · batch 1 · sequence length 128

What this result shows · SmoothQuant reduced measured memory from 369 GB to 182 GB while keeping W8A8 accuracy close to FP16 in the paper’s evaluation.

What it does not prove · This is W8A8 on a supported INT8 path. Different bit-widths, kernels, batch sizes, and models can change both latency and quality.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · SmoothQuant, ICML 2023

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01What extra information lets low-precision integers approximate full-precision values?
02In weight-only quantization, which tensors normally remain at higher precision?
03An INT4 checkpoint is 75% smaller, but the target GPU dequantizes it into BF16 before every matrix multiply. What is the safest prediction?
Answer all three, then check your reasoning.

What we can now explain

Quantization, in one sentence

Smaller models can fit on fewer GPUs and move less data per token.

Next lesson

A new bottleneck now comes into view

Can selected weights be removed entirely?Continue to Sparsification
Back to GPU Memory