Skip to Sparsification lesson
LLM Inference Visualizer

Model Optimization · Lesson 04

Sparsification

If many weights matter little, can the GPU skip them?

From the previous lesson

Carry forward: Quantization

Map weights and activations into lower-precision formats, inspect scaling and error, and connect reduced storage to memory traffic and supported GPU kernels.

Why this lesson comes next

If many weights contribute little, can we skip them?

Compare irregular zeros with hardware-friendly sparse patterns and reveal why a sparse model becomes faster only when its storage format and kernels agree.
Terms, if you need themPruning · Unstructured · 2:4 sparsity
Pruning
Setting selected model weights to zero according to an importance rule.
Unstructured
Zeros may appear anywhere, which is flexible but harder for hardware to accelerate.
2:4 sparsity
Exactly two of every four eligible weights remain non-zero, creating a predictable hardware-friendly pattern.

Keep in mind: A high zero count alone does not guarantee lower latency; quality recovery, sparse storage overhead, and compatible kernels still matter.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementA dense matrix stores every weight

Every matrix position consumes memory and participates in the multiplication.

BottleneckUnnecessary parameters

Remove weak connections, then encode the remaining pattern so hardware can truly skip work.

Metric to watchCompute + quality
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Dense

A dense matrix stores every weight

Every matrix position consumes memory and participates in the multiplication.

Live diagram

Zero weights are useful only when hardware can skip them

Teaching diagram · illustrative values unless marked as measured
Lesson 04 · 14 total

Dense kernels are regular and highly optimized.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
Dense100%Weights retained
2:450%Structured non-zeros
QualityRecheckAfter pruning

Actual benchmark

A supported 2:4 kernel sped up real BERT-Large layers

Named model · measured workload
Sparse matrix-multiply speedup× dense cuBLAS on A100
A supported 2:4 kernel sped up real BERT-Large layersDense BERT layers: 1×; 2:4 sparse layers: up to 1.6×Dense BERT layers1×2:4 sparse layersup to 1.6×
Model
BERT-Large layer shapes
Hardware
NVIDIA A100 · cuSPARSELt sparse Tensor Cores
Benchmark
FP16 2:4 sparse GEMM layer benchmarks

What this result shows · NVIDIA measured up to 1.6× speedup over dense cuBLAS for BERT-Large matrix shapes when the 2:4 layout and sparse kernel were both used.

What it does not prove · This is a layer-level kernel result, not a full LLM serving gain; non-matrix operations and unsupported shapes reduce end-to-end speedup.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA cuSPARSELt benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why might 50% zero weights produce almost no speedup?
02What makes 2:4 sparsity easier for supported hardware to accelerate?
03Two checkpoints are both 50% sparse. One uses random zeros; one uses supported 2:4 blocks. Which evaluation can establish the useful winner?
Answer all three, then check your reasoning.

What we can now explain

Sparsification, in one sentence

Compatible kernels can reduce memory traffic and matrix-multiply work.

Next chapter

A new bottleneck now comes into view

Can uneven requests share the GPU efficiently?Continue to Continuous Batching
Back to Quantization