From the previous lesson
Model Optimization · Lesson 04
Sparsification
If many weights matter little, can the GPU skip them?
Why this lesson comes next
If many weights contribute little, can we skip them?
Compare irregular zeros with hardware-friendly sparse patterns and reveal why a sparse model becomes faster only when its storage format and kernels agree.Terms, if you need themPruning · Unstructured · 2:4 sparsity+
- Pruning
- Setting selected model weights to zero according to an importance rule.
- Unstructured
- Zeros may appear anywhere, which is flexible but harder for hardware to accelerate.
- 2:4 sparsity
- Exactly two of every four eligible weights remain non-zero, creating a predictable hardware-friendly pattern.
Keep in mind: A high zero count alone does not guarantee lower latency; quality recovery, sparse storage overhead, and compatible kernels still matter.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.Every matrix position consumes memory and participates in the multiplication.
Remove weak connections, then encode the remaining pattern so hardware can truly skip work.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Dense
A dense matrix stores every weight
Every matrix position consumes memory and participates in the multiplication.Live diagram
Zero weights are useful only when hardware can skip them
Teaching diagram · illustrative values unless marked as measuredDense kernels are regular and highly optimized.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
A supported 2:4 kernel sped up real BERT-Large layers
- Model
- BERT-Large layer shapes
- Hardware
- NVIDIA A100 · cuSPARSELt sparse Tensor Cores
- Benchmark
- FP16 2:4 sparse GEMM layer benchmarks
What this result shows · NVIDIA measured up to 1.6× speedup over dense cuBLAS for BERT-Large matrix shapes when the 2:4 layout and sparse kernel were both used.
What it does not prove · This is a layer-level kernel result, not a full LLM serving gain; non-matrix operations and unsupported shapes reduce end-to-end speedup.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA cuSPARSELt benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
Sparsification, in one sentence
Compatible kernels can reduce memory traffic and matrix-multiply work.Next chapter