Skip to LLM Compressor lesson
LLM Inference Visualizer

Implementation Details · Lesson 09

LLM Compressor

How does a full-precision checkpoint become deployment-ready?

From the previous chapter

Carry forward: Speculative Decoding

Let a cheaper draft mechanism propose a token run, verify the run in parallel, and see how acceptance rate determines whether speculation helps.

Why this lesson comes next

How does a full-precision checkpoint become deployment-ready?

Follow calibration, compression recipes, modifier application, format export, and evaluation as an offline pipeline feeding the serving engine.
Terms, if you need themRecipe · Calibration · Artifact
Recipe
A configuration describing which compression modifiers to apply and where.
Calibration
Running representative samples to observe tensor ranges or importance before transformation.
Artifact
The saved model files, metadata, scales, and configuration delivered to a serving runtime.

Keep in mind: LLM Compressor prepares and evaluates an artifact; it does not schedule live requests or guarantee that the server will meet latency goals.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementStart with a deployment target

Pick a quantization scheme supported by the serving hardware and model architecture.

BottleneckOversized artifact

Prepare the model offline: choose a recipe, calibrate if needed, transform, evaluate, and export.

Metric to watchSize + quality
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Choose

Start with a deployment target

Pick a quantization scheme supported by the serving hardware and model architecture.

Live diagram

An offline pipeline from checkpoint to compressed artifact

Teaching diagram · illustrative values unless marked as measured
Lesson 09 · 14 total

The recipe should begin with the runtime you intend to use.

1 / 4
Optional practitioner sectionApply the optimizations from the previous lessonsShow code

Standard use case

Apply the optimizations from the previous lessons

2 recipes · compare the trade-offs

LLM Compressor uses the same oneshot entry point for different model changes. Compare a deployment-friendly FP8 quantization recipe with a SparseGPT recipe that removes two of every four eligible weights.

  1. 01
    Load the checkpoint

    Start from the same Transformers model and tokenizer used for evaluation.

  2. 02
    Choose one modifier

    Quantization changes number formats; sparsification changes which weight connections remain.

  3. 03
    Validate the real target

    Measure quality and confirm that the serving runtime and GPU have kernels for the exported format.

Quantization · FP8 weights + activationsquantize_fp8.py
Official reference
from transformers import AutoModelForCausalLM, AutoTokenizerfrom llmcompressor import oneshotfrom llmcompressor.modifiers.quantization import QuantizationModifiermodel_id = "meta-llama/Meta-Llama-3-8B-Instruct"model = AutoModelForCausalLM.from_pretrained(model_id)tokenizer = AutoTokenizer.from_pretrained(model_id)recipe = QuantizationModifier(    targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])oneshot(model=model, recipe=recipe)model.save_pretrained("./llama-3-8b-fp8", save_compressed=True)tokenizer.save_pretrained("./llama-3-8b-fp8")

This maps to the Quantization lesson: Linear weights become FP8 and input activations are dynamically scaled per token. No calibration dataset is needed for this dynamic recipe.

Sparsification · SparseGPT 2:4 weightssparsify_2of4.py
Official reference
from llmcompressor import oneshotfrom llmcompressor.modifiers.pruning import SparseGPTModifier# model + tokenizer are loaded as in the FP8 example# calibration_data contains representative, tokenized promptsrecipe = SparseGPTModifier(    sparsity=0.5,    mask_structure="2:4",    targets="Linear",    ignore=["re:.*lm_head"],)oneshot(    model=model, dataset=calibration_data, recipe=recipe,    num_calibration_samples=512, max_seq_length=2048,)model.save_pretrained("./llama-3-8b-2of4", save_compressed=True)

This maps to the Sparsification lesson: SparseGPT uses calibration activations to keep two weights out of each group of four. Current LLM Compressor docs warn that sparse 2:4 export is no longer supported for vLLM serving, so treat this as an offline recipe unless your target runtime provides a compatible sparse kernel.

03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
InputFP checkpointTransformers model
RecipeModifiersCompression plan
OutputDeployableSaved artifact

Actual benchmark

Distributed AWQ shortened a real Llama compression run

Named model · measured workload
AWQ compression time · lower is betterminutes
Distributed AWQ shortened a real Llama compression runSingle GPU: 7.02 min; 4-GPU DDP: 2.40 minSingle GPU7.02 min4-GPU DDP2.40 min
Model
Llama-3-8B
Hardware
1 GPU vs 4 GPUs; accelerator type not reported
Benchmark
LLM Compressor AWQ calibration + export run

What this result shows · The project’s release benchmark reports a 2.9× compression-time speedup and peak memory falling from 10.20 GB to 4.99 GB per process with four-way DDP.

What it does not prove · This measures the offline artifact-building pipeline, not serving latency; the release does not name the calibration corpus or GPU model.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · LLM Compressor release benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why should a compression recipe begin with the deployment target?
02What is calibration data used to observe?
03A compressed artifact is 60% smaller and passes offline perplexity checks, but the serving runtime lacks its kernel. Is it deployment-ready?
Answer all three, then check your reasoning.

What we can now explain

LLM Compressor, in one sentence

The exported artifact is smaller and ready for supported serving kernels.

Next lesson

A new bottleneck now comes into view

How does the artifact become a live server?Continue to vLLM & SGLang
Back to Speculative Decoding