From the previous chapter
Implementation Details · Lesson 09
LLM Compressor
How does a full-precision checkpoint become deployment-ready?
Why this lesson comes next
How does a full-precision checkpoint become deployment-ready?
Follow calibration, compression recipes, modifier application, format export, and evaluation as an offline pipeline feeding the serving engine.Terms, if you need themRecipe · Calibration · Artifact+
- Recipe
- A configuration describing which compression modifiers to apply and where.
- Calibration
- Running representative samples to observe tensor ranges or importance before transformation.
- Artifact
- The saved model files, metadata, scales, and configuration delivered to a serving runtime.
Keep in mind: LLM Compressor prepares and evaluates an artifact; it does not schedule live requests or guarantee that the server will meet latency goals.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.Pick a quantization scheme supported by the serving hardware and model architecture.
Prepare the model offline: choose a recipe, calibrate if needed, transform, evaluate, and export.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Choose
Start with a deployment target
Pick a quantization scheme supported by the serving hardware and model architecture.Live diagram
An offline pipeline from checkpoint to compressed artifact
Teaching diagram · illustrative values unless marked as measuredThe recipe should begin with the runtime you intend to use.
Optional practitioner sectionApply the optimizations from the previous lessonsShow code
Standard use case
Apply the optimizations from the previous lessons
LLM Compressor uses the same oneshot entry point for different model changes. Compare a deployment-friendly FP8 quantization recipe with a SparseGPT recipe that removes two of every four eligible weights.
- 01Load the checkpoint
Start from the same Transformers model and tokenizer used for evaluation.
- 02Choose one modifier
Quantization changes number formats; sparsification changes which weight connections remain.
- 03Validate the real target
Measure quality and confirm that the serving runtime and GPU have kernels for the exported format.
from transformers import AutoModelForCausalLM, AutoTokenizerfrom llmcompressor import oneshotfrom llmcompressor.modifiers.quantization import QuantizationModifiermodel_id = "meta-llama/Meta-Llama-3-8B-Instruct"model = AutoModelForCausalLM.from_pretrained(model_id)tokenizer = AutoTokenizer.from_pretrained(model_id)recipe = QuantizationModifier( targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])oneshot(model=model, recipe=recipe)model.save_pretrained("./llama-3-8b-fp8", save_compressed=True)tokenizer.save_pretrained("./llama-3-8b-fp8")This maps to the Quantization lesson: Linear weights become FP8 and input activations are dynamically scaled per token. No calibration dataset is needed for this dynamic recipe.
from llmcompressor import oneshotfrom llmcompressor.modifiers.pruning import SparseGPTModifier# model + tokenizer are loaded as in the FP8 example# calibration_data contains representative, tokenized promptsrecipe = SparseGPTModifier( sparsity=0.5, mask_structure="2:4", targets="Linear", ignore=["re:.*lm_head"],)oneshot( model=model, dataset=calibration_data, recipe=recipe, num_calibration_samples=512, max_seq_length=2048,)model.save_pretrained("./llama-3-8b-2of4", save_compressed=True)This maps to the Sparsification lesson: SparseGPT uses calibration activations to keep two weights out of each group of four. Current LLM Compressor docs warn that sparse 2:4 export is no longer supported for vLLM serving, so treat this as an offline recipe unless your target runtime provides a compatible sparse kernel.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Distributed AWQ shortened a real Llama compression run
- Model
- Llama-3-8B
- Hardware
- 1 GPU vs 4 GPUs; accelerator type not reported
- Benchmark
- LLM Compressor AWQ calibration + export run
What this result shows · The project’s release benchmark reports a 2.9× compression-time speedup and peak memory falling from 10.20 GB to 4.99 GB per process with four-way DDP.
What it does not prove · This measures the offline artifact-building pipeline, not serving latency; the release does not name the calibration corpus or GPU model.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · LLM Compressor release benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
LLM Compressor, in one sentence
The exported artifact is smaller and ready for supported serving kernels.Next lesson