Reuse past state or commit more useful output per expensive target pass.
A systems journey · one prompt to a production cluster
Understand what makes
LLM inference fast
Follow one request through memory, model compression, scheduling, benchmarking, and distributed serving. Every lesson starts with a bottleneck and ends with the metric it changes.
New to inference systems?
Start with one request, then widen the system
Each lesson follows the same prompt and answers three questions: what hurts, what changes, and which metric moves.- 00FoundationHow does generation work?
- 01MemoryWhat occupies the GPU?
- 02ModelCan the artifact shrink?
- 03RuntimeCan requests share better?
- 04ProductionCan the service scale safely?
What happens before we optimize anything?
Inference Foundations
Build the serving mental model first: watch a causal transformer generate one token at a time, decide when running locally is useful, and learn the latency, throughput, and quality objectives used to judge the system.
What is consuming the machine?
Memory Model
Begin with one request. See why saved attention work makes decoding faster, why the cache grows with every token, and why fitting a model is not the same as serving it efficiently.
Can the model itself become cheaper?
Model Optimization
Change the artifact before serving begins. Store fewer bits, remove selected weights, and keep quality, kernel support, and real wall-clock performance visible as separate concerns.
Can requests share work and hardware better?
Inference Optimization
Keep the accelerator useful under uneven traffic. Schedule at token boundaries, allocate cache in pages, reuse exact prefixes, and verify multiple draft tokens per target-model pass.
Continuous Batching
Why wait for the slowest request in a fixed batch?
PagedAttention
How can variable KV caches grow without large contiguous regions?
Prefix Caching
Why recompute the same system prompt for every user?
Speculative Decoding
Can one expensive target pass accept several output tokens?
Where do these mechanisms live in real software?
Implementation Details
Finish the journey at the software boundary. Prepare a compressed artifact, serve it efficiently, measure it under realistic load, and coordinate replicas across a cluster.
LLM Compressor
How does a full-precision checkpoint become deployment-ready?
vLLM & SGLang
How do two modern runtimes build a model server?
GuideLLM
How do we know whether a deployment is fast enough?
llm-d
How do model servers become one production cluster?
Course synthesis
Different techniques solve different bottlenecks
They compose; they are not interchangeable.Change the artifact before serving; kernel support decides whether smaller also means faster.
Keep execution slots useful and allocate growing KV state without large reservations.
Skip prefill only when the exact token prefix and model state are compatible.
Serve one model, measure its SLO-safe capacity, then coordinate multiple replicas.