Skip to lesson catalog

A systems journey · one prompt to a production cluster

Understand what makes
LLM inference fast

Follow one request through memory, model compression, scheduling, benchmarking, and distributed serving. Every lesson starts with a bottleneck and ends with the metric it changes.

5 chapters14 lessons≈ 101 minutes0 / 14 completed

New to inference systems?

Start with one request, then widen the system

Each lesson follows the same prompt and answers three questions: what hurts, what changes, and which metric moves.
  1. 00FoundationHow does generation work?
  2. 01MemoryWhat occupies the GPU?
  3. 02ModelCan the artifact shrink?
  4. 03RuntimeCan requests share better?
  5. 04ProductionCan the service scale safely?
Begin with Autoregressive Generation
CH. 00

What happens before we optimize anything?

Inference Foundations

Build the serving mental model first: watch a causal transformer generate one token at a time, decide when running locally is useful, and learn the latency, throughput, and quality objectives used to judge the system.

CH. 01

What is consuming the machine?

Memory Model

Begin with one request. See why saved attention work makes decoding faster, why the cache grows with every token, and why fitting a model is not the same as serving it efficiently.

CH. 02

Can the model itself become cheaper?

Model Optimization

Change the artifact before serving begins. Store fewer bits, remove selected weights, and keep quality, kernel support, and real wall-clock performance visible as separate concerns.

CH. 03

Can requests share work and hardware better?

Inference Optimization

Keep the accelerator useful under uneven traffic. Schedule at token boundaries, allocate cache in pages, reuse exact prefixes, and verify multiple draft tokens per target-model pass.

CH. 04

Where do these mechanisms live in real software?

Implementation Details

Finish the journey at the software boundary. Prepare a compressed artifact, serve it efficiently, measure it under realistic load, and coordinate replicas across a cluster.

Course synthesis

Different techniques solve different bottlenecks

They compose; they are not interchangeable.
Repeated token workKV Cache · Speculative Decoding

Reuse past state or commit more useful output per expensive target pass.

Model bytesQuantization · Sparsification

Change the artifact before serving; kernel support decides whether smaller also means faster.

Shared GPU capacityContinuous Batching · PagedAttention

Keep execution slots useful and allocate growing KV state without large reservations.

Repeated promptsPrefix Caching

Skip prefill only when the exact token prefix and model state are compatible.

Production operationvLLM · GuideLLM · llm-d

Serve one model, measure its SLO-safe capacity, then coordinate multiple replicas.