From the previous lesson
Implementation Details · Lesson 12
llm-d
How do model servers become one production cluster?
Why this lesson comes next
How do model servers become one production cluster?
Add load- and prefix-aware routing, prefill/decode disaggregation, autoscaling, workload policy, and failure handling above model-server replicas.Terms, if you need themDisaggregation · KV transfer · RDMA+
- Disaggregation
- Running prefill and decode on separately managed worker pools.
- KV transfer
- Moving the prompt's per-layer Key and Value blocks to the selected decode worker.
- RDMA
- A low-overhead network path that can move data directly between device or host memory regions.
Keep in mind: Disaggregation helps only when phase isolation outweighs KV-transfer latency and when both pools remain balanced.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.The endpoint picker selects a prefill worker and a decode worker using queue state, policy, and cache locality.
Route one request through compute-heavy prefill, transfer its KV state, then continue bandwidth-heavy decode elsewhere.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Pair
The router chooses two workers for one request
The endpoint picker selects a prefill worker and a decode worker using queue state, policy, and cache locality.Live diagram
One request, two specialized pools, one KV handoff
Teaching diagram · illustrative values unless marked as measured- Input
- All prompt tokens
- Dominant work
- Large QKV + MLP matrix multiplies
- SLO pressure
- Time to first token
layers 1…LV blocks
layers 1…L
- Bulk data plane
- K/V tensors for every prompt token
- Small control plane
- host · port · request ID · block handles
- Typical path
- Decoder pulls over RDMA with NIXL
- Input
- Transferred KV + current token
- Dominant work
- Repeated HBM reads of model + KV
- SLO pressure
- Inter-token latency
Long compute bursts no longer pause latency-sensitive decode loops. Prefill and decode replicas can scale independently and use different tensor-parallel layouts.
Intermediate activations are discarded. The request body is forwarded as control input, but the expensive payload is the per-layer KV cache.
Fast RDMA-class networking matters. Slow KV movement, poor pool balance, or short prompts can make a combined prefill/decode worker faster.
The same client request now has a compute phase and a token-streaming phase.
Optional practitioner sectionDeploy the routed vLLM baseline that P/D extendsShow code
Standard use case
Deploy the routed vLLM baseline that P/D extends
This official quickstart creates the routing and worker foundation. It does not by itself reproduce the separate prefill/decode benchmark shown above; P/D adds specialized pools and a KV-transfer data plane on top.
- 01Load the recipe environment
The repository helper supplies the chart locations and version used by the matching quickstart.
- 02Install the router
The optimized baseline enables an inference-aware route in front of the model servers.
- 03Create the baseline workers
The Kubernetes recipe starts combined GPU-backed vLLM replicas. Treat this as the prerequisite baseline before configuring separate prefill and decode pools.
export GUIDE_NAME=quickstartexport NAMESPACE=llm-d-quickstartsource guides/env.shhelm install "${GUIDE_NAME}" "${ROUTER_STANDALONE_CHART}" \ -f guides/recipes/router/base.values.yaml \ -f guides/optimized-baseline/router/optimized-baseline.values.yaml \ -n "${NAMESPACE}" --version "${ROUTER_CHART_VERSION}"kubectl apply -n "${NAMESPACE}" \ -k guides/optimized-baseline/modelserver/gpu/vllm/base/This optional code is intentionally the official routed baseline, not a claim that P/D is enabled. Follow the llm-d disaggregation guide for the pool and NIXL configuration used by your cluster.
Router deep dive
How does the router choose a worker?
A routing profile composes filters, scorers, and a picker; it is not limited to one fixed load-balancing rule.The proxy receives the request, while the Endpoint Picker uses request metadata and live serving state to choose an endpoint. In disaggregated serving, it can run separate profiles for the prefill and decode pools.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
llm-d P/D disaggregation raised GPT-OSS throughput
- Model
- OpenAI GPT-OSS on vLLM
- Hardware
- AWS ml.p6-b200.48xlarge · NVIDIA B200
- Benchmark
- 1,024 input + 1,024 output tokens · concurrency up to 128
What this result shows · AWS measured up to 70% more tokens per second when llm-d separated prefill and decode pools and transferred KV over EFA with NIXL.
What it does not prove · The gain depends on concurrency, pool ratio, model shape, and network speed; short prompts can be faster on a combined worker.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · AWS llm-d disaggregated inference benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
llm-d, in one sentence
Each phase can scale and specialize independently when the latency saved exceeds the KV transfer cost.Journey complete