Skip to llm-d lesson
LLM Inference Visualizer

Implementation Details · Lesson 12

llm-d

How do model servers become one production cluster?

From the previous lesson

Carry forward: GuideLLM

Generate repeatable production-shaped traffic, collect latency distributions and throughput, locate saturation, and compare configurations against an SLO.

Why this lesson comes next

How do model servers become one production cluster?

Add load- and prefix-aware routing, prefill/decode disaggregation, autoscaling, workload policy, and failure handling above model-server replicas.
Terms, if you need themDisaggregation · KV transfer · RDMA
Disaggregation
Running prefill and decode on separately managed worker pools.
KV transfer
Moving the prompt's per-layer Key and Value blocks to the selected decode worker.
RDMA
A low-overhead network path that can move data directly between device or host memory regions.

Keep in mind: Disaggregation helps only when phase isolation outweighs KV-transfer latency and when both pools remain balanced.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementThe router chooses two workers for one request

The endpoint picker selects a prefill worker and a decode worker using queue state, policy, and cache locality.

BottleneckCluster coordination

Route one request through compute-heavy prefill, transfer its KV state, then continue bandwidth-heavy decode elsewhere.

Metric to watchTail latency + scale
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Pair

The router chooses two workers for one request

The endpoint picker selects a prefill worker and a decode worker using queue state, policy, and cache locality.

Live diagram

One request, two specialized pools, one KV handoff

Teaching diagram · illustrative values unless marked as measured
Lesson 12 · 14 total
01 · Incoming request“A quick brown fox”
→
02 · llm-d RouterSelect one prefill + decode pairqueue · load · policy · KV affinity
→
03 · Stream from decodejumps → over → the → lazy → dog
Prefill InferencePoolProcess the whole prompt together
FLOPs-bound
Aquickbrownfox
Input
All prompt tokens
Dominant work
Large QKV + MLP matrix multiplies
SLO pressure
Time to first token
vLLMvLLMvLLM
The handoffKV cache for the prompt
K blocks
layers 1…L
V blocks
layers 1…L
Bulk data plane
K/V tensors for every prompt token
Small control plane
host · port · request ID · block handles
Typical path
Decoder pulls over RDMA with NIXL
Decode InferencePoolGenerate one token per step
Bandwidth-bound
Read all saved K/V→append “jumps”↻
Input
Transferred KV + current token
Dominant work
Repeated HBM reads of model + KV
SLO pressure
Inter-token latency
vLLMvLLMvLLM
Why separate the pools?Different bottlenecks, different scaling

Long compute bursts no longer pause latency-sensitive decode loops. Prefill and decode replicas can scale independently and use different tensor-parallel layouts.

What does not make the bulk trip?Model weights stay resident in both pools

Intermediate activations are discarded. The request body is forwarded as control input, but the expensive payload is the per-layer KV cache.

When does it help?Only when isolation wins more than transfer costs

Fast RDMA-class networking matters. Slow KV movement, poor pool balance, or short prompts can make a combined prefill/decode worker faster.

The same client request now has a compute phase and a token-streaming phase.

1 / 4
Optional practitioner sectionDeploy the routed vLLM baseline that P/D extendsShow code

Standard use case

Deploy the routed vLLM baseline that P/D extends

Shell · practical starting point

This official quickstart creates the routing and worker foundation. It does not by itself reproduce the separate prefill/decode benchmark shown above; P/D adds specialized pools and a KV-transfer data plane on top.

  1. 01
    Load the recipe environment

    The repository helper supplies the chart locations and version used by the matching quickstart.

  2. 02
    Install the router

    The optimized baseline enables an inference-aware route in front of the model servers.

  3. 03
    Create the baseline workers

    The Kubernetes recipe starts combined GPU-backed vLLM replicas. Treat this as the prerequisite baseline before configuring separate prefill and decode pools.

Deploy the routed vLLM baseline that P/D extendsdeploy-llm-d.sh
Official reference
export GUIDE_NAME=quickstartexport NAMESPACE=llm-d-quickstartsource guides/env.shhelm install "${GUIDE_NAME}" "${ROUTER_STANDALONE_CHART}" \  -f guides/recipes/router/base.values.yaml \  -f guides/optimized-baseline/router/optimized-baseline.values.yaml \  -n "${NAMESPACE}" --version "${ROUTER_CHART_VERSION}"kubectl apply -n "${NAMESPACE}" \  -k guides/optimized-baseline/modelserver/gpu/vllm/base/

This optional code is intentionally the official routed baseline, not a claim that P/D is enabled. Follow the llm-d disaggregation guide for the pool and NIXL configuration used by your cluster.

02B

Router deep dive

How does the router choose a worker?

A routing profile composes filters, scorers, and a picker; it is not limited to one fixed load-balancing rule.

The proxy receives the request, while the Endpoint Picker uses request metadata and live serving state to choose an endpoint. In disaggregated serving, it can run separate profiles for the prefill and decode pools.

Routing choicePrefer this whenWhat to watch
Queue + load awareRequests vary in length and workers accumulate uneven queues or token load.A short queue can still hide one very large request; combine more than one load signal.
Prefix-cache awareChats, agents, or shared system prompts repeatedly reuse long prefixes.Cache stickiness can create a hotspot, so retain a load gate or load score.
Session affinityFollow-up turns should return to the worker that handled the same session.Simple stickiness knows less than precise KV events and may concentrate active users.
Latency / SLO awareYou have an explicit TTFT or ITL objective and a usable latency predictor.Predictions and SLO metadata must be trustworthy; stale estimates make bad choices.
LoRA affinityDifferent requests need adapters and some workers already hold the requested adapter.Adapter locality must still be balanced against queue and memory pressure.
Separate P/D profilesPrefill and decode run in specialized pools with different bottlenecks.The router must select both phases, and the KV transfer must justify the split.
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
Prefill targetTTFTCompute-bound burst
Decode targetITL / TPOTBandwidth-bound loop
HandoffKV blocksTransfer must stay fast

Actual benchmark

llm-d P/D disaggregation raised GPT-OSS throughput

Named model · measured workload
Relative tokens per second× standard vLLM
llm-d P/D disaggregation raised GPT-OSS throughputStandard vLLM: 1×; llm-d P/D pools: up to 1.7×Standard vLLM1×llm-d P/D poolsup to 1.7×
Model
OpenAI GPT-OSS on vLLM
Hardware
AWS ml.p6-b200.48xlarge · NVIDIA B200
Benchmark
1,024 input + 1,024 output tokens · concurrency up to 128

What this result shows · AWS measured up to 70% more tokens per second when llm-d separated prefill and decode pools and transferred KV over EFA with NIXL.

What it does not prove · The gain depends on concurrency, pool ratio, model shape, and network speed; short prompts can be faster on a combined worker.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · AWS llm-d disaggregated inference benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why can prefill and decode benefit from separate worker pools?
02What remains resident in both prefill and decode pools?
03Short prompts run on a slow network and the decode pool often waits for KV transfers. What should the operator test next?
Answer all three, then check your reasoning.

What we can now explain

llm-d, in one sentence

Each phase can scale and specialize independently when the latency saved exceeds the KV transfer cost.

Journey complete

Connect the full serving system

Return to the complete inference journey.Continue to Course complete
Back to GuideLLM