From the previous lesson
Implementation Details · Lesson 10
vLLM & SGLang
How do two modern serving engines assemble the same inference techniques?
Why this lesson comes next
How do two modern runtimes build a model server?
Trace one request through both engines, compare paged and radix-organized KV reuse, and learn how to choose through a controlled workload benchmark.Terms, if you need themToken budget · RadixAttention · Goodput+
- Token budget
- The maximum prompt and decode token work admitted in one scheduler iteration.
- RadixAttention
- SGLang's radix-tree organization for retaining and reusing KV prefixes shared across requests.
- Goodput
- Throughput from requests that actually satisfy the chosen latency objective.
Keep in mind: Both optimize model-server execution, but cluster-wide routing, autoscaling, and cross-replica policy still require a higher orchestration layer.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.vLLM and SGLang expose OpenAI-compatible serving, tokenize the prompt, and create runtime request state.
vLLM and SGLang run the same outer loop: admit work, allocate KV, execute tokens, stream results, and reclaim memory. They organize some internal state differently.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Enter
Both engines turn an API request into token work
vLLM and SGLang expose OpenAI-compatible serving, tokenize the prompt, and create runtime request state.Live diagram
Trace the same request through vLLM and SGLang
Teaching diagram · illustrative values unless marked as measuredSignature idea: PagedAttention manages KV in blocks; automatic prefix caching can reuse matching hashed blocks.
Signature idea: RadixAttention stores reusable KV prefixes in a radix tree, which is useful when requests repeatedly share prefixes.
No universal winner: compare TTFT, ITL, throughput, memory use, and operational fit using the same model, precision, prompts, concurrency, and hardware.
The client contract can look nearly identical even when the internal engine differs.
Optional practitioner sectionLaunch the same OpenAI-compatible endpoint with either engineShow code
Standard use case
Launch the same OpenAI-compatible endpoint with either engine
Both commands load one model server and accept the same style of client request. Keep model, precision, context length, and hardware fixed when benchmarking them.
- 01Hold the comparison constant
Use the same checkpoint, quantization, maximum context, GPU allocation, prompt distribution, and concurrency.
- 02Exercise the same API
Both servers can sit behind an OpenAI-compatible client, so client behavior need not decide the engine.
- 03Compare SLO goodput
Measure TTFT, ITL, throughput, memory use, failures, and the number of requests meeting the SLO.
vllm serve ./llama-3-8b-fp8 \ --served-model-name fox-model \ --host 0.0.0.0 --port 8000 \ --max-model-len 4096vLLM loads the checkpoint, profiles available memory, and runs its scheduler, KV manager, workers, and sampler behind the API.
python -m sglang.launch_server \ --model-path ./llama-3-8b-fp8 \ --served-model-name fox-model \ --host 0.0.0.0 --port 8000 \ --context-length 4096SGLang exposes the same client boundary while its scheduler and RadixAttention runtime manage execution and reusable prefixes.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
vLLM served real chat-length traces faster than TGI
- Model
- LLaMA-13B
- Hardware
- 1× NVIDIA A100 40GB
- Benchmark
- Input/output lengths sampled from ShareGPT
What this result shows · The vLLM project measured 2.2–2.5× TGI throughput on LLaMA-13B while serving ShareGPT-shaped requests.
What it does not prove · This is an end-to-end engine result combining scheduling, kernels, and PagedAttention; it does not isolate a single subsystem.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM launch benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
vLLM & SGLang, in one sentence
Choose the engine that produces higher SLO-compliant goodput for the real workload.Next lesson