Skip to vLLM & SGLang lesson
LLM Inference Visualizer

Implementation Details · Lesson 10

vLLM & SGLang

How do two modern serving engines assemble the same inference techniques?

From the previous lesson

Carry forward: LLM Compressor

Follow calibration, compression recipes, modifier application, format export, and evaluation as an offline pipeline feeding the serving engine.

Why this lesson comes next

How do two modern runtimes build a model server?

Trace one request through both engines, compare paged and radix-organized KV reuse, and learn how to choose through a controlled workload benchmark.
Terms, if you need themToken budget · RadixAttention · Goodput
Token budget
The maximum prompt and decode token work admitted in one scheduler iteration.
RadixAttention
SGLang's radix-tree organization for retaining and reusing KV prefixes shared across requests.
Goodput
Throughput from requests that actually satisfy the chosen latency objective.

Keep in mind: Both optimize model-server execution, but cluster-wide routing, autoscaling, and cross-replica policy still require a higher orchestration layer.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementBoth engines turn an API request into token work

vLLM and SGLang expose OpenAI-compatible serving, tokenize the prompt, and create runtime request state.

BottleneckEnd-to-end execution

vLLM and SGLang run the same outer loop: admit work, allocate KV, execute tokens, stream results, and reclaim memory. They organize some internal state differently.

Metric to watchGoodput
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Enter

Both engines turn an API request into token work

vLLM and SGLang expose OpenAI-compatible serving, tokenize the prompt, and create runtime request state.

Live diagram

Trace the same request through vLLM and SGLang

Teaching diagram · illustrative values unless marked as measured
Lesson 10 · 14 total
Same client request“A quick brown fox”prompt + sampling parameters
Choose an enginevLLM or SGLangBenchmark the exact model, hardware, and traffic shape
Same streamed resultjumps → over → the → lazy → dogfinish → release KV state
Serving enginevLLM
General-purpose runtime and ecosystem
01 · text → token IDsInput processor
02 · continuous batchingScheduler
03 · paged blocks + APCKV manager
04 · GPU forward passModel executor
05 · choose next tokenSampler
schedule → execute → sample → repeat ↻

Signature idea: PagedAttention manages KV in blocks; automatic prefix caching can reuse matching hashed blocks.

Serving engine + programming layerSGLang
Runtime and language co-designed for LLM programs
01 · request → token IDsTokenizer manager
02 · continuous batchingScheduler
03 · radix prefix treeRadixAttention
04 · GPU forward passModel runner
05 · sample / constrainSampler
schedule → execute → sample → repeat ↻

Signature idea: RadixAttention stores reusable KV prefixes in a radix tree, which is useful when requests repeatedly share prefixes.

Large overlapBoth support continuous batching, paged KV memory, prefix caching, chunked prefill, speculative decoding, structured outputs, quantization, parallel execution, and OpenAI-compatible APIs.
DecisionvLLM emphasisSGLang emphasis
Core abstractionFlexible model-serving engine with broad integrations and execution backends.Serving runtime co-designed with a frontend for multi-call, structured LLM programs.
Prefix reusePaged KV blocks with hash-based automatic prefix caching.RadixAttention organizes reusable prefixes in a radix tree.
Structured generationStructured-output backends such as XGrammar or Guidance.Constrained decoding integrated with its program/runtime execution model.
How to choosePrefer when its model, hardware, integrations, and operational tooling fit your stack best.Prefer when its prefix-heavy or structured-program path benchmarks better for your workload.

No universal winner: compare TTFT, ITL, throughput, memory use, and operational fit using the same model, precision, prompts, concurrency, and hardware.

The client contract can look nearly identical even when the internal engine differs.

1 / 4
Optional practitioner sectionLaunch the same OpenAI-compatible endpoint with either engineShow code

Standard use case

Launch the same OpenAI-compatible endpoint with either engine

2 recipes · compare the trade-offs

Both commands load one model server and accept the same style of client request. Keep model, precision, context length, and hardware fixed when benchmarking them.

  1. 01
    Hold the comparison constant

    Use the same checkpoint, quantization, maximum context, GPU allocation, prompt distribution, and concurrency.

  2. 02
    Exercise the same API

    Both servers can sit behind an OpenAI-compatible client, so client behavior need not decide the engine.

  3. 03
    Compare SLO goodput

    Measure TTFT, ITL, throughput, memory use, failures, and the number of requests meeting the SLO.

vLLM serverserve-vllm.sh
Official reference
vllm serve ./llama-3-8b-fp8 \  --served-model-name fox-model \  --host 0.0.0.0 --port 8000 \  --max-model-len 4096

vLLM loads the checkpoint, profiles available memory, and runs its scheduler, KV manager, workers, and sampler behind the API.

SGLang serverserve-sglang.sh
Official reference
python -m sglang.launch_server \  --model-path ./llama-3-8b-fp8 \  --served-model-name fox-model \  --host 0.0.0.0 --port 8000 \  --context-length 4096

SGLang exposes the same client boundary while its scheduler and RadixAttention runtime manage execution and reusable prefixes.

03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
SharedServing loopBatch + execute
vLLMPaged blocksBroad engine ecosystem
SGLangRadix treeProgram/runtime co-design

Actual benchmark

vLLM served real chat-length traces faster than TGI

Named model · measured workload
Serving throughput× Hugging Face TGI
vLLM served real chat-length traces faster than TGITGI: 1×; vLLM: 2.2–2.5×TGI1×vLLM2.2–2.5×
Model
LLaMA-13B
Hardware
1× NVIDIA A100 40GB
Benchmark
Input/output lengths sampled from ShareGPT

What this result shows · The vLLM project measured 2.2–2.5× TGI throughput on LLaMA-13B while serving ShareGPT-shaped requests.

What it does not prove · This is an end-to-end engine result combining scheduling, kernels, and PagedAttention; it does not isolate a single subsystem.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM launch benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01In both vLLM and SGLang, which component decides what token work runs next?
02Why is a request in either engine not just one model forward call?
03What is the clearest cache-organization difference highlighted in this lesson?
Answer all three, then check your reasoning.

What we can now explain

vLLM & SGLang, in one sentence

Choose the engine that produces higher SLO-compliant goodput for the real workload.

Next lesson

A new bottleneck now comes into view

How do we measure the live server?Continue to GuideLLM
Back to LLM Compressor