Skip to GuideLLM lesson
LLM Inference Visualizer

Implementation Details · Lesson 11

GuideLLM

How do we know whether a deployment is fast enough?

From the previous lesson

Carry forward: vLLM & SGLang

Trace one request through both engines, compare paged and radix-organized KV reuse, and learn how to choose through a controlled workload benchmark.

Why this lesson comes next

How do we know whether a deployment is fast enough?

Generate repeatable production-shaped traffic, collect latency distributions and throughput, locate saturation, and compare configurations against an SLO.
Terms, if you need themTTFT · SLO · Goodput
TTFT
Time to first token: queueing and prefill delay before streaming begins.
SLO
A measurable service objective, such as p95 TTFT below a chosen threshold.
Goodput
The rate of requests or tokens completed while meeting the SLO.

Keep in mind: A benchmark predicts production only when its model, prompt lengths, output lengths, arrival pattern, and hardware resemble real traffic.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementDescribe production-like requests

Choose prompt lengths, output lengths, data, and multi-turn behavior that resemble real traffic.

BottleneckUnreliable measurement

Increase realistic load until latency crosses the SLO; the last passing point is safe goodput.

Metric to watchSLO goodput
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Shape

Describe production-like requests

Choose prompt lengths, output lengths, data, and multi-turn behavior that resemble real traffic.

Live diagram

Read an illustrative sweep, then inspect a measured run

Teaching diagram · illustrative values unless marked as measured
Lesson 11 · 14 total

A benchmark is meaningful only when its request shape is realistic.

1 / 4
Optional practitioner sectionSweep concurrency against the live vLLM endpointShow code

Standard use case

Sweep concurrency against the live vLLM endpoint

Shell · practical starting point

Point GuideLLM at the server from the previous lesson and sweep through increasing load using an actual prompt dataset. Keep the same dataset, seed, and limits when comparing configurations.

  1. 01
    Choose the target

    The backend points at any compatible inference endpoint, including the local vLLM server.

  2. 02
    Sweep the load

    The sweep profile tests multiple concurrency points instead of reporting one flattering number.

  3. 03
    Compare with the SLO

    Use the generated JSON and CSV results to compare TTFT, token latency, throughput, and goodput.

Sweep concurrency against the live vLLM endpointbenchmark.sh
Official reference
guidellm run \  --backend kind=openai_http,target=http://localhost:8000 \  --profile kind=sweep \  --constraint kind=max_requests,count=1000 \  --data '{"kind":"huggingface","source":"anon8231489123/ShareGPT_Vicuna_unfiltered","split":"train"}' \  --seed kind=static,value=42

GuideLLM writes benchmarks.json and benchmarks.csv by default, so the same run can feed a report or regression check.

03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
TTFTFirst tokenPrefill + queue
ITLToken gapDecode pace
GoodputWithin SLOUseful throughput

Actual benchmark

A real benchmark exposed the tail, not just the average

Named model · measured workload
End-to-end request latencyseconds · lower is better
A real benchmark exposed the tail, not just the averageMedian: 7.91 s; P99: 9.53 sMedian7.91 sP999.53 s
Model
Qwen2.5-72B-AWQ
Hardware
Remote OpenAI-compatible endpoint; GPU not disclosed
Benchmark
10 ShareGPT requests; throughput profile

What this result shows · GuideLLM reported 68.3 output tok/s, 259.9 ms TTFT, and a 9.53 s p99 request latency; only 3 of 10 requests completed inside the short run.

What it does not prove · This small public reproduction is evidence of what the tool records, not a capacity recommendation; use longer runs and your own SLO before deployment.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · GuideLLM public benchmark output

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Why should a benchmark match production prompt and output lengths?
02What does the saturation knee represent?
03Config A completes 100 req/s but only 70 meet the SLO; Config B completes 85 req/s and 82 meet it. Which has higher goodput?
Answer all three, then check your reasoning.

What we can now explain

GuideLLM, in one sentence

Load sweeps reveal saturation and the SLO-safe operating point.

Next lesson

A new bottleneck now comes into view

How do multiple servers act as one service?Continue to llm-d
Back to vLLM & SGLang