From the previous lesson
Implementation Details · Lesson 11
GuideLLM
How do we know whether a deployment is fast enough?
Why this lesson comes next
How do we know whether a deployment is fast enough?
Generate repeatable production-shaped traffic, collect latency distributions and throughput, locate saturation, and compare configurations against an SLO.Terms, if you need themTTFT · SLO · Goodput+
- TTFT
- Time to first token: queueing and prefill delay before streaming begins.
- SLO
- A measurable service objective, such as p95 TTFT below a chosen threshold.
- Goodput
- The rate of requests or tokens completed while meeting the SLO.
Keep in mind: A benchmark predicts production only when its model, prompt lengths, output lengths, arrival pattern, and hardware resemble real traffic.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.Choose prompt lengths, output lengths, data, and multi-turn behavior that resemble real traffic.
Increase realistic load until latency crosses the SLO; the last passing point is safe goodput.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Shape
Describe production-like requests
Choose prompt lengths, output lengths, data, and multi-turn behavior that resemble real traffic.Live diagram
Read an illustrative sweep, then inspect a measured run
Teaching diagram · illustrative values unless marked as measuredA benchmark is meaningful only when its request shape is realistic.
Optional practitioner sectionSweep concurrency against the live vLLM endpointShow code
Standard use case
Sweep concurrency against the live vLLM endpoint
Point GuideLLM at the server from the previous lesson and sweep through increasing load using an actual prompt dataset. Keep the same dataset, seed, and limits when comparing configurations.
- 01Choose the target
The backend points at any compatible inference endpoint, including the local vLLM server.
- 02Sweep the load
The sweep profile tests multiple concurrency points instead of reporting one flattering number.
- 03Compare with the SLO
Use the generated JSON and CSV results to compare TTFT, token latency, throughput, and goodput.
guidellm run \ --backend kind=openai_http,target=http://localhost:8000 \ --profile kind=sweep \ --constraint kind=max_requests,count=1000 \ --data '{"kind":"huggingface","source":"anon8231489123/ShareGPT_Vicuna_unfiltered","split":"train"}' \ --seed kind=static,value=42GuideLLM writes benchmarks.json and benchmarks.csv by default, so the same run can feed a report or regression check.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
A real benchmark exposed the tail, not just the average
- Model
- Qwen2.5-72B-AWQ
- Hardware
- Remote OpenAI-compatible endpoint; GPU not disclosed
- Benchmark
- 10 ShareGPT requests; throughput profile
What this result shows · GuideLLM reported 68.3 output tok/s, 259.9 ms TTFT, and a 9.53 s p99 request latency; only 3 of 10 requests completed inside the short run.
What it does not prove · This small public reproduction is evidence of what the tool records, not a capacity recommendation; use longer runs and your own SLO before deployment.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · GuideLLM public benchmark output04 · Quiz
Can you reason through the concept?
What we can now explain
GuideLLM, in one sentence
Load sweeps reveal saturation and the SLO-safe operating point.Next lesson