Skip to Speculative Decoding lesson
LLM Inference Visualizer

Inference Optimization · Lesson 08

Speculative Decoding

Can one expensive target pass accept several tokens?

From the previous lesson

Carry forward: Prefix Caching

Reuse content-addressed KV blocks for exact shared prefixes, then examine cache hits, eviction, tenant boundaries, and cache-aware routing.

Why this lesson comes next

Can one expensive target pass accept several output tokens?

Let a cheaper draft mechanism propose a token run, verify the run in parallel, and see how acceptance rate determines whether speculation helps.
Terms, if you need themDraft model · Target model · Acceptance rate
Draft model
A cheaper model that proposes several likely next tokens.
Target model
The authoritative model that verifies every committed output token.
Acceptance rate
The fraction of draft proposals that match what the target model would accept.

Keep in mind: Speculation can lose when draft cost is high or acceptance is low, and its benefit often shrinks at large serving batch sizes.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementThe small model predicts three display tokens ahead

For “A quick brown fox,” the draft proposes “jumps over a” before the target runs.

BottleneckSerial target passes

Let a cheap model propose three words, then use one target-model pass to verify the whole run.

Metric to watchITL
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Draft 3

The small model predicts three display tokens ahead

For “A quick brown fox,” the draft proposes “jumps over a” before the target runs.

Live diagram

Draft several tokens, verify them together

Teaching diagram · illustrative values unless marked as measured
Lesson 08 · 14 total
Teaching simulation · illustrative timingDraft exactly 3 display tokens aheadFor readability, one displayed word stands in for one tokenizer token. The measured benchmark below uses real tokenization.
Committed context
Aquickbrownfoxwaiting for generation
Small draft modelPredict 3 words ahead3 ms per draft word

Fast guesses; never committed without verification.

Large target modelVerify all 3 positions together45 ms per target pass
One causal input sent to the target

The prefix and all three draft words enter one forward pass. Each position can only see tokens to its left.

Three shifted next-token checks computed in parallel

Run the target to accept or correct the draft.

Ready: the draft proposes “jumps over a”.
Acceptance → illustrative serving performanceHigher acceptance amortizes each target pass over more output tokens
Teaching assumptions: 45 ms target + 3 × 3 ms draft = 54 ms per round
Expected words / target pass2.311 + h + h²
Effective ITL23.4 msBaseline 45 ms
Throughput42.8 tok/sBaseline 22.2 tok/s
Speedup1.93×Net win
Example TTFT SLO ≤ 250 ms180 ms · pass
Example ITL SLO ≤ 25 ms23.4 ms · pass
Example throughput SLO ≥ 40 tok/s42.8 tok/s · pass

Expected committed tokens per target pass: E = 1 + h + h². This teaching estimate assumes the same independent acceptance probability h at every position; a real draft model has position- and prompt-dependent acceptance. Verification stops at the first miss.

The displayed words are cheap proposals, not committed output; real tokenizers may split a word into multiple tokens.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
Draft costMust be lowCheap proposals
AcceptanceMust be highUseful draft run
QualityPreservedTarget verifies

Actual benchmark

A real draft model accelerated Llama 3 on ShareGPT

Named model · measured workload
Relative decoding speed× normal vLLM decoding
A real draft model accelerated Llama 3 on ShareGPTTarget only: 1×; Draft + verify: up to 1.5×Target only1×Draft + verifyup to 1.5×
Model
Llama-3-70B + Qwama-0.5B-Instruct draft
Hardware
4× NVIDIA H100
Benchmark
ShareGPT · QPS 1

What this result shows · vLLM measured up to 1.5× faster token generation when the small Qwama draft model proposed tokens for Llama-3-70B.

What it does not prove · The same report reached 2.8× with n-gram speculation on CNN/DailyMail; acceptance rate and draft cost determine the gain.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM speculative decoding benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01Who decides which drafted tokens become final output?
02Why are draft tokens after the first rejection discarded?
03Draft A costs 3 ms/token with 80% acceptance; Draft B costs 20 ms/token with 85% acceptance. What must be compared before choosing B?
Answer all three, then check your reasoning.

What we can now explain

Speculative Decoding, in one sentence

High acceptance produces more output tokens per target-model pass.

Next chapter

A new bottleneck now comes into view

How is a compressed artifact prepared?Continue to LLM Compressor
Back to Prefix Caching