From the previous lesson
Inference Optimization · Lesson 08
Speculative Decoding
Can one expensive target pass accept several tokens?
Why this lesson comes next
Can one expensive target pass accept several output tokens?
Let a cheaper draft mechanism propose a token run, verify the run in parallel, and see how acceptance rate determines whether speculation helps.Terms, if you need themDraft model · Target model · Acceptance rate+
- Draft model
- A cheaper model that proposes several likely next tokens.
- Target model
- The authoritative model that verifies every committed output token.
- Acceptance rate
- The fraction of draft proposals that match what the target model would accept.
Keep in mind: Speculation can lose when draft cost is high or acceptance is low, and its benefit often shrinks at large serving batch sizes.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.For “A quick brown fox,” the draft proposes “jumps over a” before the target runs.
Let a cheap model propose three words, then use one target-model pass to verify the whole run.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Draft 3
The small model predicts three display tokens ahead
For “A quick brown fox,” the draft proposes “jumps over a” before the target runs.Live diagram
Draft several tokens, verify them together
Teaching diagram · illustrative values unless marked as measuredFast guesses; never committed without verification.
The prefix and all three draft words enter one forward pass. Each position can only see tokens to its left.
Run the target to accept or correct the draft.
Expected committed tokens per target pass: E = 1 + h + h². This teaching estimate assumes the same independent acceptance probability h at every position; a real draft model has position- and prompt-dependent acceptance. Verification stops at the first miss.
The displayed words are cheap proposals, not committed output; real tokenizers may split a word into multiple tokens.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
A real draft model accelerated Llama 3 on ShareGPT
- Model
- Llama-3-70B + Qwama-0.5B-Instruct draft
- Hardware
- 4× NVIDIA H100
- Benchmark
- ShareGPT · QPS 1
What this result shows · vLLM measured up to 1.5× faster token generation when the small Qwama draft model proposed tokens for Llama-3-70B.
What it does not prove · The same report reached 2.8× with n-gram speculation on CNN/DailyMail; acceptance rate and draft cost determine the gain.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · vLLM speculative decoding benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
Speculative Decoding, in one sentence
High acceptance produces more output tokens per target-model pass.Next chapter