Start here
Inference Foundations · Lesson 00.1
Autoregressive Generation
How does a transformer turn one prompt into a streamed answer?
Why this lesson comes next
How does a transformer turn one prompt into a streamed answer?
Separate prompt prefill from the token-by-token decode loop, then follow logits, token selection, append, and stop conditions using one sentence.Terms, if you need themToken · Logits · EOS+
- Token
- A model vocabulary unit. A visible word may be one token, several tokens, or part of a token.
- Logits
- Raw scores for every possible next token before they are converted into probabilities.
- EOS
- End-of-sequence: a special token that tells generation to stop.
Keep in mind: The diagram uses whole words for readability. Real tokenizers may split “quick” or “jumps” into smaller pieces, but the one-token-at-a-time loop is the same.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.The prompt ‘A quick brown fox’ is split into tokens and mapped to integer IDs before the transformer runs.
Process the prompt once, choose one next token, append it, and repeat until the model emits a stop token.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Tokenize
Text becomes model-readable token IDs
The prompt ‘A quick brown fox’ is split into tokens and mapped to integer IDs before the transformer runs.Live diagram
One prompt becomes a loop of next-token decisions
Teaching diagram · illustrative values unless marked as measuredPrefill: process the four prompt tokens together and build the initial state.
Tokens are model units, not guaranteed to be whole words.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Batching multiple autoregressive streams raised aggregate output
- Model
- Llama 3.1 8B
- Hardware
- NVIDIA Jetson AGX Thor
- Benchmark
- vLLM · 2,048 input + 128 output tokens
What this result shows · NVIDIA measured 41.3 tokens/s with one active stream and 150.8 tokens/s at concurrency eight, showing why serving engines interleave many independent decode loops.
What it does not prove · This is aggregate throughput, not proof that each individual request becomes faster; batching can increase queueing or per-request latency.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Jetson LLM benchmark04 · Quiz
Can you reason through the concept?
What we can now explain
Autoregressive Generation, in one sentence
Know which work controls first-token delay and which work controls the streaming pace.Next lesson