Skip to Autoregressive Generation lesson
LLM Inference Visualizer

Inference Foundations · Lesson 00.1

Autoregressive Generation

How does a transformer turn one prompt into a streamed answer?

Start here

Begin with one request before optimizing the system

Follow a single prompt into the model and watch its answer appear one token at a time. That loop becomes the shared baseline for every memory, scheduling, and serving technique that follows.

Why this lesson comes next

How does a transformer turn one prompt into a streamed answer?

Separate prompt prefill from the token-by-token decode loop, then follow logits, token selection, append, and stop conditions using one sentence.
Terms, if you need themToken · Logits · EOS
Token
A model vocabulary unit. A visible word may be one token, several tokens, or part of a token.
Logits
Raw scores for every possible next token before they are converted into probabilities.
EOS
End-of-sequence: a special token that tells generation to stop.

Keep in mind: The diagram uses whole words for readability. Real tokenizers may split “quick” or “jumps” into smaller pieces, but the one-token-at-a-time loop is the same.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementText becomes model-readable token IDs

The prompt ‘A quick brown fox’ is split into tokens and mapped to integer IDs before the transformer runs.

BottleneckSerial token dependency

Process the prompt once, choose one next token, append it, and repeat until the model emits a stop token.

Metric to watchTTFT + ITL
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Tokenize

Text becomes model-readable token IDs

The prompt ‘A quick brown fox’ is split into tokens and mapped to integer IDs before the transformer runs.

Live diagram

One prompt becomes a loop of next-token decisions

Teaching diagram · illustrative values unless marked as measured
Lesson 00.1 · 14 total
Input text“A quick brown fox”
Tokenizer
A32quick1974brown14198fox39935
One prompt passPrefill all 4 tokens togetherbuild context + initial K/V state
Prepare the promptRun prefill before decoding begins
Prompt pass
Current history
Aquickbrownfox
waiting…
waiting…
waiting…
waiting…
Model stateKVprompt context ready
The same loop repeatsEach committed token becomes input to the next step
1 / 17
Aquickbrownfox→

Prefill: process the four prompt tokens together and build the initial state.

Tokens are model units, not guaranteed to be whole words.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
PrefillPrompt togetherBuild initial state
DecodeOne step at a timeAppend and repeat
DependencySerialNext token needs the past

Actual benchmark

Batching multiple autoregressive streams raised aggregate output

Named model · measured workload
Output throughputtokens/s on one edge system
Batching multiple autoregressive streams raised aggregate outputConcurrency 1: 41.3 tok/s; Concurrency 8: 150.8 tok/sConcurrency 141.3 tok/sConcurrency 8150.8 tok/s
Model
Llama 3.1 8B
Hardware
NVIDIA Jetson AGX Thor
Benchmark
vLLM · 2,048 input + 128 output tokens

What this result shows · NVIDIA measured 41.3 tokens/s with one active stream and 150.8 tokens/s at concurrency eight, showing why serving engines interleave many independent decode loops.

What it does not prove · This is aggregate throughput, not proof that each individual request becomes faster; batching can increase queueing or per-request latency.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Jetson LLM benchmark

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01What does one autoregressive decode step produce?
02Why is the prompt normally processed differently from later output tokens?
03The model emits EOS after “dog”. What should the serving loop do?
Answer all three, then check your reasoning.

What we can now explain

Autoregressive Generation, in one sentence

Know which work controls first-token delay and which work controls the streaming pace.

Next lesson

A new bottleneck now comes into view

When is self-hosting worth the operational work?Continue to Why Local Models?
Course map