Skip to Continuous Batching lesson
LLM Inference Visualizer

Inference Optimization · Lesson 05

Continuous Batching

Why wait for the slowest request in a fixed batch?

From the previous chapter

Carry forward: Sparsification

Compare irregular zeros with hardware-friendly sparse patterns and reveal why a sparse model becomes faster only when its storage format and kernels agree.

Why this lesson comes next

Why wait for the slowest request in a fixed batch?

Admit and retire sequences at iteration boundaries so requests with different arrival times and output lengths can continuously share the GPU.
Terms, if you need themIteration · Batch slot · Queueing
Iteration
One scheduling and model-execution cycle that usually emits one token per active sequence.
Batch slot
One active sequence position sharing a GPU forward pass with other sequences.
Queueing
Time a request waits before the scheduler admits its next work.

Keep in mind: Continuous batching improves utilization, but an overloaded or unfair scheduler can still violate latency SLOs.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementRequests arrive at different times

Each prompt and output length is different, so a fixed group rarely finishes together.

BottleneckIdle batch slots

At each token boundary, remove finished sequences and immediately fill their slots with waiting requests.

Metric to watchThroughput + queueing
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Arrive

Requests arrive at different times

Each prompt and output length is different, so a fixed group rarely finishes together.

Live diagram

Refill a batch at every token step

Teaching diagram · illustrative values unless marked as measured
Lesson 05 · 14 total

Generation is iterative: one decode step produces one token per active sequence.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
Batch slotsRefilledEvery iteration
GPU idleLowerFewer empty rows
ThroughputHigherMore useful tokens

Actual benchmark

Iteration-level scheduling raised GPT-3 serving throughput

Named model · measured workload
Throughput at the same latency× FasterTransformer
Iteration-level scheduling raised GPT-3 serving throughputRequest-level batch: 1×; ORCA iteration scheduling: 36.9×Request-level batch1×ORCA iteration scheduling36.9×
Model
GPT-3 175B
Hardware
Distributed GPU cluster reported in the ORCA paper
Benchmark
Generative serving · iteration-level scheduling

What this result shows · ORCA measured 36.9× higher throughput than FasterTransformer at the same latency by scheduling and rebuilding the batch at iteration boundaries.

What it does not prove · ORCA also uses selective batching and distributed execution, so 36.9× is an end-to-end system result rather than a universal multiplier.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · ORCA, OSDI 2022

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01When can a continuous-batching scheduler safely change the active batch?
02What causes wasted rows in a static generation batch?
03A scheduler keeps the GPU at 99% utilization, but interactive requests wait behind long prefills and miss TTFT targets. What should change?
Answer all three, then check your reasoning.

What we can now explain

Continuous Batching, in one sentence

More active slots raise throughput and reduce avoidable queueing.

Next lesson

A new bottleneck now comes into view

How should growing KV caches be allocated?Continue to PagedAttention
Back to Sparsification