From the previous chapter
Inference Optimization · Lesson 05
Continuous Batching
Why wait for the slowest request in a fixed batch?
Why this lesson comes next
Why wait for the slowest request in a fixed batch?
Admit and retire sequences at iteration boundaries so requests with different arrival times and output lengths can continuously share the GPU.Terms, if you need themIteration · Batch slot · Queueing+
- Iteration
- One scheduling and model-execution cycle that usually emits one token per active sequence.
- Batch slot
- One active sequence position sharing a GPU forward pass with other sequences.
- Queueing
- Time a request waits before the scheduler admits its next work.
Keep in mind: Continuous batching improves utilization, but an overloaded or unfair scheduler can still violate latency SLOs.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.Each prompt and output length is different, so a fixed group rarely finishes together.
At each token boundary, remove finished sequences and immediately fill their slots with waiting requests.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Arrive
Requests arrive at different times
Each prompt and output length is different, so a fixed group rarely finishes together.Live diagram
Refill a batch at every token step
Teaching diagram · illustrative values unless marked as measuredR1: 8 · R2: 5 · R3: 3 · R4: 6 cycles · 22 / 32 useful slots
R1 to R4 keep the same cycle counts; R5 and R6 raise use to 30 / 32 slots.
Generation is iterative: one decode step produces one token per active sequence.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Iteration-level scheduling raised GPT-3 serving throughput
- Model
- GPT-3 175B
- Hardware
- Distributed GPU cluster reported in the ORCA paper
- Benchmark
- Generative serving · iteration-level scheduling
What this result shows · ORCA measured 36.9× higher throughput than FasterTransformer at the same latency by scheduling and rebuilding the batch at iteration boundaries.
What it does not prove · ORCA also uses selective batching and distributed execution, so 36.9× is an end-to-end system result rather than a universal multiplier.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · ORCA, OSDI 202204 · Quiz
Can you reason through the concept?
What we can now explain
Continuous Batching, in one sentence
More active slots raise throughput and reduce avoidable queueing.Next lesson