Skip to Why Local Models + SLOs lesson
LLM Inference Visualizer

Inference Foundations · Lesson 00.2

Why Local Models + SLOs

When should we run locally, and how do we prove the service is good enough?

From the previous lesson

Carry forward: Autoregressive Generation

Separate prompt prefill from the token-by-token decode loop, then follow logits, token selection, append, and stop conditions using one sentence.

Why this lesson comes next

When should we run locally, and how do we prove the service is good enough?

Choose between hosted, self-hosted, and on-device inference, then define TTFT, token pace, end-to-end latency, goodput, reliability, and quality targets for that deployment.
Terms, if you need themSelf-hosted · SLO · Goodput
Self-hosted
Your team operates the inference runtime on hardware it controls, locally or in its cloud account.
SLO
A measurable service objective, such as 90% of requests receiving their first token within 300 ms.
Goodput
The rate of useful requests or tokens that satisfy the selected latency, reliability, and quality gates.

Keep in mind: Local is not automatically cheaper, safer, or faster, and metric names alone are not an SLO. State the workload, percentile, hardware, measurement boundary, and quality gate.

01

Current setup & bottleneck

See where the existing system loses time or capacity

Start with the unoptimized path before introducing the technique.
Before the improvementA hosted API crosses an external service boundary

The provider operates the model and hardware, while prompts, outputs, availability, pricing, and model changes depend on that service.

BottleneckDeployment choice + success criteria

First choose who owns the model and data boundary; then prove that deployment meets explicit latency, reliability, quality, and cost objectives.

Metric to watchCost + latency + goodput
02

Improvement visualization

Apply the technique and follow what changes

Advance the live diagram one idea at a time.

Boundary

A hosted API crosses an external service boundary

The provider operates the model and hardware, while prompts, outputs, availability, pricing, and model changes depend on that service.

Live diagram

Choose the deployment boundary, then define what good looks like

Teaching diagram · illustrative values unless marked as measured
Lesson 00.2 · 14 total
Hosted APIOperations move to a provider
Your appprompt + output
cross network →
Provider model
  • Fast access to managed models
  • Elastic capacity and upgrades
  • Usage, policy, and data boundary depend on provider
Local / self-hostedData and runtime stay in your boundary
Your appprivate link →Your model server
  • Control model version, logs, and retention
  • Can work offline or near the data source
  • You own capacity, updates, security, and failures
Decision matrixChoose from the workload, not from ideology
QuestionLocal tends to fitHosted tends to fit
Data boundarySensitive, offline, low-latency edgeApproved external processing
DemandPredictable and well utilizedBursty or rapidly changing
Model needTask-sized model passes qualityFrontier capability is required
TeamCan operate GPU serving reliablyWants managed infrastructure
Choose localControl outweighs operations
Choose hostedCapability and elasticity win
Choose hybridKeep steady/private work local; burst elsewhere

Local is a deployment choice, not a guarantee of lower cost, stronger security, or better quality.

Interactive loadWatch latency, throughput, and goodput separate
One request · “A quick brown fox …”User-visible timeline
Send0 ms
TTFT184 ms
jumpsfirst token
ITL36 ms
overtoken 2
ITL38 ms
thetoken 3
ITL35 ms
dogfinal token
TTFT · initial wait184 mssend → first token
ITL · each gap36 mstoken → next token
TPOT · average pace36 ms/token(E2E − TTFT) ÷ later tokens
E2E · complete answer364 mssend → final token
Raw throughput72 tok/sall completed token work
SLO attainment100%
requests passing every gate
Request goodput4.0 req/suseful completions within the SLO
Responsivenessp90 TTFT ≤ 300 mspassing
Streaming pacep90 TPOT ≤ 50 mspassing
ReliabilityAvailability ≥ 99.9%passing
QualityTask evaluation ≥ targetpassing

Teaching simulation: metric definitions and relationships are exact; the slider values are illustrative, not a benchmark result. The measured model result appears below.

Hosted inference usually minimizes infrastructure work and maximizes access to frontier models.

1 / 4
03

How the numbers usually move

Connect the mechanism to measurable outcomes

Read the direction first, then inspect source-specific results and trade-offs.
BoundaryLocal · hosted · hybridControl + responsibility
User latencyTTFT · ITL · E2EWait + stream + finish
Fleet goodputWork within SLOUseful capacity

Actual benchmark

Task-sized models can change local generation speed dramatically

Named model · measured workload
Single-stream output throughputtokens/s · same edge system
Task-sized models can change local generation speed dramaticallyLlama 3.3 70B: 4.7 tok/s; Llama 3.1 8B: 41.3 tok/sLlama 3.3 70B4.7 tok/sLlama 3.1 8B41.3 tok/s
Model
Llama 3.3 70B vs Llama 3.1 8B
Hardware
NVIDIA Jetson AGX Thor
Benchmark
vLLM · concurrency 1 · 2,048 input + 128 output tokens

What this result shows · On the same Jetson AGX Thor, NVIDIA reports 41.3 tokens/s for Llama 3.1 8B versus 4.7 tokens/s for Llama 3.3 70B, illustrating the performance effect of model size at the edge.

What it does not prove · The models have different capabilities and quality. Choose the smaller model only after it meets the task’s accuracy and safety requirements.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Jetson LLM benchmark

Actual benchmark

Meeting both TTFT and TPOT objectives changed usable capacity

Named model · measured workload
SLO-attaining request rate× vLLM at 90% attainment
Meeting both TTFT and TPOT objectives changed usable capacityvLLM baseline: 1×; DistServe: 4.3×vLLM baseline1×DistServe4.3×
Model
OPT-66B
Hardware
NVIDIA A100 GPU cluster reported in the paper
Benchmark
LongBench summarization · joint TTFT + TPOT SLO · 90% attainment

What this result shows · For OPT-66B summarization, DistServe sustained 4.3× the vLLM request rate while meeting the paper’s TTFT and TPOT objectives for at least 90% of requests.

What it does not prove · The multiplier comes from one system, workload, and SLO definition. Your goodput must be measured with your own request lengths, arrival pattern, quality gate, and hardware.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · DistServe, OSDI 2024

04 · Quiz

Can you reason through the concept?

3 conceptual questions
01What is the clearest privacy benefit of local inference?
02When can a hosted API remain the better engineering choice?
03Two systems deliver 100 requests/s. System A has 92% SLO attainment; System B has 70%. Which has more SLO-qualified request goodput?
Answer all three, then check your reasoning.

What we can now explain

Why Local Models + SLOs, in one sentence

Choose the boundary deliberately and optimize for SLO-attaining goodput instead of one flattering average.

Next chapter

A new bottleneck now comes into view

Now reduce repeated work inside the autoregressive loop.Continue to KV Cache
Back to Autoregressive Generation