From the previous lesson
Inference Foundations · Lesson 00.2
Why Local Models + SLOs
When should we run locally, and how do we prove the service is good enough?
Why this lesson comes next
When should we run locally, and how do we prove the service is good enough?
Choose between hosted, self-hosted, and on-device inference, then define TTFT, token pace, end-to-end latency, goodput, reliability, and quality targets for that deployment.Terms, if you need themSelf-hosted · SLO · Goodput+
- Self-hosted
- Your team operates the inference runtime on hardware it controls, locally or in its cloud account.
- SLO
- A measurable service objective, such as 90% of requests receiving their first token within 300 ms.
- Goodput
- The rate of useful requests or tokens that satisfy the selected latency, reliability, and quality gates.
Keep in mind: Local is not automatically cheaper, safer, or faster, and metric names alone are not an SLO. State the workload, percentile, hardware, measurement boundary, and quality gate.
Current setup & bottleneck
See where the existing system loses time or capacity
Start with the unoptimized path before introducing the technique.The provider operates the model and hardware, while prompts, outputs, availability, pricing, and model changes depend on that service.
First choose who owns the model and data boundary; then prove that deployment meets explicit latency, reliability, quality, and cost objectives.
Improvement visualization
Apply the technique and follow what changes
Advance the live diagram one idea at a time.Boundary
A hosted API crosses an external service boundary
The provider operates the model and hardware, while prompts, outputs, availability, pricing, and model changes depend on that service.Live diagram
Choose the deployment boundary, then define what good looks like
Teaching diagram · illustrative values unless marked as measuredcross network →Provider model
- Fast access to managed models
- Elastic capacity and upgrades
- Usage, policy, and data boundary depend on provider
- Control model version, logs, and retention
- Can work offline or near the data source
- You own capacity, updates, security, and failures
Local is a deployment choice, not a guarantee of lower cost, stronger security, or better quality.
Teaching simulation: metric definitions and relationships are exact; the slider values are illustrative, not a benchmark result. The measured model result appears below.
Hosted inference usually minimizes infrastructure work and maximizes access to frontier models.
How the numbers usually move
Connect the mechanism to measurable outcomes
Read the direction first, then inspect source-specific results and trade-offs.Actual benchmark
Task-sized models can change local generation speed dramatically
- Model
- Llama 3.3 70B vs Llama 3.1 8B
- Hardware
- NVIDIA Jetson AGX Thor
- Benchmark
- vLLM · concurrency 1 · 2,048 input + 128 output tokens
What this result shows · On the same Jetson AGX Thor, NVIDIA reports 41.3 tokens/s for Llama 3.1 8B versus 4.7 tokens/s for Llama 3.3 70B, illustrating the performance effect of model size at the edge.
What it does not prove · The models have different capabilities and quality. Choose the smaller model only after it meets the task’s accuracy and safety requirements.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · NVIDIA Jetson LLM benchmarkActual benchmark
Meeting both TTFT and TPOT objectives changed usable capacity
- Model
- OPT-66B
- Hardware
- NVIDIA A100 GPU cluster reported in the paper
- Benchmark
- LongBench summarization · joint TTFT + TPOT SLO · 90% attainment
What this result shows · For OPT-66B summarization, DistServe sustained 4.3× the vLLM request rate while meeting the paper’s TTFT and TPOT objectives for at least 90% of requests.
What it does not prove · The multiplier comes from one system, workload, and SLO definition. Your goodput must be measured with your own request lengths, arrival pattern, quality gate, and hardware.Independent case study · compare the bars inside this card, not numbers across different lessons.Read the original source · DistServe, OSDI 202404 · Quiz
Can you reason through the concept?
What we can now explain
Why Local Models + SLOs, in one sentence
Choose the boundary deliberately and optimize for SLO-attaining goodput instead of one flattering average.Next chapter