Latency budgeting
p50/p99 targets, per-hop budgets, tail latency mitigation.
This page gets you ready to put numbers on a request path before the interviewer finds the slow hop for you.
Read this if your last attempt…
- You didn't compute a per-hop latency budget
- You said "p99 is fine" without stating the p99
- You added network hops without budgeting them
- Your design sends serial calls when parallel would save time
The concept
A latency budget is the end-to-end target broken into per-hop allowances. Pick the user-visible target first (e.g. 200 ms p95 for a feed load), then subtract the unavoidable (client RTT, TLS handshake, edge → origin) and allocate what's left across your internal hops.
The three distinctions that matter
- p50 vs p99 vs p99.9 — p50 is what most users see; p99 is what your heaviest users hit every session; p99.9 is where the tail problems hide. Sizing a distributed system on p50 almost always gives you an unhappy p99.
- Serial vs parallel fan-out — N serial calls cost
sum(latencies); N parallel calls costmax(latencies). When calls are independent, parallelise them. - Tail amplification — if a request fans out to N backends in parallel, its latency is the max of N random latencies, so its percentiles come from each backend's deeper tail. For independent backends, P(all N finish under T) = F(T)ᴺ, where F is one backend's latency CDF. For the request's p99 you need F(T)ᴺ = 0.99, so F(T) = 0.99^(1/N). At N=100 that's ≈ 0.9999: the request p99 is each backend's p99.99. Correlated slowness (shared GC, a hot network link) makes it worse.
User target → subtract client & edge → allocate across app → storage → return. Every hop needs a number, not a vibe.
Tail-latency techniques — cost vs savings.
| Technique | Saves | Costs |
|---|---|---|
| Hedged requests (fire second at p95) | p99 / p99.9 | ~5% extra requests when tails are independent; cap it and hedge only idempotent reads |
| Request cancellation | Steady-state load from hedging | Cancellation token plumbing |
| Parallel fan-out (fanning independent calls) | Serial latency sum | More concurrent connections |
| Timeouts at every hop | Unbounded stragglers | Must tune carefully — too tight = false fails |
| Circuit breakers | Latency under partial failure | Complexity; must exercise in chaos testing |
How interviewers grade this
- You state the user-visible target and the percentile (e.g. "p95 < 200 ms").
- Every hop in your diagram has a latency number.
- You identify the serial critical path and propose parallelisation where possible.
- You distinguish p50 sizing from p99 sizing — and size infrastructure for the higher percentile.
- You have a tail-latency plan (hedging, cancellation, circuit-breaking) when fan-out is wide.
Variants
Budget-first design
Pick user target, subtract fixed costs, allocate the rest to hops.
The discipline: start with a number, end with an allocation. Any hop that can't fit its budget is a design problem you see before launch, not after.
Pros
- +Forces a concrete constraint
- +Surfaces overbudget hops early
Cons
- −Takes 5 minutes of interview time
Choose this variant when
- Any new design
- Any performance-critical workflow
Parallel fan-out
Fire independent calls concurrently; wait for all (or quorum).
Cuts the sum to the max. But the max of N calls is set by each backend's deeper tail — the more backends, the worse the request p99 unless you hedge, cap fan-out, or accept partial results.
Pros
- +Sum of latencies → max of latencies
- +Essential for feeds, search, aggregation
Cons
- −Amplifies tail latency as N grows
- −Wastes work on cancelled calls
Choose this variant when
- Independent calls
- Aggregation / fan-out queries
Hedged requests
After p95 elapses, fire a duplicate; first to return wins.
The technique from Dean and Barroso's "The Tail at Scale". Hedging at the backend's p95 adds roughly 5% more requests when slow responses are independent, and can cut p99 sharply in systems with a slow tail. Guardrails: hedge only idempotent reads, cap the hedge rate, and suppress hedges during correlated slowdowns.
Pros
- +p99/p99.9 reduction without capacity changes
- +Works behind any replicated backend
Cons
- −Slightly higher steady load
- −Needs idempotent requests and a hedge-rate cap
- −Makes correlated incidents worse if uncapped
Choose this variant when
- Replicated read path
- Fan-out > 10
- Tail is dominated by a few slow shards
Worked example
Target: feed-load p95 < 300 ms end-to-end.
- Client RTT: 50 ms (fixed).
- TLS + edge: 30 ms.
- Remaining server budget: 220 ms.
Server path (serial):
- Auth check: 5 ms (cached token).
- Feed-compose service: ~100 ms (4 parallel backend calls, max of them).
- Timeline (user tweets): p99 80 ms. - Follow graph: p99 40 ms. - Ads: p99 60 ms. - Feed rank ML: p99 100 ms.
- Hydration (parallel): 40 ms.
- Serialization + response: 10 ms.
Total server: 5 + 100 + 40 + 10 = 155 ms. Under budget with 65 ms of slack. (Summing per-hop p99s like this is deliberately conservative; confirm the real end-to-end p95 from traces.)
With 4 parallel backends, the compose step's p99 is at least the slowest backend's p99, and for independent backends it's set by each backend's slightly deeper tail (0.99^(1/4) ≈ 0.9975, so roughly each one's p99.75). If one backend regresses to 200 ms p99, the whole request p99 jumps with it. Add hedging on the Feed rank ML call (slowest tail) to pull p99 back.
Good vs bad answer
Interviewer probe
“What's your latency target and how did you allocate it?”
Weak answer
"It'll be fast. Redis is fast, gRPC is fast. Should be under a second."
Strong answer
"p95 < 300 ms end-to-end. After 80 ms of client + edge, server has 220 ms. Auth 5, feed compose 100 (4 parallel calls, max of them), hydration 40, serialize 10. Total 155 — 65 ms slack. Feed rank is the tail offender at p99 100 ms; I'd hedge it after 80 ms elapsed. Everything else is well-behaved."
Why it wins: Names a percentile, allocates per hop, identifies the tail offender, proposes a specific mitigation.
When it comes up
- Right after non-functional requirements — whenever you agree to a latency SLO
- During deep-dive on a read path with any fan-out or remote calls
- When the interviewer asks "how fast does this need to be?"
- When you propose adding a service, cache layer, or remote hop
- Whenever "p99" or "tail latency" enters the conversation
Order of reveal
- 1Commit to a target AND a percentile. "Let's target p95 < 300 ms for the feed load. I'll allocate a per-hop budget so we can see if the design fits before we get deep into components."
- 2Subtract the unavoidable first. "Client RTT ~50 ms and edge/TLS ~30 ms are fixed. That leaves ~220 ms for the server path."
- 3Allocate per hop and point at the slack. "Auth 5, feed compose 100, hydration 40, serialize 10 — that's 155 ms with 65 ms of slack. Every hop has a number, not a vibe."
- 4Identify the critical path and parallelise. "These three calls are independent, so they run in parallel — the cost is max(80, 40, 60) = 80, not the 180 sum."
- 5Call out the tail explicitly. "At 4-way fan-out the overall p99 is at least the worst backend's p99, and it's driven by each backend's slightly deeper tail. Feed rank is the slow backend, so I'd hedge that one after 80 ms."
- 6Add timeouts and a fallback. "Every remote hop gets a timeout tighter than its budget, plus a fallback: stale cache on the feed service, degraded result on ads."
Signature phrases
- “Pick a target AND a percentile” — Prevents the "it'll be fast" hand-wave that weak candidates fall into.
- “Serial is sum, parallel is max” — One-line mental model for restructuring the call graph.
- “Wide fan-out amplifies tails” — Shows you know the Tail-at-Scale result, not just the average case.
- “Hedge the slow backend, not everything” — Signals calibrated use — hedging isn't free.
- “A hop without a timeout is a p99 disaster waiting to happen” — Concrete operational discipline.
- “Size for the percentile you committed to” — Catches the common p50-sizing mistake.
Likely follow-ups
?“Walk me through a hedged request in detail — when does the second fire, what happens if both return, and what's the downside?”Reveal
Trigger: start a timer when the first request goes out. If it hasn't returned by some quantile of the latency distribution (typically p95 of that backend), fire a second request to a different replica.
Resolution: take whichever response arrives first. Cancel the other in-flight request (tied requests take this further — the second backend cancels itself if it sees the first one is already executing).
Effect on load: in steady state only ~5% of requests spawn a hedge (since only ~5% exceed p95). The backend sees about 5% extra traffic, not double.
Downside 1 — duplicate work if backends aren't idempotent. Reads are fine; writes need careful thought or idempotency keys.
Downside 2 — correlated hedging. If all backends are slow together (GC pause, thundering herd), hedging fires everywhere at once and amplifies the incident. The fix is a cap: never more than X% of in-flight requests hedged, back off under load.
Downside 3 — metric pollution. Your "request latency" histogram now counts whichever response arrived first, which makes the underlying backend p99 harder to see. Instrument both.
?“Why does a 100-way fan-out have a p99 much worse than any individual backend's p99?”Reveal
The overall latency is max(L₁, L₂, ..., L₁₀₀) where each Lᵢ is drawn from the backend's latency distribution.
P(max < T) = P(L < T)¹⁰⁰. So for the overall p99 = max < T₉₉, we need P(L < T)¹⁰⁰ = 0.99, which means P(L < T) ≈ 0.9999 — we need each backend at its p99.99, not its p99, to get an overall p99.
Concretely: if each backend is p99 = 100 ms but p99.99 = 500 ms, the overall p99 of a 100-way fan-out is ~500 ms, not 100 ms.
Three mitigations:
- 1Reduce fan-out — aggregate at an intermediate tier so each request only spreads to e.g. 10 backends, then 10 of those aggregators merge.
- 2Hedge the slow tail — fire a duplicate to a replica once you hit the p95 of the individual distribution.
- 3Tolerate partial results — return after N-K backends respond (quorum). Works only if the missing K are tolerable in the result (e.g., search ranking).
?“How do you set the timeout for a hop? What's the right policy?”Reveal
Two constraints compete:
- Too tight → false failures on the slow tail; users see errors when the backend would have eventually answered.
- Too loose → the hop drags the whole request into its worst tail, defeating the budget.
Starting rule: timeout = backend p99.9 × 1.2, clamped to the remaining budget. Measure the p99.9 in production, don't guess.
Propagate a deadline, not just a timeout. Every request carries "you have X ms left". Each hop computes its timeout from the remaining deadline and the rest of the call graph. That way a slow early hop tightens the later hops automatically and the request either completes or fails fast.
Layer with retries carefully. Retry + timeout can multiply: a 1 s timeout with 2 retries can block up to 3 s. Either use a global deadline that caps all attempts, or set the retry budget as a percentage of the remaining deadline.
Fallback plan: every hop needs one. Stale cache, degraded result, or a cached default. A hard failure on a non-critical hop should never fail the whole request.
?“Your p95 is fine but p99 is terrible. Where do you look first?”Reveal
Signal that something happens to ~1% of requests but not the rest. Five usual suspects in order:
- 1GC pauses / JIT compilation / cold caches. Check GC logs; long pauses spike p99 without touching p50. Mitigation: tune GC (ZGC/Shenandoah for JVM), pre-warm caches on deploy.
- 1One slow shard or replica. The 1% of requests that hash to a degraded node. Check per-shard latency histograms, not just global. Fix: hedge, remove the bad replica, or rebalance.
- 1Lock / connection pool contention. Check pool wait times. If p99 wait >> p50 wait, the pool is undersized or there's a slow query holding a connection.
- 1Tail of a dependency. Your p99 is often someone else's p99.9. Drill down to the slowest hop in tracing; the bad actor is usually one call.
- 1Request size outliers. A few requests do much more work (large result sets, fat payloads). Segment the latency histogram by payload size.
Tool: distributed tracing with per-span percentiles is the fastest way to find the culprit. Global histograms only tell you p99 is bad — per-span tells you where.
?“How much budget should I spend on the database?”Reveal
A useful rule: cache hits ~1 ms, primary-key reads ~5-10 ms, indexed queries ~20-30 ms, unindexed or aggregations ~100 ms+. Anything you can push to a cache gets ~1 ms; the DB budget should cover only the cache misses.
For a 220 ms server budget with a 90% cache hit ratio:
- Hot path: 1-2 ms (cache hit, ~90% of requests).
- Cold path: 20-50 ms (one indexed DB query, ~10% of requests).
- Amortised per-request DB cost: ~0.9×1 + 0.1×30 = ~4 ms.
If your design needs more than one DB query on the hot path, either:
- 1Pre-compute and cache the composite result.
- 2Materialise a denormalised view you read in one query.
- 3Fan out to read replicas in parallel — p99 becomes max, not sum.
And if you're doing writes: writes are strictly more expensive than reads. A p95 < 100 ms on a write path usually means write-ahead log append + async materialisation, not synchronous commit-through to every index.
Common mistakes
The average user's experience is not the problem. Size on the percentile you've committed to — usually p95 or p99.
Three serial 50 ms calls is 150 ms; three parallel is 50 ms. Fan out anything independent.
A hop without a timeout drags the whole request into the slow tail. Every remote call needs a timeout that fits the hop's budget.
The p99 of a 100-way fan-out is set by each backend's ~p99.99 (for independent backends, 0.99^(1/100) ≈ 0.9999). Hedge, cap fan-out, pre-aggregate, or accept partial results.
Practice drills
User target is 500 ms p99. Client RTT 100 ms, TLS 50 ms. You have 3 serial calls server-side at 80, 120, 60 ms p99. Over or under budget?Reveal
Under, with about 90 ms of slack, if you budget by summing: 100 + 50 + 80 + 120 + 60 = 410 ms. That sum is a conservative planning number, not the true p99 of the total: for independent hops the end-to-end p99 is usually lower than the sum of the p99s, because the hops rarely all hit their tail on the same request. Correlated slowness (a GC pause, a congested link) can push it back up, so confirm with end-to-end measurements. 90 ms is thin slack; parallelise the calls if they are independent (the server part becomes ~max(80, 120, 60) = 120 ms), or cache the slowest call.
Interviewer: "you have a 50-way fan-out; each backend is p99 = 50 ms. What's the overall p99?"Reveal
Not 50 ms. For independent backends, P(all 50 finish under T) = F(T)⁵⁰, where F is one backend's latency CDF. For the request p99 you need F(T)⁵⁰ = 0.99, so F(T) = 0.99^(1/50) ≈ 0.9998: the request p99 is each backend's p99.98. The answer depends on how heavy the tail is beyond p99 — if a backend's p99.98 is 150 ms, the request p99 is about 150 ms. Mitigations: hedge, reduce fan-out, pre-aggregate, or return partial results.
You add a new hop costing 20 ms p99. Your budget was already tight. What do you do?Reveal
Options in order: (1) see if it can be parallelised with an existing hop (free); (2) make it async if the result isn't needed in the response (free); (3) cache it if its inputs are stable (cheap); (4) push back on the feature until the budget is renegotiated (political, often correct).
Cheat sheet
- •Pick a target + percentile. "p95 < 300 ms", not "it'll be fast".
- •Every hop gets a budget number.
- •Serial = sum. Parallel = max. Prefer parallel for independent work.
- •Wide fan-out amplifies tails: request p99 needs each backend at 0.99^(1/N). Hedge the slow backends.
- •Summing per-hop p99s is a conservative budget, not the true end-to-end p99. Measure with traces.
- •Timeout at every hop. Default to < budget so failures are bounded.
- •Watch p99 and p99.9 separately — they are different problems.
Practice this skill
These problems exercise Latency budgeting. Try one now to apply what you just learned.
Read this if