LLM serving and cost
Tokens as the unit of latency and cost: batching, streaming, safe caching, routing and fallbacks, quotas and the usage ledger.
Serving an LLM is a capacity problem with language-shaped inputs: you meter tokens, protect latency, and choose when to stream, cache, batch, route, or shed load. This page gets you ready to explain the serving path and the bill without quoting a vendor price sheet.
Read this if your last attempt…
- Your last answer said “use an LLM API” and stopped before cost or latency
- You cannot explain why a short prompt can still produce a slow answer
- You mentioned caching but did not scope it by tenant and permissions
- You have not tied rate limits, quotas and billing to a usage ledger
The concept
Tokens are the cost and latency unit
Treat an LLM request as a token budget. Prompt tokens are the input: system instructions, conversation history, retrieved passages, tool results and the user message. Output tokens are the generated answer. Meter tokens in, tokens out and tokens per minute per tenant, not just requests per second.
The request is authenticated, checked against tenant budgets, routed, optionally cached, scheduled onto accelerator capacity, streamed back, and recorded in the usage ledger.
Serving levers: pick the lever that matches the bottleneck.
| Lever | Helps most with | Main risk |
|---|---|---|
| Output token cap | Tail latency and spend on verbose answers. | Answers become incomplete if the product copy does not set expectations. |
| Streaming with SSE | Perceived latency and early cancellation. | The backend still pays for generated tokens until work is stopped. |
| Continuous batching | Accelerator utilisation under mixed request lengths. | Fairness and memory pressure if one tenant floods the scheduler. |
| Exact response cache | Repeated deterministic tasks and public help content. | Stale or leaked answer when the key omits tenant, permission or data version. |
| Semantic cache | Near-duplicate public questions. | Wrong answer or cross-user leakage when the similarity match is too broad. |
| Prompt / prefix cache | Repeated system prompts, tools and long shared context. | It saves prefix work, not final-answer validation or access control. |
| Task-aware routing | Cost control across easy and hard tasks. | Quality regression if the router is not backed by evaluation data. |
| Tenant token budget | Runaway spend and abuse. | Poor UX if the limit fails closed without a downgrade path. |
How interviewers grade this
- You split latency into TTFT and decode time, then connect output length to user-visible delay.
- You meter prompt tokens and output tokens per request, per tenant, not just request count.
- You explain continuous batching as filling accelerator slots while other requests are decoding.
- You cache exact responses safely and treat semantic caching as a permission-scoped optimisation, not a free win.
- You route simple tasks to cheaper tiers, keep deadlines, and define a fallback or degradation path.
- You name the usage ledger as the source for billing, limits, abuse investigation and capacity planning.
Variants
Continuous batching scheduler
Keep the accelerator busy by admitting new sequences as old ones finish.
The scheduler admits new prompts while other requests are already decoding. Finished sequences leave the batch, and waiting requests join without waiting for a whole fixed batch to drain.
Use a scheduler in front of accelerator workers. It groups prompt work where possible, then keeps the decode loop full as streams finish at different lengths. The scheduler watches queue age, tenant fairness, max batch size and KV-cache memory pressure. This is the main conceptual upgrade from fixed interval batching.
Pros
- +Better utilisation for mixed output lengths.
- +Lower head-of-line blocking than fixed batches.
- +Gives one place to enforce tenant fairness.
Cons
- −Harder scheduler logic.
- −Needs live metrics on memory and queue age.
- −Over-admission can raise TTFT for everyone.
Choose this variant when
- Chat, assistants and mixed interactive workloads where response lengths vary widely.
Layered cache stack
Combine exact response caching, guarded semantic caching and provider prefix caching.
An LLM cache key is not just the prompt. It includes tenant, user or role, permission scope, model settings, tool data version and policy version before a response can be reused.
Exact response caching is the default safe layer when the full key is known. Semantic caching is reserved for low-risk near duplicates and must be scoped by tenant and permission. Provider prefix caching helps repeated context and tool definitions, but it does not reuse final answers. A strong design names which layer is used for which request class.
Pros
- +Can cut both latency and token spend.
- +Separates safe exact hits from fuzzy semantic hits.
- +Works well with long shared prompts.
Cons
- −Invalidation is product-specific.
- −Semantic matches can be wrong.
- −Cache hit metrics need to feed the usage ledger.
Choose this variant when
- Help-centre answers, repeated extraction jobs, long shared prompts and public documentation queries.
Task-aware model router
Route by task risk, context length, tenant plan and current load.
The router decides whether a request is simple classification, extraction, rewrite, search-grounded answer, or hard synthesis. Easy tasks take cheaper tiers and shorter output limits. Hard tasks get more capable routes when the tenant budget and deadline allow it. The router also owns timeouts, fallback provider choice and graceful degradation.
Pros
- +Direct cost control.
- +Keeps high-capacity routes for the requests that need them.
- +Creates a clear place for fallbacks and deadlines.
Cons
- −A bad router silently lowers answer quality.
- −Needs evaluation data from /learn/ai-evaluation.
- −Complexity grows with every provider and feature.
Choose this variant when
- Multi-feature assistants, tiered products, or any system with enough volume that routing mistakes matter.
Worked example
Numbers in this section are illustrative.
Scenario: an AI support assistant for a multi-tenant SaaS product. Each answer can use a system prompt, recent conversation, retrieved docs and a generated response. Numbers in this section are illustrative.
1. Start with token maths. Suppose an average request has 1,500 prompt tokens and 300 output tokens. At illustrative prices of $2 per million prompt tokens and $8 per million output tokens, that is about $0.003 input plus $0.0024 output, or $0.0054 per request. At 100,000 requests per day, the daily model-call cost is about $540 before cache hits, retries, tools and overhead.
2. Set latency targets. Suppose the target is TTFT under 1 second and complete response under 8 seconds. At 40 output tokens per second after TTFT, a 300-token answer takes about 8.5 seconds. A 120-token answer takes about 4 seconds, so concise defaults are a capacity control.
3. Serve the common path cheaply. Exact-cache common help answers with tenant, locale, role, docs version and policy version in the key. Use provider prefix caching for stable instructions. Route simple rewrite, sentiment and topic tasks to a small tier; reserve larger routes for long-context troubleshooting or ambiguous synthesis.
4. Batch, stream and meter. Continuous batching keeps accelerator slots full while fairness rules prevent one tenant from taking them all. SSE shows the first token quickly. The ledger records tenant_id, feature, prompt_tokens, output_tokens, cache_hit, route, provider_call_count, latency, status and cost estimate, so budgets, invoices and support all read the same facts.
A credible interview summary: “I control cost with token caps, safe caching, model routing and tenant budgets. I control latency with TTFT monitoring, streaming, continuous batching and output limits. The usage ledger ties both together.”
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“You are designing an LLM support assistant. How do you keep latency and cost under control?”
Weak answer
“I would stream the response and use caching. If it gets expensive, we can switch to a smaller model.”
Strong answer
“I would budget it in tokens. Each request records prompt tokens, output tokens, cache hit, route and tenant in a usage ledger. For latency I track TTFT versus decode time, stream with SSE, and cap output length by task. For throughput I use continuous batching with tenant fairness. For cost I check exact cache, use provider prefix caching for stable prompts, and allow semantic caching only with tenant and permission scope. The router sends simple extraction and classification to a cheaper tier, reserves larger routes for hard synthesis, and has timeouts plus fallback. Tenant budgets read from the ledger.”
Why it wins: It names the unit of control, the request path, the cache safety rule, the batching strategy, routing, fallbacks and the ledger that makes the policies enforceable.
When it comes up
- Any prompt with AI assistant, chatbot, agent, summariser or generation API.
- The interviewer asks how you control the bill or prevent one tenant from consuming the service.
- A design includes streaming responses or long context windows.
- The system has free-tier users, enterprise tenants or abuse risk.
- The interviewer asks how model choice changes by task.
Order of reveal
- 11. Start with tokens. “I meter prompt tokens and output tokens per request and per tenant. Requests per second is not enough for LLM serving.”
- 22. Split latency. “I track TTFT separately from decode time, because output length drives how long the user waits after the first token.”
- 33. Protect the hot path. “The path is quota check, safe cache, router, scheduler, stream and ledger. Each step has a timeout or policy.”
- 44. Use capacity levers. “Continuous batching keeps accelerator slots full, while output caps and tenant fairness keep one workload from taking over.”
- 55. Cache safely. “Exact cache keys include tenant, permissions and data version. Semantic cache is permission-scoped and conservative.”
- 66. Close with governance. “The usage ledger powers billing, limits, alerts and abuse response, so the cost story is enforceable.”
Signature phrases
- ““The unit is tokens, not requests.”” — Frames cost and capacity correctly.
- ““TTFT is the first promise; output length is the tail.”” — Separates perceived latency from full completion.
- ““I cache only what I can key safely.”” — Shows the permission-leak risk.
- ““The ledger is the contract between product, billing and abuse controls.”” — Connects metering to operations.
Likely follow-ups
?“How do you choose between small and large model tiers?”Reveal
Classify by task risk and evidence need. Extraction, classification, routing and simple rewrites usually go to a smaller tier with strict output caps. Ambiguous synthesis, long context and high-value workflows can use a larger tier. I would validate that router on the eval set from /learn/ai-evaluation and keep a fallback when the chosen tier times out or fails quality checks.
?“What do you cache?”Reveal
First exact responses with a full safety key: tenant, permission scope, prompt hash, model settings, data version and policy version. Then provider prefix caching for repeated system prompts and tool definitions. Semantic caching only for low-risk public or tenant-scoped content, with conservative similarity thresholds and freshness limits.
?“How do you stop one tenant from taking all capacity?”Reveal
Use per-tenant request and token budgets at the gateway, plus scheduler fairness on active sequences and KV-cache memory. The scheduler should see tenant identity, not just anonymous requests. The ledger records every attempt so budget counters, invoices and support investigations agree.
?“What happens when the provider is slow?”Reveal
Each route has a deadline. Before the deadline, I can retry once if the error is transient and safe. After that I degrade: shorter answer, smaller route, fallback provider, stale exact cache, async completion or a clear failure. I avoid fallback chains that exceed the user deadline and multiply cost.
Code examples
CREATE TABLE llm_usage_ledger (
request_id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
user_hash TEXT,
feature TEXT NOT NULL,
route TEXT NOT NULL,
prompt_tokens INTEGER NOT NULL,
output_tokens INTEGER NOT NULL,
cache_result TEXT NOT NULL, -- miss, exact_hit, semantic_hit, prefix_hit
provider_calls INTEGER NOT NULL,
latency_ms INTEGER NOT NULL,
estimated_cost_usd NUMERIC(12, 6) NOT NULL,
status TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX llm_usage_ledger_tenant_day
ON llm_usage_ledger (tenant_id, created_at);type Route = "small" | "large" | "async";
function chooseRoute(input: {
task: "classify" | "rewrite" | "answer" | "synthesise";
promptTokens: number;
tenantTokensLeft: number;
deadlineMs: number;
}): Route {
if (input.tenantTokensLeft < 2_000) return "small";
if (input.deadlineMs < 1_500) return "small";
if (input.task === "classify" || input.task === "rewrite") return "small";
if (input.promptTokens > 12_000) return "async";
return "large";
}Common mistakes
Prompt trimming helps, but long generated answers can dominate both latency and spend. Put max output tokens, answer style and stop conditions into the design. For many products, “short answer first, expand on request” is a capacity decision as well as a UX decision.
A cache hit that ignores tenant, role, document ACL or data version can leak private content. Exact caches still need a full key. Semantic caches need an even stricter scope because the lookup is fuzzy.
Streaming improves perceived latency and supports cancellation. It does not reduce accelerator work for tokens already generated. Keep measuring TTFT, full completion time and tokens generated after cancellation.
A fallback that waits for one provider timeout and then repeats the same long request elsewhere can double cost and still miss the user deadline. Give each route a budget and degrade deliberately: shorter answer, stale cache, async job or clear failure.
If limits read from Redis counters while billing reads from provider invoices, disputes are hard to debug. The request ledger should be the shared source for budgets, invoices, abuse review and capacity planning.
One short classification request and one long report are not equal. Track prompt tokens, output tokens, active sequences and KV-cache memory pressure. Otherwise the capacity plan looks healthy until a few long outputs fill the accelerators.
Practice drills
Numbers in this section are illustrative.
Why can output length dominate latency even when the prompt is large?Reveal
Prompt processing affects TTFT, but generated tokens are produced step by step after the first token. A long answer keeps the user waiting through many decode steps, so max output tokens and concise answer design are latency controls.
What belongs in a safe exact-response cache key?Reveal
Tenant, user role or permission scope, prompt hash, model settings, tool or retrieval data version, locale and policy version. If any of those dimensions can change the answer, omitting it can serve stale or private data to the wrong caller.
How is provider prefix caching different from response caching?Reveal
Prefix caching reuses work for a repeated prompt prefix such as system instructions or tool schemas. It does not reuse the final answer. Response caching stores and replays an answer, so it needs stricter correctness and permission checks.
A tenant hits its monthly token budget during a support chat. What are reasonable choices?Reveal
Fail clearly, ask the user to shorten input, route to a cheaper tier, cap output more aggressively, queue low-priority work, or require an admin to raise the budget. Which one fits depends on the product plan and whether the request is urgent.
Deep dives
GPU and accelerator capacity at interview depth
Accelerator capacity is where many LLM-serving answers get vague. At interview depth, you do not need to size a real cluster, but you should know what consumes the expensive hardware.
There are two phases. Prefill processes the prompt tokens and builds attention state. It can be heavy for long prompts, especially with retrieved context. Decode produces output tokens step by step. Decode is sensitive to active sequence count because each sequence carries KV-cache state and needs repeated small steps until it finishes.
The scheduler is managing a two-dimensional bin. One side is compute: how many token steps can the accelerators run per second under the current batch. The other side is memory: how much KV cache the active prompts and generated tokens occupy. Long prompts, long outputs and many concurrent streams all push memory. When memory is tight, the system may reduce batch size, evict low-priority work, spill, or reject requests.
Continuous batching is the serving answer to variable output length. Fixed batches wait for the slowest request. Continuous batching lets completed sequences leave and new ones enter, which keeps the accelerator busier and reduces head-of-line blocking. Fairness still matters: if you admit requests purely by arrival time, a noisy tenant can consume slots and cache memory while smaller tenants queue.
The practical capacity metric is tokens in flight by tenant and route. Track queued prompt tokens, active output tokens, TTFT, tokens per second, cache memory pressure and cancellation rate. Requests per second hides the difference between a 50-token classification request and, for example, a 5,000-token report.
Cache safety checklist
An LLM cache key is not just the prompt. It includes tenant, user or role, permission scope, model settings, tool data version and policy version before a response can be reused.
Caching can cut cost quickly, but the failure modes are sharper than in a normal HTTP cache. The response may depend on private data, current permissions, retrieval freshness, tool results, model settings and safety policy. The cache key has to represent those dependencies.
Exact caching starts with a strict key. Include tenant, user role or permission set, locale, prompt hash, model tier, temperature or sampling settings, tool version, retrieval corpus version and policy version. You can loosen the key only after you know which dimensions do not affect the answer.
Semantic caching needs a second gate because the key is fuzzy. A near match can be a different intent: “refund policy for admins” and “refund policy for customers” may embed close to each other but require different answers. For private data, cross-user leakage is the biggest risk. Scope the embedding index by tenant and permission, and prefer semantic caching for public docs or low-risk boilerplate.
Provider prefix caching is safer because it does not reuse an answer. It reuses work for a repeated prefix, such as system instructions or tool schemas. You still pay attention to cache invalidation: when the system prompt, policy text or tool list changes, the prefix should change too.
The interview phrasing is: “I cache only what I can key safely. Exact responses are easy to reason about, semantic hits need tenant and permission scoping, and provider prefix caching helps repeated context without crossing users.”
Cheat sheet
- •Meter tokens, not just requests: prompt tokens, output tokens and tokens per tenant.
- •Latency = TTFT plus decode time; output length is usually the tail.
- •Continuous batching fills accelerator slots as streams finish.
- •KV-cache memory makes long prompts and many streams a capacity problem.
- •SSE is a good fit for one-way token streams; use /learn/protocol-choice for trade-offs.
- •Exact cache keys include tenant, permission, prompt, settings, data version and policy version.
- •Semantic caches can be wrong or leaky; scope them tightly.
- •Route easy tasks to cheaper tiers, but validate the router on /learn/ai-evaluation.
- •Fallbacks need deadlines and deliberate degradation.
- •The usage ledger powers billing, quotas, abuse review and capacity planning.
Practice this skill
No problem is tagged directly to LLM serving and cost yet. These published problems still exercise the same interview category.
Read this if