Service mesh (Envoy, Istio)
A service-to-service traffic layer that adds identity, mTLS, routing policy and telemetry through proxies rather than application libraries.
Also worth naming: Istio · Envoy mesh · Linkerd · Consul service mesh · ambient mesh
A service mesh is useful when internal calls have become their own platform problem. The trick is to know which concerns belong in the mesh, and which still belong in code or at the edge.
What it is
A service mesh is infrastructure for east-west traffic: calls between services inside your platform. In the common Istio sidecar model, each workload gets an Envoy proxy beside it; app code talks locally, and the proxy handles mutual TLS, service discovery, load balancing, retries, timeouts, policy checks and telemetry for the network hop. Istio describes this split as a data plane of Envoy proxies and a control plane that configures them in its architecture docs.
Do not draw it as the public front door. A load balancer chooses healthy instances, and an API gateway handles north-south client policy such as authentication, quotas and request shaping. A mesh handles service-to-service identity, routing and observation once traffic is already inside the platform. You often run all three: load balancer and gateway at the edge, mesh between services.
The safest interview framing is: "I use the mesh for transport security, workload identity, traffic policy and hop-level telemetry, while business retries, idempotency and product fallbacks still live in the application." The mesh makes the network programmable, but it does not make a bad dependency contract safe.
When to reach for it
Reach for this when…
- You have many microservices, and every team is re-implementing mTLS, retries, timeouts, routing and telemetry differently.
- Service-to-service identity and least-privilege policy matter more than IP-based network rules.
- You need canaries, weighted routing, locality-aware load balancing or outlier detection without rebuilding clients.
- You want consistent hop-level metrics and traces across polyglot services.
Not really this pattern when…
- A monolith or small service set can solve this with a gateway, a normal load balancer and disciplined libraries.
- The issue is public API policy, WAF, client auth or per-plan quotas; use an API gateway instead.
- The issue is only client-side balancing for one gRPC service; a gRPC resolver and client-side load-balancing policy may be cheaper.
- Your team cannot operate the control plane, policy lifecycle, proxy upgrades and incident debugging that come with a mesh.
How it works
1. Data plane vs control plane. The data plane is the set of proxies that actually see traffic. Envoy’s router matches a request to an upstream cluster, gets a connection pool, and forwards it; the same routing docs cover timeouts, retries, traffic splitting and weighted clusters in Envoy HTTP routing. The control plane turns platform state and mesh policy into proxy config. In Istio, istiod provides service discovery, configuration and certificate management, then sends Envoy config to proxies at runtime according to the Istio architecture docs.
2. xDS is the config stream, not the data path. Envoy’s xDS protocol describes discovery requests and responses, resource versions, ACK/NACK handling and updates when subscribed resources change. xDS tells proxies about listeners (LDS), routes (RDS), clusters (CDS), endpoints (EDS) and TLS secrets (SDS); Envoy names those services in its dynamic configuration overview. Existing proxies can keep serving with their last accepted config if the control plane is briefly down, but new pods, endpoints and policy changes may not converge.
3. Sidecar mode puts one Envoy beside each workload. Every inbound and outbound call crosses the local sidecar, then usually the peer sidecar. That gives rich L4 and L7 policy at the cost of one proxy per pod, extra hops in the request path, and application lifecycle coupling when sidecars are injected or upgraded.
Applications talk to local Envoy proxies. Istiod pushes xDS config and certificates; the proxies enforce mTLS, routing, policy and telemetry on each hop.
4. Ambient mode splits L4 and L7. Istio’s ambient overview says ambient uses a per-node L4 ztunnel, and optionally a waypoint Envoy proxy for L7 features. The current sidecar-or-ambient docs list ambient support as Stable for Kubernetes single-cluster, and the Istio project’s GA announcement dated 2024-11-07 says ambient reached General Availability in Istio 1.24 with ztunnel, waypoints and APIs marked Stable. The same announcement lists multi-cluster, multi-network and VM support as future work, so check your target version.
Ztunnel is the per-node secure overlay. Add a waypoint proxy when a namespace or service needs HTTP routing, request retries, L7 authorisation or request metrics.
5. Identity is workload identity, not only encryption. SPIFFE defines a workload and a SPIFFE ID such as spiffe://acme.com/billing/payments in the SPIFFE concepts docs; Istio’s mTLS verification docs show workload principals in the same URI form, for example spiffe://cluster.local/ns/default/sa/curl in ambient mTLS verification. Istio issues X.509 certificates to workloads and uses mTLS for peer authentication; its security docs describe automatic key and certificate provisioning and rotation in the identity flow. This is why mesh policy can say "orders may call payments" instead of naming pod IPs.
6. The mesh gives you hop-level controls, not product correctness. Istio traffic management covers virtual services, destination rules, retries, timeouts, circuit breakers, fault injection and percentage-based routing in its traffic management docs. Envoy outlier detection passively ejects hosts whose failures or latency are unlike peers according to the outlier detection docs. Those are powerful, but they work at the network hop. Idempotency, user-visible fallbacks and business decisions still belong with the service contract; see failure-mode analysis for the retry and circuit-breaker thinking.
Performance envelope
Service mesh envelope: exact figures below are from Istio 1.24 docs where stated; latency and capacity must still be measured on your workload.
| Concern | Sourced anchor | Design implication |
|---|---|---|
| Sidecar CPU and memory | [Istio 1.24](https://istio.io/v1.24/docs/ops/deployment/performance-and-scalability/) reports about 0.20 vCPU and 60 MB for one sidecar proxy with 2 worker threads at 1000 HTTP req/s, 1 KB payload. | Multiply by pod count; small services can pay more for the sidecar than for the app. |
| Waypoint and ztunnel resource use | [Istio 1.24](https://istio.io/v1.24/docs/ops/deployment/performance-and-scalability/) reports about 0.25 vCPU / 60 MB for one waypoint and about 0.06 vCPU / 12 MB for one ztunnel under the same benchmark shape. | Ambient can shift fixed cost from every pod to shared node or namespace components. |
| Latency overhead | Istio says sidecar, ztunnel and waypoint proxies sit on the data path, and every enabled feature increases proxy path length in the [1.24 performance page](https://istio.io/v1.24/docs/ops/deployment/performance-and-scalability/). | Budget extra hops and proxy work; do not put mesh retries on an already tight single-digit-ms path without measurement. |
| Telemetry cost | The same [Istio 1.24 page](https://istio.io/v1.24/docs/ops/deployment/performance-and-scalability/) says telemetry filters such as logging, tracing and metrics are known to have a moderate impact. | Use [observability](/learn/observability) deliberately: sample traces, bound label cardinality and avoid full debug logging on hot paths. |
| Control plane scale | [Istio 1.24 performance docs](https://istio.io/v1.24/docs/ops/deployment/performance-and-scalability/) say `istiod` CPU and memory scale with config changes, deployment changes and the number of proxies connected. | Scope config in large meshes; otherwise every proxy can carry routes and clusters it does not need. |
Capabilities in interviews
mTLS and workload identity
Encrypt service-to-service traffic and authenticate the workload on both sides.
A mesh can make every in-mesh call mutual TLS without each app loading certificates or choosing cipher suites. Istio’s security docs say peer authentication uses mTLS for service-to-service authentication, with each service receiving a strong identity and automated key/certificate generation, distribution and rotation in Istio security. The Istio policy API names matter: `PeerAuthentication` controls peer mTLS mode, and `AuthorizationPolicy` controls workload access decisions.
The useful design move is to attach policy to identity. For example, "the checkout service account may call payments on port 443" is more stable than allowing a pod IP range. Still link this to security and privacy: mTLS proves the peer and protects traffic in transit, but it does not replace resource-level authorisation in the service.
Choose this variant when
- Zero-trust service-to-service traffic
- Polyglot services that should not all own certificate code
- Least privilege by workload identity rather than IP
Retries, timeouts, circuit breakers and outlier detection
Apply hop-level resilience policy consistently, while keeping business safety in the app.
The caller uses one service name. Mesh config sends most traffic to the stable subset, a small share to the canary, and can eject failing endpoints.
Istio lets you set timeouts and retry policies on virtual services, and connection-pool/circuit-breaker settings on destination rules in its traffic management docs. Envoy circuit breakers cap cluster connections, pending requests, requests and active retries in the circuit breaking docs; outlier detection is separate passive health ejection. Envoy also supports retry budgets to limit retry traffic and avoid retry storms in the retry semantics docs.
The trap: mesh-level retries multiply with application-level retries. For example, if app code retries twice and Envoy also retries twice, one user request can become several downstream attempts before you notice. Keep one owner for retries, cap attempts, add jitter, propagate deadlines, and only retry idempotent operations. Link the detailed reasoning to failure-mode analysis.
Choose this variant when
- Transient network failures
- Known-safe idempotent reads
- Per-host failure isolation through outlier detection
Canaries and weighted routing
Shift traffic by percentage or request attributes without teaching every client about versions.
A virtual service can split traffic across service subsets, such as 90 percent to v1 and 10 percent to v2, while clients keep calling the stable host name; Istio shows this pattern in traffic management. That makes canaries and A/B tests a routing concern rather than a client rollout concern.
Use it with observability and rollback. A canary is only useful if you compare error rate, latency and saturation by version, then stop or roll forward based on those signals.
The caller uses one service name. Mesh config sends most traffic to the stable subset, a small share to the canary, and can eject failing endpoints.
Choose this variant when
- Canary releases
- Header- or user-based routing during migration
- Blue-green or staged rollouts behind one service name
Locality-aware load balancing
Prefer nearby healthy endpoints and shift away from unhealthy zones.
Envoy separates distributed load balancing, where the proxy chooses among endpoints it knows, from global load balancing, where the control plane assigns priorities, locality weights and endpoint health in the load balancing overview. In a multi-zone service, that means callers can usually stay in-zone, then spill over when local capacity or health is poor. Pair this with load balancing, because the same trade-offs still apply: locality lowers latency and cross-zone cost, but it can overload a small zone if you do not reserve failover headroom.
Choose this variant when
- Multi-zone clusters
- Cross-zone cost or latency matters
- You can tolerate controlled spillover during a zone incident
Hop-level observability
Get consistent request, connection and proxy metrics even when app stacks differ.
Envoy says it emits downstream, upstream and server statistics, including counters, gauges and histograms, in its statistics docs. Istio sidecars and waypoints can add request-level RED metrics, access logs and trace spans without changing every service framework.
What you do not get for free: business events, semantic success, user identifiers with safe cardinality, payload-level audit fields, useful span names inside your code, or a user-facing SLO. The mesh can tell you "payments returned 503 from checkout". It cannot tell you whether the checkout was a gift-card purchase unless the app emits that context. Link the operating model to observability.
Choose this variant when
- Polyglot fleet with inconsistent instrumentation
- You need per-hop latency and error-rate visibility
- You still plan application metrics for business outcomes
Sidecar-less adoption with ambient mode
Start with L4 mTLS and telemetry through ztunnel, then add waypoints for L7 policy.
Istio ambient mode uses ztunnel for the L4 secure overlay and optional waypoint proxies for L7 features, according to the ambient overview. The ambient data plane docs say L4 traffic between ambient workloads is secured by mTLS over HBONE, and a waypoint is needed for L7 policies, request routing and L7 load balancing. This is a better fit when you mainly want mTLS and identity first, and only some namespaces need advanced traffic management.
Choose this variant when
- You want mesh security without a sidecar in every pod
- You can accept current ambient platform limits
- L7 features are needed selectively rather than everywhere
Operating knobs
Sidecar vs ambient mode
Choose sidecars when you need the most mature, broadest Istio feature set, especially multi-cluster, multi-network or VM coverage. Choose ambient when you are on Kubernetes, single-cluster support fits, and you want to avoid per-pod sidecars for L4 security and telemetry. Add waypoints only where L7 policy, retries, traffic splitting or request metrics are needed. The docs currently say ambient is Stable for single-cluster use, while multi-cluster, multi-network and VM support remain outside ambient support in the sidecar-or-ambient page.
Who owns retries and deadlines
Pick one place to own each retry. A mesh retry is useful for a safe, idempotent read to another replica. Application code should own retries that need an idempotency key, compensation, product fallback or user-visible message. Either way, propagate an end-to-end deadline so inner timeouts are shorter than outer timeouts, and use retry budgets when the proxy owns retries.
Trust domain and certificate authority
A mesh identity is only as useful as its issuance boundary. Decide the trust domain, CA integration, certificate lifetime, rotation path and policy for cross-cluster or cross-tenant calls. SPIFFE trust domains map to trust roots, and X.509-SVIDs are signed by an authority in that trust domain in the SPIFFE concepts docs.
Config scope
Large meshes suffer when every proxy receives every route and endpoint. Use namespace scoping, sidecar resources or ambient waypoint scope so a checkout proxy does not carry search, analytics and admin-only config. Istio 1.24 performance docs recommend configuration scoping at large scale in the control plane section.
Versus the alternatives
Service mesh vs nearby answers.
| Need | Service mesh | API gateway | Load balancer | Library / gRPC client LB |
|---|---|---|---|---|
| Traffic direction | East-west service-to-service | North-south client-to-service | Instance selection for a pool | One client stack calling one service family |
| Identity | Workload identity, mTLS and service-level policy | Client identity, API keys, JWTs and coarse API policy | Usually not identity-aware beyond TLS termination | Depends on the app and platform |
| Routing | Weighted subsets, locality, outlier detection and retries between services | Path, host, version and product API routing | L4/L7 balancing and health checks | Resolver and client-side policy, often enough for gRPC |
| Best answer when | Many services need uniform internal traffic policy | Public API control and per-client quotas matter | You only need healthy instance distribution | A small fleet needs cheap client-side balancing without mesh ops |
| Main cost | Proxy resources, latency, policy complexity and control-plane operations | Every public request crosses it, so it must stay thin | Less app-aware; may not solve service identity | Each language and team must keep behaviour consistent |
- In interviews, draw the narrowest component that solves the problem. A mesh is not a replacement for a gateway or a load balancer.
Failure modes & gotchas
If the application retries and the mesh retries, one failing dependency receives multiplied attempts during the moment it is least able to handle them. Envoy explicitly recommends retry budgets or active-retry circuit breakers to avoid retry storms in its retry docs. Decide whether the app or the mesh owns retries, cap attempts, add jitter and propagate deadlines.
A sidecar per pod means CPU, memory, startup, upgrade and debugging cost per pod. The Istio 1.24 benchmark gives a concrete reminder: one sidecar proxy at the documented benchmark point consumes about 0.20 vCPU and 60 MB in the Istio 1.24 performance page. That may be fine for a large service, but painful for many tiny pods.
Ztunnel gives L4 mTLS, L4 authorisation and TCP telemetry; it does not parse workload HTTP headers. Istio’s ambient overview says waypoints are required for advanced traffic management, L7 authorisation, telemetry and VirtualService routing in ambient mode. If you need path, method, JWT claim or request-duration policy, put a waypoint in the path.
A bad route, broad authorisation policy or broken certificate path can be distributed quickly to many proxies. xDS ACK/NACK helps proxies reject invalid resources, but a valid bad policy is still bad policy. Use staged rollout of mesh config, linting, policy tests, canaries and fast rollback just as you would for application code.
In production
Istio project announcement
Ambient mode reaches General Availability in Istio 1.24
The Istio GA post says ambient mode reached General Availability in v1.24, with ztunnel, waypoints and APIs marked Stable by the Istio Technical Oversight Committee. It frames the architecture as shared L4 node proxies plus optional L7 waypoints, so teams can start with mTLS, L4 authorisation and telemetry, then add request routing and richer policy where needed. This is a project announcement rather than a customer production report; the same post keeps the boundary honest by listing future work such as multi-cluster, multi-network and VM support.
Checkout platform mesh migration (illustrative)
Use mesh policy for internal calls, not public API policy
The platform has a clear split: public clients enter through the API gateway, while internal services call each other through the mesh. The mesh is valuable because every internal call gets workload identity, mTLS and consistent telemetry. It is not asked to validate customer sessions, calculate discounts or decide whether a payment should be retried after a business decline.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“You have many microservices and want secure calls, canaries and consistent retries. Would you add a service mesh?”
Weak answer
"Yes, I would add Istio. It gives mTLS, retries, observability and canary deployments automatically, so services do not need to care about networking anymore."
Strong answer
"I would consider a mesh because the concerns are east-west and repeated across many services: workload identity with mTLS, traffic policy, canaries, locality-aware load balancing and hop-level telemetry. I would keep the gateway for north-south client policy and keep business idempotency and fallbacks in the apps. I would start with a narrow rollout: one namespace, strict mTLS, no mesh retries until deadlines and app retries are audited, then add weighted routing and outlier detection where we can measure success. If we only had a few gRPC services, I would use client-side load balancing and a shared resilience library instead."
Why it wins: It places the mesh on east-west traffic, names the risks, separates gateway/app responsibilities and proposes an incremental rollout rather than treating the mesh as magic.
Interview playbook
When it comes up
- The system has many microservices and repeated internal traffic policy.
- The interviewer asks how services authenticate each other or roll out canaries safely.
- You need to compare API gateways, load balancers, client libraries and meshes.
- You mention retries, mTLS, outlier detection or service-to-service observability.
Order of reveal
- 11. Draw the boundary. The mesh is for service-to-service traffic inside the platform; the gateway remains the public API boundary.
- 22. Split planes. Data plane proxies see traffic; the control plane pushes xDS config and certificates.
- 33. State the security model. mTLS gives workload identity and encryption in transit. Authorisation policy uses that identity, but service code still enforces resource permissions.
- 44. Add traffic controls carefully. Weighted routes and outlier detection are useful. Retries need one owner, deadlines and budgets to avoid amplification.
- 55. Pay the bill. I account for proxy CPU, memory, latency and operational complexity before saying yes.
Signature phrases
- “East-west in the mesh, north-south at the gateway.” — Keeps the diagram clean.
- “mTLS authenticates the workload; it does not authorise the row.” — Prevents overclaiming security.
- “Retries multiply unless one layer owns them.” — Catches the classic mesh outage pattern.
- “Ambient gives L4 by default; waypoints buy L7.” — Shows current Istio architecture knowledge.
Likely follow-ups
?“What do you get for free from the mesh, and what not?”Reveal
You get consistent hop-level telemetry such as request/connection metrics, access logs and, with L7 proxies, trace participation. You do not get business metrics, semantic success, useful span names inside code, safe high-cardinality attributes or a user-facing SLO. The app still emits domain signals and correlation context.
?“When would you choose a library instead?”Reveal
When the problem is narrow: one gRPC client stack needs resolver-based client-side load balancing, retries and deadlines, and the fleet is small enough to keep libraries consistent. A mesh earns its cost when several languages and teams need shared policy, identity and telemetry.
?“What changes in ambient mode?”Reveal
The L4 secure overlay moves to ztunnel per node, and L7 features move to waypoint proxies when needed. That lowers per-pod overhead, but you only get L7 routing, request retries, L7 auth and request metrics through waypoints. It is Stable for single-cluster use in the current docs; check unsupported features before choosing it.
Worked example
Numbers in this section are illustrative.
Setup. You are designing checkout for a retailer with separate cart, inventory, payment and fulfilment services. The platform team wants mTLS, canary deploys and consistent telemetry across Java, Go and Node services.
Mesh boundary. I keep clients going through the API gateway for auth, quotas and request validation. Inside the platform, I put checkout, inventory and payment in the mesh. That gives service identity and mTLS between workloads without adding certificate code to every app. I start in one namespace and require strict mTLS only after permissive telemetry shows all callers are meshed.
Traffic policy. I use weighted routing for canaries, for example sending a small slice of checkout-to-payment traffic to v2 while dashboards compare error rate, latency and saturation by version. For retries, I pick one owner. If checkout already retries payment with an idempotency key, the mesh does not retry that call. For safe inventory reads, the mesh may do one bounded retry inside the caller deadline.
Failure controls. I add outlier detection for unhealthy payment endpoints and short timeouts so slow failures do not fill checkout worker pools. The product fallback remains in code: payment failure returns a clear retryable checkout state; it is not hidden by the mesh. Connection-pool circuit breakers cap queued work; outlier detection ejects bad endpoints; user-visible behaviour stays explicit.
Observability. Mesh metrics show checkout to payment latency and status by destination, version and response class. The app adds business metrics such as checkout_authorised_total, payment_declined_total and idempotency-key dedupe counts. Together those answer both "which hop is slow?" and "did users actually buy tickets?"
Cost and rollout. I measure proxy CPU, memory and p99 latency before broad rollout. If most services need only L4 security, I consider ambient: ztunnel for mTLS and TCP telemetry, then waypoints only for checkout and payment where L7 routing and authorisation are needed.
Cheat sheet
- •Mesh = east-west service-to-service policy; gateway = north-south client policy; load balancer = healthy instance selection.
- •Data plane proxies see traffic. Control plane (`istiod`) pushes xDS config and certs.
- •Sidecar mode: Envoy per workload. Ambient: ztunnel per node; waypoint for L7.
- •mTLS authenticates workloads and encrypts traffic; app/resource authorisation still matters.
- •Traffic tools: timeouts, retries, connection circuit breakers, outlier detection, weighted routes, locality.
- •Retry storms happen when app retries multiply with mesh retries. Choose one owner and use budgets.
- •Free observability is hop-level. Business metrics, SLOs and useful spans still come from apps.
- •Costs: proxy CPU/memory, extra hops, config complexity, policy rollout and incident debugging.
Drills
Numbers in this section are illustrative.
A team says, "We have Istio mTLS, so service authorisation is solved." What do you say?Reveal
mTLS authenticates the peer workload and encrypts traffic in transit. You still need authorisation policy for which workload may call which service, and the service still needs resource-level checks such as whether this user can refund this order. Mesh identity is a strong input, not the whole decision.
The app retries a payment call twice with an idempotency key. Should the mesh also retry?Reveal
Usually no. The app owns the business-safe retry because it has the idempotency key and understands the payment result. If the mesh also retries, attempts multiply and can create a storm. Use mesh retries for safe idempotent reads or connection failures where the app is not already retrying, and keep everything inside the caller deadline.
In ambient mode, why do you need a waypoint for path-based policy?Reveal
Ztunnel handles L4 traffic: mTLS, L4 authorisation and TCP telemetry. It does not parse HTTP headers or paths. A waypoint is an Envoy L7 proxy, so it can enforce path, method, JWT-claim routing, L7 authorisation and request metrics.
When is gRPC client-side load balancing a better answer than a mesh?Reveal
When the need is narrow: one service family, one client stack, service discovery and balancing are enough, and you do not need shared mTLS policy, cross-language telemetry or central traffic rules. A mesh earns its cost when the concern repeats across many services and teams.
What it is