Feature flags and experiments
Flags evaluated locally from a cached ruleset, safe defaults and kill switches, consistent bucketing, exposure events and sample-ratio checks.
Feature flags let you change behaviour after deploy without turning your runtime into a remote-control dependency. This page gets you ready to design local evaluation, fast kill switches, experiment bucketing and flag cleanup.
Read this if your last attempt…
- You put a network call on every flag check.
- Your browser SDK can see rules that mention private segments or plan names.
- You can ramp traffic but cannot explain whether users stay in the same bucket.
- Your incident plan says "toggle the flag" but does not say how quickly the toggle propagates.
The concept
Feature flags are runtime decisions with governance around them. Application code asks "which value should this context get?" and the SDK returns a typed value. OpenFeature standardises this shape: evaluation methods take a flag key and a default value, and detailed evaluations can return value, variant, reason and error metadata (OpenFeature flag evaluation). Providers are the vendor or local resolver behind those calls and can wrap a vendor SDK, REST client or local file (OpenFeature providers).
Evaluation is on the hot path
Treat a flag check like a branch in your code, not like a remote API call. Server-side SDKs should evaluate from a local cache of flag rules and segments. One product's approach is LaunchDarkly: server-side and edge SDKs evaluate internally using embedded evaluation rules, and server-side SDKs keep transmitted flag data in local caches by default (Flag evaluation rules, Choosing an SDK type).
Application code asks the SDK for a typed value. The SDK evaluates from a cached ruleset and returns a coded fallback if evaluation cannot complete.
Where should a flag be evaluated?
| Approach | Use it when | Watch out for |
|---|---|---|
| Server-side local evaluation | Trusted backend hot paths. | Warm the cache, monitor freshness and keep safe fallbacks. |
| Client-side evaluated values | A browser or mobile app needs UI decisions for one context. | Do not ship secrets, private segments or full rules to the device. |
| Edge evaluation | The decision must happen before origin. | The edge store becomes part of release latency. |
| Remote evaluation per request | Low-volume tools where freshness matters more than latency. | It can turn a flag outage into an app outage. |
| Build-time config | The value is environment setup. | Changing it requires redeploying, so it cannot be a kill switch. |
- Most interview designs use server-side local evaluation for backends and evaluated values for clients.
How interviewers grade this
- You state that flag evaluation in the application hot path is local to the SDK cache.
- You separate the control plane that edits rules from the data plane that evaluates flags.
- You name the safe default for each flag and what happens when the provider is stale or unavailable.
- You explain propagation latency for streaming, polling and client-side SDKs.
- You do not send full rulesets, private segments or SDK keys to browsers.
- You include exposure events, SRM checks and stale-flag cleanup in the design.
Variants
Release flags
Hide unfinished code until the rollout is ready.
Use release flags to decouple deploy from release. They should be short lived, owned by the team shipping the change, and removed after the rollout is complete. The cleanup task is part of the release plan, not a future refactor.
Pros
- +Reduce deploy risk.
- +Support gradual rollout and rollback without a new build.
- +Let multiple teams merge behind disabled paths.
Cons
- −Create dead branches if nobody removes them.
- −Can hide integration bugs if the disabled path is not exercised.
Choose this variant when
- A new feature is ready to merge but not ready for every user.
Experiment flags
Assign users to variants so product impact can be measured.
Experiment flags need stable bucketing, exposure logging and analysis guardrails. Assignment decides what a user would receive. Exposure records that the user actually saw the variant. Keep assignment and exposure separate so the denominator matches the question you are answering.
Pros
- +Measure impact before committing.
- +Support ramping while preserving earlier buckets.
- +Surface SRM and event-quality problems.
Cons
- −Need careful analytics plumbing.
- −Overlapping experiments can confound results on the same surface.
Choose this variant when
- You need evidence that a product change helps users or the business.
Operational flags and kill switches
Disable risky behaviour during an incident.
Operational flags protect systems. They turn off a dependency call, degrade an expensive feature, cap concurrency or force a safer algorithm. They need clear ownership, monitoring, rehearsal and a propagation path that is faster than the incident they are meant to stop.
Pros
- +Reduce blast radius without redeploying.
- +Let on-call staff choose a safer mode.
- +Pair naturally with failure-mode analysis.
Cons
- −Unsafe defaults can make the incident worse.
- −Untested switches may not work when needed.
Choose this variant when
- A code path can hurt availability, data quality or cost under stress.
Permission and entitlement flags
Control access by account, plan or contract.
Entitlement flags often live close to auth and billing. They can use the same SDK shape as release flags, but they are long lived business policy. Keep them auditable, test them like access control, and avoid exposing private plan logic in clients.
Pros
- +Centralise access decisions.
- +Support sales, beta and contract-specific access.
- +Audit changes that affect customer rights.
Cons
- −A bad fallback can leak or block access.
- −They need ownership beyond the feature team.
Choose this variant when
- The value represents an account right rather than a temporary release decision.
Worked example
Numbers in this section are illustrative.
Scenario: a checkout team is shipping a new fraud-scoring call. The interviewer asks how you roll it out safely and run an experiment.
Flag shape. Create checkout.fraud_score_v2 with off, shadow and enforce. The coded fallback is off because blocking checkout during a flag outage is worse than skipping the new check. The checkout service evaluates locally from the cached ruleset.
Rollout plan. Start with internal accounts, then a small customer segment, then an illustrative ramp from 5% to 10% to 25%. Stable bucketing keeps the first 5% in treatment when allocation increases. If fraud risk is account-scoped, bucket by account.
Rule distribution. Backend SDKs stream rule changes. Each process reports ruleset version and age. If a process is stale for an illustrative few minutes, remove it from service before relying on that flag for enforcement. Browser code receives display decisions, not the full rule graph.
Kill switch. The same flag has an emergency target: set all contexts to off. The runbook says who can change it, what alert triggers it, and how to confirm propagation. If only polling is available, the polling interval is the worst-case control-loop delay.
Experiment measurement. Log exposure when checkout renders the treatment. Assignment in the SDK is not exposure. Conversion events use the same account key. Check planned versus observed split before outcome metrics, because SRM can mean the denominator is broken.
Cleanup. After rollout, replace the branch with the winning path, remove old segments and archive the flag. If the team wants a permanent operations switch, create a smaller kill switch with its own owner and fallback.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“How do you design feature flags so the flag service outage does not take down your checkout service?”
Weak answer
"The checkout service calls the flag service to see if the feature is on. If the service is down, we will retry or use the dashboard value."
Strong answer
"The checkout process uses a server-side SDK that keeps the latest good ruleset locally. A check is a local function call with a coded fallback, for example off for risky fraud enforcement. Rule changes arrive out of band through streaming, or polling with versions. Each process exposes ruleset version and freshness. Browsers get evaluated values for the current user, not the full ruleset or secrets. The flag service can be down and checkout still runs with the last good ruleset or fallback."
Why it wins: It separates hot-path evaluation from rule distribution, names the outage behaviour, protects client-side secrets and gives operations a way to detect stale SDKs.
When it comes up
- A prompt mentions gradual rollout, A/B testing, beta access, kill switches or safe deployment.
- You introduce a risky dependency that should be easy to disable.
- The interviewer asks what happens when the config or flag service is down.
- A browser or mobile client wants to use feature flags.
Order of reveal
- 1Start with local evaluation. “The app does not call the flag service for every check. The SDK evaluates locally from a cached ruleset.”
- 2Name the failure behaviour. “If the provider is unavailable, I use the last good ruleset, and if evaluation still fails I return the coded fallback.”
- 3Explain distribution. “Rules change out of band through streaming where I need fast propagation, or polling with versions where simplicity matters.”
- 4Handle clients carefully. “Server SDKs can hold rules. Browsers get evaluated values for the current user only.”
- 5Add experiment hygiene. “Stable bucketing decides assignment, exposure is logged when the user sees the variant, and SRM is a guardrail before reading results.”
- 6Close with lifecycle. “Every temporary flag has an owner, audit trail, expiry and cleanup path.”
Signature phrases
- ““Flags are a local branch with a remote control plane.”” — Separates data-plane evaluation from rule management.
- ““Last good ruleset, then coded fallback.”” — Gives a concise outage story.
- ““Assignment is not exposure.”” — Protects experiment denominators.
- ““Client SDKs get values, not secrets.”” — Shows the security boundary.
Likely follow-ups
?“What if the flag service is down for an hour?”Reveal
Existing server processes continue evaluating from their last good ruleset. New processes either bootstrap from a persistent cache or return coded fallbacks until ready. The system should alert on stale ruleset age, and the incident runbook should say whether to freeze deploys that depend on new flag values.
?“How do you roll from an illustrative 5% to 10% without reshuffling everyone?”Reveal
Hash a stable salt plus the randomization key into a fixed bucket range. The illustrative 5% treatment owns the first slice of buckets. When you increase to an illustrative 10%, you add the next slice, so the first cohort stays in treatment.
?“Can a browser evaluate targeting rules itself?”Reveal
Only if the rules are safe to expose. For normal product flags, rules can contain private attributes or segments, so the server or flag provider evaluates and the client stores evaluated values for that context.
?“How do approvals fit with a kill switch?”Reveal
Routine production targeting changes may require approvals. Emergency kill switches need a pre-approved path or a smaller set of on-call owners, otherwise the approval workflow can be slower than the incident.
Code examples
type FlagClient = {
boolVariation(flag: string, context: Record<string, unknown>, fallback: boolean): boolean;
};
function canUseFraudV2(flags: FlagClient, accountId: string) {
return flags.boolVariation(
'checkout.fraud_score_v2',
{ targetingKey: accountId, kind: 'account' },
false,
);
}
function renderCheckout(flags: FlagClient, accountId: string) {
const enabled = canUseFraudV2(flags, accountId);
if (enabled) {
logExposureOnce(accountId, 'checkout.fraud_score_v2', 'fraud-v2');
}
return enabled ? renderFraudV2Checkout() : renderDefaultCheckout();
}import { createHash } from 'node:crypto';
const BUCKETS = 10_000;
export function bucket(salt: string, userId: string): number {
const input = `${salt}:${userId}`;
const digest = createHash('sha256').update(input, 'utf8').digest();
const n = digest.readUInt32BE(0);
return n % BUCKETS;
}Common mistakes
This makes the flag service part of application availability. Use local SDK evaluation for hot paths, then update the SDK cache in the background. Remote evaluation belongs in low-volume tools, not every checkout request.
Targeting rules can reveal private segments, plan names or rollout strategy. Browsers and mobile apps should receive evaluated values for one context, not server-side flag configuration.
Your fallback is the value passed in code. If the SDK cannot evaluate, that value protects the app. Review fallbacks like failure-mode decisions: risky writes fail closed, optional UI can fail hidden, and availability controls fail open only when safer.
A switch that waits for long polling, cache expiry and manual approval may be too slow. State the expected propagation path, measure it, and rehearse it.
A user can be assigned and still miss the surface. Log exposure when the variant is shown, then deduplicate by user, flag and variant. Otherwise the denominator includes users who could not be affected.
Old flags create dead branches, hidden policy and interactions with new features. Give release and experiment flags an owner, expiry and cleanup ticket. Long-lived operational and entitlement flags need ownership and tests.
Changing the bucketing salt can reshuffle users between variants and contaminate metrics. Treat salts and bucket algorithms as versioned experiment configuration.
Practice drills
Numbers in this section are illustrative.
Why is a remote flag call on every request risky?Reveal
It puts the flag service in the request availability and latency path. A flag-service outage, slow network or rate limit can now hurt the application path you were trying to protect. Use local SDK evaluation for hot checks and refresh rules in the background.
Your SDK starts before it has downloaded flags. What should happen?Reveal
It returns coded fallbacks until a ruleset is available. If it has a persistent cache from a previous run, it may use that last good ruleset while marking freshness. Either way, evaluation must not crash the business request.
Where do you log experiment exposure?Reveal
At the point the user can actually see or experience the variant. Log once per user, flag and variant for the analysis window. Do not count plain assignment as exposure.
A 50/50 experiment observes 60/40 exposure counts. What do you do?Reveal
Treat it as an SRM until proven otherwise. Run a chi-square check, inspect randomization keys, redirects, bot filters and event instrumentation, and do not trust the outcome metrics until the root cause is understood.
Deep dives
Consistent bucketing
Consistent bucketing is how a rollout keeps a user in the same variant across services, deploys and SDK languages. Pick a stable randomization key, usually user id for user-facing UI or account id for account-wide behaviour. Combine it with a flag or experiment salt, encode the input exactly, hash it with a named algorithm, then map the digest into a fixed bucket range.
For example, define the input as UTF-8 bytes for experimentSalt + ":" + userId, hash with SHA-256, read the first unsigned 32-bit big-endian integer, and map it into buckets 0-9,999. In an illustrative ramp, buckets 0-499 receive a 5% treatment. When you ramp to an illustrative 10%, you add buckets 500-999, so the original 5% stays in treatment. This is why a fixed bucket space is better than re-randomising at each ramp.
Do not use a language's built-in hash for bucketing. Built-in hashes can differ by language, runtime version, process seed or object representation. Specify the hash function, input string, character encoding, byte order, bucket range, salt and modulo or division rule. Then publish golden test vectors so every SDK returns the same bucket for the same input.
Changing the salt reshuffles users because the hash input changes. That can be useful when a new independent experiment must not inherit old assignments, but it is dangerous during a running experiment because users move between variants. Treat salt changes as starting a new experiment iteration.
LaunchDarkly documents one product's rollout approach. Its UI docs describe percentage rollouts as assigning a context to one of 100,000 buckets by context key and context kind. Its server-side evaluation rules are more specific: choose the rollout context kind and bucket attribute, default the attribute to the context key, then compute SHA1 over flagKey.salt.attributeValue or seed.attributeValue, take the first 15 hex characters, convert to base 10, and divide by 0xFFFFFFFFFFFFFFF (Flag evaluation rules).
Overlapping experiments need an explicit policy. If two experiments affect unrelated surfaces, a user can be in both and you analyse them separately. If they affect the same decision or metric, use mutual exclusion, layers, or a shared namespace so a user enters at most one of the conflicting tests. If you ignore overlap, one treatment can change the other treatment's denominator or effect size.
Exposure events
Exposure is the moment a context actually experiences a variant. Assignment is only the result of bucketing or targeting. A user can be assigned to treatment, leave before the component renders, be redirected away, have the feature hidden by a downstream condition, or use an API path that never touches the changed surface. Counting those users as exposed dilutes the denominator and can bias the result.
Log exposure at the boundary where the variant can affect behaviour. For UI, that is usually render of the component, not page load if the component might never appear. For an API, it is the code path that uses the variant, not a configuration lookup at request start. Deduplicate per user, flag and variant for the analysis window so retries, React re-renders or repeated polling do not multiply exposure counts.
Keep exposure and conversion events on the same randomization key. If assignment is by account, conversion should be attributed by account or by a user-to-account join with clear rules. Mixing user-level exposure with account-level outcomes can create correlation errors that look like product impact.
OpenFeature hooks can run before, after, on error and finally around flag evaluation (OpenFeature hooks). Hooks are useful for telemetry, but do not blindly emit exposure from every successful evaluation. A service may evaluate a flag to decide which data to fetch, then return a response that the user does not see. Put exposure logging close to the user-visible or business-effective point.
LaunchDarkly's Experimentation docs describe evaluation events and metric events separately: SDKs send feature evaluation events when a flag is evaluated, and metric events when the user or system action occurs; for hosted metrics, evaluations determine exposure and audience membership, while a track call records the metric event (Experimentation and metric events). The design lesson is to make the denominator explicit before comparing outcomes.
Sample-ratio mismatch
Sample-ratio mismatch, or SRM, means the observed split does not match the planned split. If an experiment is planned as an illustrative 50/50 split and exposures arrive as 60/40, the first question is not "which variant won?" It is "why did randomization or measurement fail?"
The usual check is a chi-square goodness-of-fit test against the planned allocation. Counts are compared to expected counts for each variant. A small p-value means the observed split is unlikely under the planned split, so the experiment has a data-quality problem. Do this check on the analysis denominator, usually exposed users or accounts, not only on converters.
Typical causes are mundane and serious. One variant may redirect users before the exposure event fires. Bot filtering may remove one path more than another. A client may log exposure only for the new component. A backend may bucket by user id while the analytics pipeline deduplicates by account id. Mobile and web clients may use different salts or SDK versions. A ramp can change traffic allocation without starting a clean iteration.
An SRM can invalidate the result because the treatment and control populations are no longer comparable in the way the test planned. Fabijan et al. describe SRM as a symptom of data-quality issues in online controlled experiments and warn that ignoring it can make a bad product modification look good or the reverse (Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners, KDD 2019).
The interview answer is: pause interpretation, quantify the mismatch, trace assignment and exposure separately, segment by platform and path, inspect redirects and filters, and fix instrumentation before shipping a decision. If the root cause affected only a known period or path, you may restart the experiment or create a clean iteration. Do not cherry-pick around the mismatch unless the exclusion rule was pre-declared and defensible.
Cheat sheet
- •Evaluate flags locally in SDKs on hot paths.
- •Distribute rules out of band through streaming or polling with versions.
- •Use last good ruleset first, coded fallback second.
- •A kill switch needs measured propagation and a rehearsed owner path.
- •Target with context attributes and reusable segments.
- •Server SDKs can hold rules; clients get values for one context.
- •Stable bucketing needs specified hash, salt and encoding.
- •Exposure is logged when the user sees the variant, not when assigned.
- •Check SRM before trusting experiment results.
- •Clean up stale release and experiment flags.
Practice this skill
No problem is tagged directly to Feature flags and experiments yet. These published problems still exercise the same interview category.
Read this if