Security, authorization and multi-tenancy
Authentication vs authorization, tenant isolation models, encryption in transit, at rest and end to end, key management, least privilege and audit.
Security and privacy answers should sound like a set of boundaries you can enforce, audit and explain. This page gets you ready to choose tenant isolation, authorization, encryption, key handling and privacy controls without hand-waving.
Read this if your last attempt…
- Your last design said "use OAuth" but did not say where access is enforced
- You scoped API requests by tenant but forgot background jobs, caches or SQL queries
- You used "encrypted" as a blanket answer without naming what it protects
- You were asked about deletion, residency or audit logs and gave a compliance-only answer
The concept
Authentication and authorization are different decisions. Authentication proves who the caller is. Authorization decides whether that subject may perform a specific action on a specific resource in a specific tenant. Sessions, tokens, OAuth, OIDC, CSRF and the short RBAC / ABAC / ReBAC overview live in Auth and sessions; this page assumes identity exists and focuses on enforcement.
Authorization lives at every boundary
Treat authorization as a pipeline, not one middleware call. A policy decision point (PDP) evaluates a request such as "user 42 may export invoice 9 for tenant A". A policy enforcement point (PEP) blocks or allows the action. Put PEPs at the API handler, service method, repository helper, database policy, cache key builder and job runner.
The request is authenticated once, then policy and tenant isolation are enforced at the API boundary, service method, data query and audit path.
Tenant isolation models: choose the boundary you can operate and test.
| Model | Blast radius | Noisy neighbours | Cost and operations |
|---|---|---|---|
| Shared table with tenant_id + RLS | Largest if policy coverage, request roles or cache keys are wrong. | High unless quotas, rate limits and workload controls are tenant-aware. | Lowest cost; one schema and simple migrations, but tests must prove every path is scoped. |
| Schema per tenant | Medium: namespace mistakes can be contained, but the database engine is shared. | Medium: tenants share CPU, memory, connections and maintenance windows. | Moderate cost; migrations, grants and search_path handling become product work. |
| Database per tenant | Lower when credentials, backups and admin paths are separated. | Lower than a shared database, although the host or managed service can still be shared. | High cost; provisioning, monitoring, migrations and failover multiply by tenant count. |
| Cluster per tenant | Lowest product blast radius when network and operations are also isolated. | Lowest, because capacity can be reserved per tenant. | Highest cost; suited to regulated, very large or contractually isolated tenants. |
How interviewers grade this
- You separate authentication from authorization, then link to the session/token design rather than re-teaching it.
- You name PDP and PEP, and you put PEPs at API, service and data-query boundaries.
- You compare tenant isolation by blast radius, noisy neighbours and cost.
- You distinguish TLS, mTLS, at-rest encryption and end-to-end encryption by what each protects.
- You explain KMS, envelope encryption, rotation, per-tenant keys and crypto-shredding with caveats.
- You include least privilege, service identities, secrets, audit logs and privacy controls in the main design.
Variants
Shared table plus row-level security
One schema, every tenant-owned row has tenant_id, and the database enforces policy.
Shared rows are cheapest and need strong policy coverage. Schemas reduce accidental mixing but share a database. Separate databases or clusters cost more and isolate failures better.
This is the default for many SaaS products because it keeps migrations, analytics and operations simple. The safe version uses a verified tenant context, constrained database roles, RLS policies on tenant-owned tables and tests that try to read or write a neighbour tenant. RLS is defence in depth, not a licence to pass a superuser connection to the request path.
Pros
- +Lowest operational cost
- +Easy cross-tenant product analytics when allowed
- +One migration path for the whole fleet
Cons
- −A missing tenant filter can affect many tenants
- −Noisy neighbours share the same database resources
- −Privileged roles or raw SQL can bypass the intended guardrail
Choose this variant when
- Long tail of small tenants
- Early SaaS product with strict test coverage
- Data classes that are not contractually isolated
Schema per tenant
Each tenant has a namespace inside one database.
A schema boundary reduces accidental row mixing and lets you migrate or restore one tenant namespace more directly. It also creates thousands of copies of the same objects if the fleet grows. The interview risk is forgetting search_path, grants and migration orchestration.
Pros
- +Clearer namespace boundary than shared rows
- +Tenant-specific restore or migration can be easier
- +Still cheaper than many databases
Cons
- −Shared database resources remain shared fate
- −Many schemas make migrations slower and riskier
- −Wrong search_path or grants can cross the boundary
Choose this variant when
- Medium tenant count
- Customers need clearer namespace separation
- You can invest in migration automation
Database or cluster per tenant
Put the biggest or most regulated tenants behind separate credentials, backups and capacity.
This is a business-tier and risk-tier choice. It is attractive for regulated tenants, custom maintenance windows, data residency and strong operational isolation. The cost is a control plane: provisioning, schema changes, backups, failover, observability and support tooling all become multi-fleet problems.
Pros
- +Smaller blast radius for query bugs and admin mistakes
- +Better noisy-neighbour control
- +Easier to explain to high-assurance customers
Cons
- −Highest cost
- −Slower provisioning unless automated
- −Operational drift can create its own security risk
Choose this variant when
- Large enterprise tenants
- Strict residency or isolation contracts
- Tenants that need dedicated capacity
Per-tenant key hierarchy
Use tenant-scoped wrapping keys so one key decision affects one tenant.
The app encrypts data with a data key, stores the encrypted data key next to the ciphertext, and asks KMS to unwrap only for authorized tenant-scoped work.
A per-tenant key hierarchy gives you a smaller compromise window and a clean story for tenant offboarding. The data path still needs authorization: KMS decrypt permission must include service identity, tenant, purpose and environment. Crypto-shredding is not enough if plaintext copies live in search indexes, logs or exports.
Pros
- +Limits key compromise blast radius
- +Supports tenant offboarding and crypto-shredding
- +Gives audit events per tenant key
Cons
- −More KMS policies and quotas to manage
- −Key loss is tenant data loss
- −Derived stores must use the same lifecycle discipline
Choose this variant when
- Sensitive tenant data
- Enterprise offboarding requirements
- Blob or document stores where envelope encryption is natural
Worked example
Numbers in this section are illustrative.
Scenario: design security for a B2B document workspace. Numbers in this section are illustrative: assume 20,000 tenants, a long tail of small teams, and 25 regulated tenants that pay for stronger isolation.
Identity. Users authenticate through the existing access plus refresh design from Auth and sessions. The token gives subject, issuer and selected tenant. It is not the final permission decision.
Request path. The API boundary validates the token, derives tenant context from server-verified membership and calls a policy decision point with subject, action, resource and tenant. The service method enforces the result before it reads data. The repository receives a required tenant context and refuses unscoped access. Every SQL query includes tenant scope; shared tables also have PostgreSQL RLS as a second line of defence.
Isolation. Small tenants share tables with tenant_id, RLS, tenant-aware cache keys and per-tenant rate limits. The 25 regulated tenants use a database-per-tenant tier with separate credentials, backups and maintenance windows. A directory maps tenant to database so a tenant can move from shared to dedicated after a copy, validation and cutover.
Encryption and keys. Browser to edge uses TLS. Service-to-service calls use mTLS where the network boundary is shared. Documents are encrypted with envelope encryption: a data key encrypts each document object, and the encrypted data key is stored next to the ciphertext. The wrapping key is tenant-scoped in KMS. Offboarding a tenant means disable access, export if contracted, delete active rows, schedule backup expiry and destroy the tenant key only after the retention policy allows it.
Least privilege and secrets. The document service can decrypt document keys for its tenant-scoped path, but the notification worker cannot. The support tool needs a user, tenant, ticket id and reason code before it can open a read-only session. Database passwords, webhook signing keys and third-party API keys live in the secrets manager, with rotation runbooks and redaction in logs.
Audit and privacy. Each protected action writes actor, tenant, resource, action, decision, reason, source service, correlation id and outcome to an append-only audit stream. Product logs avoid document bodies and raw tokens. Data classes have retention rules: active documents, deleted documents, audit events, search indexes and exports each have a deletion path. Residency is enforced by tenant home region, storage bucket, KMS key region, job queue and analytics pipeline.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“How would you secure a multi-tenant SaaS data model?”
Weak answer
"We use OAuth, encrypt the database and put tenant_id on every table. Admins can debug issues if needed."
Strong answer
"OAuth tells me who the caller is; it does not prove they can touch this tenant or document. I derive tenant context from verified membership, then enforce policy at the API boundary and again at the service/data layer. For the long tail I would use shared tables with tenant_id, constrained DB roles and PostgreSQL RLS, plus tests that try cross-tenant reads and writes. Bigger regulated tenants can move to database-per-tenant using a directory. Encryption is layered: TLS in transit, at-rest encryption for storage loss, and envelope encryption with tenant-scoped KMS keys for documents. Audit logs record actor, tenant, resource, decision and reason without secrets. Privacy means minimisation, retention, deletion and residency."
Why it wins: It separates identity from policy, places enforcement in more than one layer, compares isolation tiers, names what encryption does and does not protect, and includes audit and privacy instead of treating them as add-ons.
When it comes up
- B2B SaaS, enterprise tenants or admin/support tooling
- Prompts mentioning privacy, data residency, regulated customers or tenant isolation
- Any system with sensitive documents, messages, payments or health-like data
- Follow-ups after authentication: "what can this user access?"
Order of reveal
- 11. Split identity from access. Authentication gives me a subject; authorization is subject plus action plus resource plus tenant.
- 22. Place enforcement points. I enforce at the API boundary, service method and data query, with a policy decision point behind them.
- 33. Choose tenant isolation. I compare shared table with RLS, schema per tenant and database or cluster per tenant on blast radius, noisy neighbours and cost.
- 44. Layer cryptography. TLS protects traffic, at-rest encryption protects stored media, and end-to-end encryption keeps the service from seeing plaintext.
- 55. Finish with operations and privacy. Least privilege, service identities, secrets, audit logs, minimisation, deletion, retention and residency are part of the design.
Signature phrases
- “AuthN says who; authZ says whether this subject can do this action to this resource in this tenant.” — Keeps the model concrete.
- “Every request gets a policy check; every query gets tenant scope.” — Shows enforcement at more than one boundary.
- “Isolation is a blast-radius and cost decision, not a yes/no feature.” — Frames the tenant model trade-off.
- “Encryption is not one control; name the attacker and the moment you protect.” — Prevents vague "we encrypt it" answers.
Likely follow-ups
?“What is the difference between RBAC and the enforcement design here?”Reveal
RBAC is one way to express policy. Enforcement is where that policy is checked and applied. You can have RBAC, ABAC or ReBAC decisions behind the same PEP pattern: route, service method, query and job worker all ask or enforce before touching data.
?“Why not put every tenant in its own database from day one?”Reveal
It gives a strong boundary, but multiplies provisioning, migrations, monitoring, backups and failover. For a long tail of small tenants, shared tables with RLS, tests and quotas may be the right cost point. Keep a directory design so high-risk tenants can move to dedicated databases later.
?“What does mTLS buy you if you already have JWTs?”Reveal
mTLS authenticates the service connection and protects the transport. JWTs carry caller claims. They answer different questions: is this connection from the expected service, and is this user or service allowed to perform this action? In a service mesh, use both rather than treating either as a full authorization system.
?“How do you prove deletion?”Reveal
Track data classes and stores, not just tables. Delete active rows, derived indexes, blobs and exports; let backups age out under documented retention; then record the deletion workflow in audit logs. Crypto-shredding can make protected ciphertext unrecoverable when key ownership and copies are controlled.
Code examples
ALTER TABLE documents ENABLE ROW LEVEL SECURITY;
ALTER TABLE documents FORCE ROW LEVEL SECURITY;
CREATE POLICY documents_tenant_isolation
ON documents
USING (tenant_id = current_setting('app.current_tenant')::uuid)
WITH CHECK (tenant_id = current_setting('app.current_tenant')::uuid);
BEGIN;
SELECT set_config('app.current_tenant', :tenant_id, true);
SELECT id, title
FROM documents
WHERE id = :document_id;
COMMIT;type DecisionInput = {
subjectId: string;
tenantId: string;
action: 'read' | 'write' | 'export';
resource: { type: 'document'; id: string };
};
async function loadDocument(input: DecisionInput) {
const decision = await policy.decide(input);
if (!decision.allow) throw new ForbiddenError(decision.reason);
return documents.findOne({
tenantId: input.tenantId,
documentId: input.resource.id,
});
}Common mistakes
A route check can be bypassed by a new endpoint, a background job, a batch export or a direct repository call. Put enforcement at the API boundary and the data boundary, then make unscoped data access hard to call.
A header or URL segment can select a tenant, but it should not prove access. Bind tenant context to authenticated membership or service authorization, and pass the verified context through the request.
RLS is not useful on ordinary request paths if the app connects as a superuser or a role that bypasses row security. Use constrained roles for requests, reserve privileged roles for migrations and test the deployed role.
TLS, at-rest encryption and end-to-end encryption protect different moments. Name the attacker: network observer, stolen disk, compromised app server, curious operator or subpoenaed service. Then choose the layer that addresses that attacker.
If the app can rewrite its own audit trail, the log is weak evidence. Write through a separate identity to an append-only store, include policy decisions and tenant context, and avoid logging secrets or unnecessary personal data.
Destroying a tenant key only helps ciphertext protected by that key. Search indexes, exports, thumbnails, logs, analytics tables and backups need their own deletion or retention path.
Practice drills
Numbers in this section are illustrative.
A user can change /tenants/a/documents/1 to /tenants/b/documents/1 and read data. What failed?Reveal
The system treated a client-supplied tenant id as proof of access. Derive tenant context from authenticated membership, check authorization for the specific document and tenant, and scope the query so document 1 is looked up only inside the verified tenant.
Why does at-rest encryption not solve a compromised app server?Reveal
The app server normally has permission to ask the database or KMS for plaintext during legitimate requests. If that server is compromised, at-rest encryption does not stop authorized decrypt calls. You still need least privilege, tenant-scoped keys, monitoring, audit logs and query-level authorization.
When would you move a tenant from shared tables to database per tenant?Reveal
Move when the tenant needs lower blast radius, dedicated capacity, custom residency, stricter backup controls or a contractually clearer boundary. The move needs a directory, data copy, validation, cutover plan, key policy change and rollback path.
What belongs in an audit log for a denied export?Reveal
Record actor, tenant, action, target resource, policy decision, denial reason, source service, request or correlation id, time and outcome. Do not log full tokens, secrets or exported content.
Deep dives
Tamper-evident audit logs
An audit log is useful only if it survives the incident you are investigating. The minimum design is append-only from the application point of view: application services can write audit events but cannot edit or delete old ones. A stronger design writes to a separate account or project, uses a stream or object store with retention locks, and periodically seals batches with a hash chain or signature. Hash chaining means each event or batch includes the hash of the previous one; changing an old record breaks the chain from that point forward. It does not stop deletion by a privileged operator, so pair it with independent storage permissions and monitoring.
Keep the schema boring and complete. Include actor, tenant, subject type, action, resource id, decision, policy version or reason, source service, client context when safe, correlation id and outcome. Store enough to reconstruct who did what without storing the sensitive payload itself. If support staff can impersonate or view customer data, make support-session start, scope, reason and end events first-class audit entries. Audit the audit system too: changes to log retention, export, access and alert settings deserve their own events.
Privacy at interview depth
Interview privacy answers should be concrete enough for engineering. Start with data minimisation: for each field, say why you collect it, where it is stored and which feature uses it. Then define retention: active data, soft-deleted data, backups, logs, analytics and exports may all have different lifetimes. Finally, define deletion: user-request deletion, tenant offboarding, legal hold and fraud hold are separate workflows. The mistake is to say "delete the row" when search indexes, object storage, data warehouse tables and logs still carry copies.
Residency is about all locations, not just the primary database. A tenant home region must cover request routing, database, object storage, KMS key, logs, queues, caches, analytics and support access. Cross-region replication can still be allowed if the contract and policy allow it. For active-active products, the cleanest privacy story is often tenant or user home-region partitioning: writes for that tenant go to the home region, and other regions read replicas or derived views under explicit rules.
Server-side request forgery (SSRF)
SSRF appears whenever your server fetches a URL chosen by someone else. Webhook delivery, link unfurling, image import, PDF generation and "fetch this URL" integrations are classic cases. The attacker is not trying to read the public URL they gave you. They are using your server as a network position. OWASP describes SSRF as an application interacting with internal or external networks on the attacker's behalf, and its examples include user-controlled webhooks and avatar URLs (OWASP SSRF Prevention Cheat Sheet).
The dangerous target is often something the attacker cannot reach directly: localhost admin ports, databases with HTTP APIs, private service names, or a cloud metadata endpoint. OWASP's 2021 Top 10 SSRF page says metadata services such as 169.254.169.254 are a common attack scenario (OWASP A10:2021 SSRF). In 2025, OWASP no longer lists SSRF as its own Top 10 category; A01 Broken Access Control names CWE-918 Server-Side Request Forgery as a notable mapped weakness (OWASP Top 10 A01:2025).
Deny-lists are brittle. OWASP warns that URL parsing is hard, redirects can bypass validation, DNS pinning or rebinding can turn a public host into an internal address, and IPv4 plus IPv6 ranges must be handled together (OWASP SSRF Prevention Cheat Sheet). Alternate IP encodings also break naive string checks. Do not check for "localhost" and call it done.
The safer pattern is layered. Accept only http and https. Resolve the host, then validate every resolved address against private, loopback, link-local, multicast and reserved ranges. Connect to the validated address rather than letting a later resolver choose a different one, while preserving the intended Host header and TLS verification. Disable redirects, or re-run the same validation on every hop. Route this traffic through a dedicated egress proxy or subnet that has no route to internal networks. Set short timeouts and response size limits so fetches cannot tie up workers.
On AWS, use IMDSv2 and require tokens where possible. AWS says IMDSv2 is session-oriented: the instance first sends a PUT to /latest/api/token with a TTL header, then includes the returned token on metadata GETs; token-required mode makes IMDSv1 requests fail. AWS also exposes a metadata response hop limit: the option accepts 1-64 hops, and the IMDSv2 PUT response defaults to an IP hop limit of 1 (EC2 IMDSv2 docs, IMDS options). That is defence in depth, not permission to skip URL and network controls.
Signing webhooks
A webhook receiver has to answer two questions before it trusts a request: did the provider send this exact body, and is this delivery fresh? IP allow-lists can help operations, but they are not enough for authenticity. The common answer is a signature over the raw body plus delivery metadata, verified with a secret that belongs to that endpoint.
The Standard Webhooks spec sends webhook-id, webhook-timestamp and webhook-signature headers. It signs id.timestamp.body (the spec names this msg_id.timestamp.payload), then serialises a base64 HMAC-SHA256 as v1,<signature> for symmetric signatures (Standard Webhooks spec). It also recommends unique signing keys per endpoint. Include the timestamp in the signed material so an attacker cannot copy a valid body and replay it later with a new time.
Verify the signature against the raw bytes you received. If middleware parses JSON and serialises it again, spaces, key order or encoding can change and a correct signature will fail. Stripe's Stripe-Signature header includes a t= timestamp and one or more v1= signatures; Stripe signs timestamp.payload, computes HMAC-SHA256, and its libraries use a default 5 minute tolerance (Stripe webhook handler docs, Stripe webhook signature docs). Its constructEvent call needs the request body string sent by Stripe, the Stripe-Signature header and the endpoint secret.
Use a constant-time comparison for MAC results that you compute yourself. The Standard Webhooks spec calls that out because ordinary string comparison can leak timing information, and it tells receivers to verify the timestamp within an allowable tolerance. Then deduplicate by event id. Tolerance limits replay time; deduplication stops the same event being processed twice because of retries, races or a malicious resend. Store the event id long enough for your retry window.
Plan rotation before the first incident. Standard Webhooks supports zero-downtime rotation by sending signatures made with the current key and an old key during an overlap, so receivers can accept either until the old secret is removed. Keep the overlap short and visible in audit logs.
If receivers should not hold a signing secret, use asymmetric signatures. Standard Webhooks defines Ed25519 as its asymmetric option, where the producer keeps the private key and consumers verify with public keys. Mutual TLS is another option when both sides can issue, rotate and validate client certificates; TLS 1.3 includes certificate-based client authentication as part of the handshake (RFC 8446).
Running untrusted code
Coding platforms, plugin systems and user scripts are hostile-input systems with a compiler attached. Treat the submitted code as able to read files, fork processes, open sockets and attack the kernel unless a boundary stops it. Language sandboxes can catch accidents, but they are a weak primary boundary because runtime escape bugs and native extensions change the threat model.
Containers are useful packaging, not a complete answer by themselves. A normal container still shares the host kernel, so a kernel bug can become a host escape. gVisor reduces that exposure by putting a Sentry between the application and the host; its security guide says application interactions with the host System API are intercepted and that no system call is passed directly to the host (gVisor security guide). Firecracker takes a different trade-off: it creates lightweight microVMs using Linux KVM, so each run gets a virtual machine barrier with low startup overhead (Firecracker docs). Its design document says Firecracker combines KVM isolation with seccomp filters, cgroups, namespaces and a jailer for defence in depth (Firecracker design).
For an interview answer, build layers. Start each run in a fresh sandbox and destroy it after the result is collected. Do not reuse sandboxes across users after a run, because leftover files, processes or warmed state can become a cross-user channel. Drop Linux capabilities, run as a non-root user, mount a read-only root filesystem and give the run a small writable scratch area that is deleted with the sandbox.
Use seccomp to reduce kernel attack surface, but describe it accurately. Linux has strict mode, which permits only read, write, _exit and sigreturn, and filter mode, where BPF filters incoming system calls (seccomp(2), Linux seccomp filter docs). The kernel docs say system call filtering is not a sandbox by itself, and the man page recommends an allow-list approach because deny-lists miss new calls and alternate representations. Pair filters with cgroup limits for CPU, memory, pids and time, plus output size caps.
Network should be off by default. If a task needs package downloads or API calls, send egress through an allow-listed proxy and log the domains. Keep platform secrets, signing keys and hidden tests outside the sandbox; pass only per-run tokens with narrow scope. The clean answer is boring: one sandbox per run, few syscalls, few files, few routes, clear limits and no secrets inside.
End-to-end encryption in messaging
End-to-end encryption means the server relays ciphertext it cannot read. The plaintext exists at the communicating endpoints, usually user devices, and the service stores or forwards encrypted messages. That protects content from a curious operator or a server breach, but it shifts product work to clients.
At interview depth, name the key agreement and ratchet rather than saying "public key crypto". Signal's X3DH establishes a shared secret for asynchronous messaging and provides forward secrecy and cryptographic deniability (Signal X3DH). Signal's PQXDH is the post-quantum variant: it adds a post-quantum key encapsulation mechanism while, in this revision, still relying on elliptic-curve identity authentication (Signal PQXDH). After that initial agreement, the Double Ratchet derives new message keys for each message; later keys cannot calculate earlier keys, and new Diffie-Hellman outputs help future keys recover after compromise (Signal Double Ratchet).
Multi-device messaging is fan-out, not one shared inbox key. If Alice sends to Bob, her device may need to encrypt a copy for each of Bob's devices and for Alice's other devices so her own device list stays in sync. Signal's Sesame spec describes this asynchronous multi-device setting and says sending to a user can require sessions to all recipient devices and the sender's other devices (Signal Sesame).
E2E does not hide all metadata. The server may still see accounts, device ids, delivery timing, message sizes, IP addresses, group membership changes and abuse signals unless the design adds separate metadata protection. Backups need a clear answer too. A server-side plaintext backup breaks the E2E promise. A client-encrypted backup keeps content private only if recovery keys, device transfer and key rotation are designed carefully; Sesame also notes that restoring from backup can roll a device back to older session state (Signal Sesame).
Server-side features get harder. Full-text search, moderation queues, spam scanning, previews, analytics and support tooling cannot read message bodies on the server unless you add client-side indexes, user reports, encrypted search schemes or explicit plaintext escape hatches. Say that trade-off plainly. E2E protects content, but it is not free privacy for every part of the product.
Key management in depth
Key management is the control plane for encryption. A KMS is usually an API service with policy, identity integration, audit and managed key material. AWS KMS says plaintext KMS keys stay inside its FIPS 140-3 Security Level 3 validated HSMs, and that applications call KMS for cryptographic operations (AWS KMS cryptography essentials). An HSM is the hardware boundary itself. AWS CloudHSM describes an HSM as a device that processes cryptographic operations and provides secure storage for keys, with customers controlling keys, algorithms and application integration (AWS CloudHSM introduction). In interviews, KMS is the managed control plane; HSM is the specialised hardware you might operate or dedicate when compliance or control requires it.
Envelope encryption keeps large data away from KMS. AWS defines it as encrypting plaintext data with a data key, then encrypting that data key under another key (AWS KMS envelope encryption). Store the encrypted data key next to the ciphertext. The app asks KMS to generate or decrypt the data key, uses the plaintext data key in memory, then clears it. If KMS latency or quotas become a bottleneck, data keys can be cached briefly, but the AWS Encryption SDK says caching should be used cautiously and with security thresholds because excessive data-key reuse is discouraged (AWS Encryption SDK data key caching).
Rotation needs precise language. AWS KMS automatic rotation for customer managed symmetric keys defaults to 365 days when enabled, and current docs allow a custom period from 90 to 2560 days (AWS KMS EnableKeyRotation API). AWS KMS key rotation changes the current key material used for new encryption, keeps older key material for decrypting existing ciphertext and does not re-encrypt data or rotate data keys (AWS KMS key rotation). On-demand rotation can start immediately without changing automatic schedules, but AWS caps it at 25 times per KMS key (AWS KMS RotateKeyOnDemand API). If you need old objects under new data keys, you need a re-encryption job and a migration plan.
Per-tenant keys reduce blast radius when access policy, encryption context and service paths are tenant-scoped. They can support crypto-shredding, but only for ciphertext controlled by that key hierarchy. Backups, caches, search indexes, exports, logs and derived analytics may survive unless their lifecycle is tied to the same tenant deletion workflow.
Access policy and audit decide whether the cryptography matters. AWS KMS says no principal has permissions to a KMS key unless access is explicitly provided, and it uses key policies, IAM policies and grants to control access (AWS KMS access control). AWS also logs KMS API calls in CloudTrail, including key management and cryptographic operations such as GenerateDataKey and Decrypt (AWS KMS CloudTrail logging). NIST SP 800-57 Part 1 gives the broader lifecycle frame for cryptographic keying material (NIST SP 800-57 Part 1 Rev. 5).
Prompt injection is a security problem
Prompt injection is the security version of a boundary mistake. In most LLM apps, instructions and data share one natural-language channel; the OWASP LLM Prompt Injection Prevention Cheat Sheet says this lack of clear separation is the common design weakness attackers exploit (OWASP LLM Prompt Injection Prevention Cheat Sheet). Treat every retrieved document, email, web page, issue, ticket, vector-search result and tool result as untrusted input, not as a new policy source.
Direct and indirect injection
OWASP names LLM01:2025 Prompt Injection as the prompt-injection entry in its LLM Top 10, and defines direct injection as user input that changes model behaviour while indirect injection comes through external content such as websites or files. Greshake et al. show why indirect injection matters: an attacker can place prompts into data likely to be retrieved, then influence tool use, API calls or data theft when an LLM-integrated app reads it (Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection).
Treat output as untrusted
The model output is also untrusted. Do not render raw output as HTML, and be careful with markdown renderers that auto-load images or make links look safe. OWASP's prompt-injection page includes an indirect-injection scenario where a model inserts an image URL that can exfiltrate private conversation data (LLM01:2025 Prompt Injection). Do not pass model output directly to a shell, SQL query, browser automation step or internal admin API. Parse it as data, validate it against a schema and make code enforce the decision.
Contain the agent
Least privilege is the main defence after the model is fooled. OWASP LLM06:2025 Excessive Agency says damaging actions can follow manipulated outputs, and ties the root causes to excessive functionality, permissions and autonomy. Give tools scoped, per-user credentials. Prefer read-only access by default. Keep a tool and argument allow-list outside the prompt. Require human confirmation for side effects such as payments, sends and deletes. Enforce per-tenant data scope in the retrieval layer and downstream systems, not by asking the model to remember the tenant.
Isolation still matters. Run tool code and plugin code in a sandbox like the one in Running untrusted code. If a tool fetches URLs, apply the egress controls from Server-side request forgery (SSRF): restricted routes, validated destinations, redirects handled carefully and short limits. A model that can choose "fetch this URL" has an SSRF-shaped edge unless the network path is constrained.
Be honest about limits
No filter reliably stops prompt injection. OWASP says it is unclear whether fool-proof prevention exists for prompt injection, so filters, prompt rules and model guardrails should reduce risk rather than be the boundary (LLM01:2025 Prompt Injection). Use AI evaluation for guardrail placement, adversarial cases and latency trade-offs. For security, design for containment: least privilege, strict schemas, sandboxing, egress limits, monitoring, rate limits and audit logs of every tool call and downstream action. OWASP LLM06 recommends logging and monitoring extension and downstream activity to spot undesirable actions (LLM06:2025 Excessive Agency).
Cheat sheet
- •Authentication proves identity; authorization decides action on resource in tenant.
- •Put PEPs at route, service, query and job boundaries.
- •Shared table + RLS is cheap; database or cluster per tenant lowers blast radius at higher cost.
- •TLS protects traffic; at-rest encryption protects stored media; E2EE hides plaintext from the service.
- •Envelope encryption = data key for data, wrapping key for data key.
- •Per-tenant keys reduce key blast radius and can support crypto-shredding.
- •Least privilege applies to humans, services, database roles and KMS grants.
- •Secrets live in a manager, rotate, and stay out of source, logs and tickets.
- •Audit actor, tenant, action, resource, decision, reason and outcome in append-only logs.
- •Privacy basics: minimise, retain deliberately, delete across stores, route by residency.
Practice this skill
No problem is tagged directly to Security, authorization and multi-tenancy yet. These published problems still exercise the same interview category.
Read this if