S3 / Blob Storage
Durable object storage for large unstructured bytes, with databases holding metadata and CDNs handling delivery.
Also worth naming: Amazon S3 · Google Cloud Storage · Azure Blob Storage · Cloudflare R2 · MinIO (self-hosted, S3-compatible)
S3 gets you ready to keep file bytes out of your app and database. You’ll design the upload path, the metadata state machine and the CDN delivery path separately.
What it is
Object storage is a service for storing large, unstructured blobs — images, videos, documents, backups, log archives, ML datasets — addressed by a key within a bucket. You PUT an object and get back a URL; you GET it by key. There are no rows, no schema, no queries — just durable, cheap, effectively unlimited file storage.
Its defining properties make it boring in the best way: it scales automatically (no capacity to provision; request rates scale per prefix, see below), extremely durable (S3 advertises eleven nines — 99.999999999% — via replication and erasure coding across devices and facilities), and cheap (cents per GB-month, far less than storing the same bytes in a database). In an interview you can treat its scalability and durability as given and spend your time on how data flows in and out.
The canonical pattern is store the blob in S3, store a pointer (URL/key) in your database. The database indexes and queries the metadata with low latency; S3 holds the heavy bytes cheaply. Uploads go direct from the client via pre-signed URLs so your app never proxies the bytes, and downloads are served through a CDN with the bucket as origin. Reach for object storage whenever a system stores files bigger than a few KB — and never use it as a primary database or for low-latency random access to structured data.
When to reach for it
Reach for this when…
- You store large files — images, video, audio, PDFs, user uploads, ML data
- You need durable, cheap, effectively unlimited capacity for unstructured bytes
- Static assets or media to serve globally (paired with a CDN)
- Backups, log/event archives, data-lake storage, or a staging area for batch processing
Not really this pattern when…
- The data is structured and you query it by field (that is a database)
- You need low-latency random reads/writes or transactions (database / cache)
- You need to modify parts of a file in place — objects are written and replaced whole
- Tiny values where a database row or cache entry is simpler and faster
How it works
A few facts cover almost every interview use:
1. Objects, not files or rows. Each object is a blob plus metadata under a key in a flat namespace (the "folders" in a key like users/42/avatar.png are just a naming convention). Objects are written and read whole — there is no in-place edit and no partial append; to "change" an object you upload a new version. This is why it is great for media and terrible as a mutable database.
2. Durability and scale are givens; design around access, not capacity. Providers replicate and erasure-code across facilities for eleven-nines durability and present effectively unlimited capacity. So in a design you don't size S3 — you focus on the upload path, the download path, and lifecycle/cost.
3. Direct-to-blob uploads via pre-signed URLs. The client requests a short-lived signed URL scoped to one object; it then PUTs the bytes straight to S3, bypassing your app entirely. The app validates and records metadata but never proxies the file — which is what makes uploads scale.
The client asks the app for permission, gets a short-lived signed URL scoped to one object, and uploads straight to S3. The app stays out of the data path, so its bandwidth and memory are spent on metadata, not file bytes.
4. Serve through a CDN, with the bucket as origin. For downloads, put a CDN in front so objects are cached at the edge near users; the origin bucket only sees cache misses. Private content uses signed URLs / signed cookies with short expiry so only authorised users can fetch.
The bucket is the durable origin of record; a CDN caches objects at edge locations near users. Downloads are served from the edge with signed URLs for private content, so the origin handles only cache misses.
5. Large uploads use multipart. Files over ~100 MB are uploaded in parallel chunks (S3 multipart / the tus protocol) so a network blip retries one part instead of the whole file — and a lifecycle rule aborts abandoned multipart uploads so they don't accrue cost.
6. Writes can be guarded. For create-only keys, complete the PUT or multipart upload with If-None-Match: *; for compare-and-swap overwrites, use If-Match with the ETag you previously read. Those checks keep concurrent clients from silently overwriting each other, but your metadata row still needs an explicit state machine.
Performance envelope
Object storage characteristics — order-of-magnitude anchors; hardware, payload size, access pattern, Region, storage class, and configuration change the numbers.
| Dimension | Number | Why it matters |
|---|---|---|
| Durability | ~11 nines (99.999999999%) | Replication + erasure coding; treat as effectively never-lost |
| Scalability | Effectively unlimited | Never a capacity concern in a design — treat as given |
| Object size | Single PUT up to 5 GB; multipart up to 50 TB | Parts are 5 MiB–5 GiB, max 10,000 parts; chunk the big ones |
| Throughput | Baseline at least 3,500 PUT/COPY/POST/DELETE + 5,500 GET/HEAD per partitioned prefix | S3 scales prefixes automatically; use natural keys first, explicit partitions only for known extreme hot spots |
| Latency | General S3 roughly 100–200 ms first-byte; S3 Express One Zone directory buckets single-digit ms | Express supports very high request rates inside one AZ; lower latency/pricing trades off the multi-AZ availability/blast-radius profile of general-purpose buckets |
| Cost | ~$0.02/GB-month (cheaper cold tiers) | Far cheaper than DB storage; tier to cut it further |
Capabilities in interviews
Media & user-upload store
Hold images, video, and documents cheaply while the database keeps queryable metadata.
The universal pattern: bytes in S3, pointer in the DB.
S3: s3://media/posts/991/video.mp4
DB: posts(id=991, ..., media_url="s3://media/posts/991/video.mp4", status="ready")The database indexes and queries the metadata (owner, status, created_at) with low latency; S3 holds the heavy bytes for cents. This is the shape used by file and media products: transactional metadata in a database, immutable or versioned bytes in object storage.
Choose this variant when
- User uploads of any kind
- Storing media referenced by a database row
- Anywhere a blob is bigger than a few KB
Direct-to-blob uploads
Pre-signed URLs (and multipart for big files) keep file bytes out of your app entirely.
The app issues a scoped, short-lived signed URL; the client uploads directly to S3:
POST /uploads → app validates auth/type/declared size → returns signed upload instructions
client → PUT/POST bytes → S3 → event/HEAD check → app marks the DB row "ready"For files over ~100 MB, the client uses multipart upload: split into chunks, upload parts in parallel, and retry only the failed part. A pre-signed PUT can scope the key, method, expiry, content type, checksum, and conditional headers, but it does not enforce a maximum byte size by itself; use a pre-signed POST policy with content-length-range for browser form uploads, or verify size after upload and reject/delete the object. The app's bandwidth bill is metadata and events, not gigabytes of media — which is the only way uploads scale past toy load.
Choose this variant when
- Any non-trivial file upload
- Large or flaky-network uploads (multipart/resumable)
- Protecting app servers from proxying bytes
Static asset & media delivery
Use the bucket as a CDN origin to serve assets and downloads globally from the edge.
S3 is the durable origin; a CDN caches objects near users so downloads are fast and the origin sees only misses:
client → CDN edge (hit) → bytes
↳ (miss) → S3 origin → cache + serveFor public assets, this is plain caching. For private content (paid video, user files), issue signed URLs / signed cookies with short expiry so the CDN only serves authorised requests. At a good hit rate this offloads most download traffic from the origin, but the exact percentage is cache-key and workload dependent.
Choose this variant when
- Serving images / video / downloads globally
- Static website or front-end asset hosting
- Paid or private media via signed URLs
Backups, archives & data lake
Cheap durable storage for backups, log archives, and analytics datasets — with lifecycle tiering.
Object storage is the default landing zone for backups, exported logs/events, and data-lake files that batch jobs (Spark, Athena, a warehouse loader) read later:
events → S3 (parquet, partitioned by date) → Athena / Spark / warehouse loadLifecycle policies move data through cheaper tiers as it ages (standard → infrequent-access → archive/Glacier) and eventually delete it, so cold data costs a fraction of hot data. As of 2026, AWS also has specialised S3 bucket types for Apache Iceberg tables, queryable object metadata, and vector indexes; name them only when the workload is analytics or AI search, not plain blob delivery.
Choose this variant when
- Database / system backups
- Log and event archival
- Data-lake / analytics staging with lifecycle tiering
Operating knobs
Storage class / lifecycle tiering
Match the class to access frequency: Standard for hot data, Infrequent-Access for occasional reads, Glacier/Archive for cold backups (cheap to store, slow/costly to retrieve). A lifecycle policy transitions objects automatically as they age and expires them at end of life — the main lever for controlling storage cost over time.
Access control & signed URLs
Buckets are private by default. Serve private content with short-lived pre-signed URLs (or CDN signed URLs/cookies) scoped to one object, rather than making the bucket public. A pre-signed PUT is a security boundary for the signed request, but it cannot enforce a max file size by itself; use a pre-signed POST policy with content-length-range, or validate size/checksum after upload and delete rejects.
Key / prefix design for throughput
AWS documents baseline request rates of at least ~3,500 PUT/COPY/POST/DELETE and ~5,500 GET/HEAD requests per second per partitioned prefix, and S3 scales automatically as load rises. Start with natural prefixes (tenant/date/object) and watch for 503 Slow Down while S3 repartitions. Only add explicit hash/shard prefixes for known extreme hot spots; randomising every key is outdated advice for most workloads.
Multipart & abandoned-upload cleanup
Use multipart upload for large files (resumable, parallel) and set a lifecycle rule to abort incomplete multipart uploads after N days — otherwise half-finished uploads silently accumulate storage and cost. Versioning (optional) protects against accidental overwrite/delete at extra storage cost.
Strong consistency & conditional writes
Object PUTs, overwrites, DELETEs, metadata reads, and LIST are strongly consistent in all AWS Regions. Use If-None-Match: * when a key must be created only if absent, and If-Match with the current ETag when overwriting only if nobody changed it (S3 documents both for PutObject, CompleteMultipartUpload and CopyObject). This reduces races in shared object namespaces; it does not make your database metadata transaction magically atomic with the upload.
Versus the alternatives
Object storage vs the alternatives.
| Dimension | S3 / Blob | Database (Postgres/Dynamo) | CDN |
|---|---|---|---|
| Stores | Large unstructured blobs | Structured rows / items | Cached copies of objects |
| Access | Whole-object GET/PUT by key | Query / index by field | Edge reads near users |
| Durability/scale | ~11 nines, unlimited | Durable, bounded by ops | Ephemeral cache |
| Cost | Cheapest per GB | Expensive per GB | Per request + egress |
| Role | Source of truth for files | Source of truth for metadata | Delivery layer over S3 |
Failure modes & gotchas
Streaming file bytes through your app server burns its bandwidth and memory and caps concurrency — ten simultaneous 500 MB uploads can pin gigabytes of RAM. Use pre-signed URLs for direct-to-S3 upload and a CDN for download; keep the app on the metadata path only.
Putting images or video as BLOB columns bloats the database, slows backups, and costs far more per GB than object storage. Store the bytes in S3 and a URL/key in the row — the database stays small and queryable.
Making a bucket public to "make it work" is a classic data-leak. Keep buckets private and serve via short-lived signed URLs scoped to one object. For uploads, remember that a pre-signed PUT cannot enforce a maximum byte size by itself; use a pre-signed POST policy with content-length-range, or validate and delete rejects after upload.
The database row and the object upload are not one transaction. A client can reserve a row and never upload, upload bytes and never call complete, or overwrite a key while metadata still points at the old expectation. Model rows as pending → ready, mark ready only after an S3 event or HeadObject verification, use conditional writes for shared keys, and reconcile old pending rows.
Since December 2020, S3 object PUTs, overwrites, DELETEs, metadata reads, and LIST are strongly consistent in all Regions. If users see stale bytes after an overwrite, first suspect CDN/browser caches, cache-control, or reused keys. Bucket configuration and delivery layers can still have propagation windows, but the object read/list model is not the old eventually-consistent one.
If the origin returns an error or you replace an object, the CDN can serve a cached error or stale bytes until TTL. Use content-hashed keys (a new version = a new key, never purge) for immutable assets, and never cache error responses with a long TTL.
Incomplete multipart uploads keep their uploaded parts (and bill for them) until explicitly aborted. Always set a lifecycle rule to abort incomplete uploads after a few days, and run a reconciler that expires stale pending DB rows and deletes orphan objects that never became visible to users.
Two clients writing the same key race; without a guard, the last completed write wins. Use content-hashed immutable keys when possible. If the key is a shared name, use If-None-Match: * for create-only writes or If-Match with the current ETag for compare-and-swap overwrites (PutObject, CompleteMultipartUpload and CopyObject support both).
In production
Video streaming origin and edge (illustrative)
Object storage plus a custom delivery edge
A large video service usually separates the storage origin from the delivery path. Object storage holds the durable encoded files, metadata tables track titles and renditions, and a CDN or private edge network places popular segments close to viewers. The origin should mostly see fills, cache misses, and administrative traffic, not every playback request.
The lesson is the boundary, not a memorised fleet size: keep heavy media bytes in durable object storage and metadata or workflow state in databases, then deliver through an edge layer so users are not streaming directly from the storage origin.
Dropbox
Built on S3 — then outgrew it at exabyte scale
Dropbox is a two-sided lesson. Its 2016 post says it had always used a hybrid architecture: metadata and web servers in Dropbox-managed data centers, and file content on Amazon S3. At that point Dropbox had surpassed 500 million signups and 500 petabytes of user data, up from about 40 petabytes in 2012.
Then Dropbox built Magic Pocket, its own storage system. The same post says Magic Pocket was serving over 90% of user data, had a goal to scale to more than 500 PB in six months, and used a high-performance network that peaked at over half a terabit per second during the buildout. This is the senior nuance: managed object storage is the right default and scales very far, but at rare exabyte-scale workloads, owning the hardware can be a rational cost/control trade.
Good vs bad answer
Interviewer probe
“Design the storage for a video platform where users upload videos up to a few GB and others stream them worldwide.”
Weak answer
"Store the videos in the database as binary so they're safe and transactional, and the app streams them out to viewers when requested."
Strong answer
"Videos go in S3, never the database — the DB just holds metadata with a pointer: videos(id, owner, status, s3_key). Upload: the client requests signed multipart upload instructions (videos are multi-GB, so multipart gives parallel chunks and resume on a dropped connection), uploads the bytes directly to S3 so our app never proxies gigabytes, and an S3 event or HEAD check flips the row from pending to ready after processing. I use a POST policy or post-upload validation for size limits, and conditional writes if a key must not already exist. Delivery: S3 is the durable origin behind a CDN, so streams are served from edge locations near viewers and the origin only sees cache misses; paid content uses signed URLs with short expiry. S3 gives us eleven-nines durability, strong object read/list consistency, and effectively unlimited capacity, with lifecycle rules to tier cold uploads and abort abandoned multipart uploads. Putting video in the database would bloat it, wreck backups, cost far more per GB, and force every byte through our app — the exact opposite of what scales."
Why it wins: Applies the bytes-in-S3/pointer-in-DB rule, uses multipart direct upload to keep the app off the data path, fronts delivery with a CDN + signed URLs, cites durability/cost/lifecycle, and explains precisely why the DB-blob approach fails.
Interview playbook
When it comes up
- Any system that stores files — images, video, documents, user uploads
- Static asset or media delivery at global scale
- Backups, log archives, or a data lake / analytics staging area
- The interviewer asks "where do the actual files live?"
Order of reveal
- 11. Bytes in S3, pointer in the DB. Large files go in object storage; the database keeps queryable metadata with a URL/key — never blobs in the DB.
- 22. Direct-to-blob upload. Clients upload straight to S3 via a pre-signed URL (multipart for big files), so the app stays off the data path.
- 33. Treat scale/durability/consistency as given. Eleven-nines durable, effectively unlimited, and strongly consistent for object writes/lists — I spend my time on data flow, not capacity folklore.
- 44. Deliver via CDN. S3 is the origin behind a CDN; downloads serve from the edge, signed URLs gate private content.
- 55. Cost via lifecycle. Lifecycle rules tier cold data to cheaper classes and clean up abandoned multipart uploads.
Signature phrases
- “Large bytes in object storage, a pointer in the database.” — The one rule that drives every file-storage decision.
- “The app never touches the file bytes — pre-signed direct upload.” — Shows you know how uploads actually scale.
- “S3 is the durable origin; the CDN is the delivery layer.” — Correctly separates storage from delivery.
- “Durability and capacity are givens — design the data flow.” — Focuses interview time where it matters.
Likely follow-ups
?“How does a multi-GB upload work without killing your servers?”Reveal
The client requests multipart signed upload instructions, splits the file into chunks (e.g. 100 MB), and uploads the parts in parallel directly to S3 — never through the app. A dropped connection retries only the failed part, not the whole file, and the client shows real progress. On completion, S3 fires an event or the app verifies with HeadObject, then post-processing (transcode, virus scan) marks the metadata row ready. The app issues URLs and records metadata; its bandwidth and memory are never spent on file bytes. A lifecycle rule aborts abandoned multipart uploads, and a reconciler expires old pending rows.
?“How do you serve private files — say paid videos — securely and fast?”Reveal
Keep the bucket private and serve through a CDN using signed URLs or signed cookies with a short expiry, scoped to the specific object and user. The client gets a time-boxed URL, the CDN validates the signature and serves from the edge (fast, global), and the origin bucket is never publicly readable. I only stream bytes through the app when I need per-request logic the CDN cannot do — like DRM license issuance — and even then only the license, not the media payload.
?“When would object storage be the wrong choice?”Reveal
When the data is structured and you query it by field, when you need low-latency random access or transactions, or when you need to modify parts of a file in place — objects are written and replaced whole, not edited. Those are database or cache jobs. Object storage is also overkill for tiny values where a row or a cache entry is simpler. It is specifically for large, unstructured, write-once-read-many blobs.
Worked example
Setup. Design storage and delivery for a video platform: users upload clips up to a few GB; millions of viewers stream them worldwide; the catalog grows without bound.
The move. Videos go in S3, never the database — the DB holds metadata with a pointer: videos(id, owner, status, s3_key, duration). Upload uses signed multipart instructions: because clips are multi-GB, the client splits the file into ~100 MB parts and uploads them in parallel, directly to S3, so the app never proxies a single byte; a dropped connection retries one part, not the whole file. On completion an S3 event or HeadObject verification triggers transcoding (into HLS renditions) and flips the row from pending to ready.
Delivery. S3 is the durable origin behind a CDN, so streams are served from edge locations near each viewer and the origin only sees cache misses; paid content uses signed URLs with short expiry so only entitled users can fetch.
Cost + durability. S3 gives ~11 nines of durability, strong object read/list consistency, and effectively unlimited capacity for ~cents/GB — I treat those as given and instead design the data flow. Lifecycle rules tier cold uploads (Standard → Infrequent-Access → Glacier) and abort abandoned multipart uploads so half-finished uploads don't silently accrue cost.
What breaks. A naive design proxies bytes through the app — ten concurrent 500 MB uploads would pin gigabytes of app RAM; the signed-upload direct path removes that entirely. The other traps are public buckets and upload/metadata races: I keep buckets private, use a pre-signed POST policy or post-upload validation for byte limits, mark rows ready only after verification, and reconcile orphan objects or stale pending rows.
The result. Multi-GB uploads that never touch the app servers, global low-latency streaming from the edge, eleven-nines durability at cents/GB, and a database that stays small and queryable — the storage shape behind essentially every media product.
Cheat sheet
- •Large unstructured bytes → object storage. Pointer (URL/key) → database. Never blobs in the DB.
- •Durability ~11 nines, effectively unlimited scale, and strong object read/list consistency — design the data flow.
- •Uploads go direct to S3 via short-lived signed PUT/POST/multipart instructions; the app never proxies bytes.
- •A pre-signed PUT cannot enforce max byte size by itself; use POST `content-length-range` or post-upload validation.
- •Files over ~100 MB → multipart (parallel, resumable). Abort abandoned uploads via lifecycle and reconcile pending rows.
- •Serve downloads via a CDN with S3 as origin; signed URLs/cookies gate private content.
- •Objects are whole-write, no in-place edit — replace to change; great for media, bad as a DB.
- •Lifecycle tiering (Standard → IA → Glacier → delete) is the main cost lever.
- •Use `If-None-Match: *` for create-if-absent; use `If-Match` + ETag for guarded overwrites.
Drills
Why store a video in S3 and only a URL in the database instead of the video itself?Reveal
Because databases are built for small, structured, queryable rows with low-latency indexed access, and object storage is built for large, cheap, durable bytes. A multi-GB BLOB column bloats the database, makes backups slow and huge, costs roughly an order of magnitude more per GB, and forces every byte through the database connection. Keeping the bytes in S3 and a pointer in the row keeps the database small and fast for the queries it is good at, while S3 handles the heavy storage and a CDN handles delivery.
Interviewer: "10,000 users upload simultaneously. Walk me through it without overloading the app."Reveal
Each client POSTs metadata; the app validates auth/type/declared size, creates a pending row, and returns signed upload instructions (multipart for large files), then is done — it does not receive the bytes. All 10,000 clients upload directly to S3 in parallel; S3 scales automatically to very high request rates, so the app stays off the byte path; clients retry with backoff on the occasional 503 Slow Down while S3 scales a hot prefix. On completion S3 fires an event or the app verifies with HeadObject, and a worker pool post-processes (transcode/scan) and flips the DB row to ready. A reconciler expires old pending rows and deletes orphan objects.
Your CDN is serving a stale image after you replaced it. How do you avoid this?Reveal
Use content-hashed keys for immutable assets: when the image changes, write a new object under a new key (e.g. avatar.<hash>.png) and update the pointer, so the new URL is a guaranteed cache miss and the old one simply ages out — no purge needed. If you must reuse the same key, explicitly invalidate/purge it at the CDN after replacing, accepting propagation delay. Either way, set sensible cache-control, and never cache error responses with a long TTL.
What it is