MongoDB
A document database for aggregate-shaped data, flexible schemas, replica-set availability and shard-key-based horizontal scale.
Also worth naming: MongoDB Atlas · MongoDB Community Server · MongoDB Enterprise Server
MongoDB is strongest when the product naturally reads and writes whole documents. The interview skill is knowing where that stops: unbounded arrays, cross-document invariants, indexes and shard keys.
What it is
MongoDB is a document database. A record is a BSON document with fields, nested documents and arrays, and MongoDB's manual says data that is accessed together should be stored together in a model built from application access patterns (data modelling). That is the reason to choose it: the product object is naturally an aggregate, not a set of rows you mostly join at read time.
The practical rule is embed when bounded and owned; reference when unbounded or independent. MongoDB's schema guidance says embedding fits data that is queried together, updated together, or archived together, while referencing fits high-cardinality children, embedded data that grows without bounds, or data written at different times in a write-heavy workload (best practices). This agrees with Data model design: a product can embed five latest reviews, but the full review list belongs in its own collection.
The hard physical limit matters in interviews. A MongoDB document has a 16 MiB maximum BSON document size, including embedded arrays and documents. That limit is not just a storage trivia fact; it is the warning against chat messages in one room document, all comments in one post document, or any array whose terminal cardinality you cannot bound.
MongoDB has modern correctness tools: single-document writes are atomic, and multi-document ACID transactions work across documents, collections, databases and shards (transactions). Still, the manual is explicit that distributed transactions cost more than single-document writes and are not a substitute for effective schema design. In an interview, do not pick MongoDB because it is schemaless; pick it because the aggregate boundary makes the common read and write one document.
When to reach for it
Reach for this when…
- The main entity is an aggregate that is usually read and written as one document
- Schema varies by tenant, product type or event payload, but access patterns are still known
- You want single-document atomicity, indexes on document fields and replica-set failover
- You need change streams to keep search, cache or analytics projections fresh
Not really this pattern when…
- The core model is relational and joins, constraints, ad-hoc SQL or multi-row transactions dominate: use PostgreSQL
- The access pattern is strictly key-based at very large scale and you want a fully managed partitioning model: use DynamoDB
- Child arrays grow without a bound or have their own access patterns: reference them, or reconsider the store
- You need a durable event log with long retention and many consumer groups: use Kafka, not only change streams
How it works
Five mechanisms explain how MongoDB behaves.
1. The document is the atomic aggregate. A single-document write is atomic even if it updates embedded sub-documents and arrays (data modelling best practices). That makes embedded, bounded children powerful: update the product and its embedded latest review snippet together. It also makes unbounded arrays dangerous, because every item lives inside the same size, growth and update boundary.
MongoDB works best when one screen or command mostly reads and writes one document aggregate. Embed bounded, owned children; reference unbounded or independently queried children.
2. Indexes are query-shaped. Compound indexes obey prefixes, can contain up to 32 fields, and usually follow the ESR guideline: equality fields first, then sort, then range, unless a highly selective range should move before sort (ESR guideline). Multikey indexes index array elements, but a compound multikey index can include at most one indexed array field per document, hashed indexes cannot be multikey, and a multikey index cannot be the shard key index (multikey indexes).
3. A replica set gives high availability, not magic strong reads from every node. The primary accepts writes and records them in the oplog; secondaries copy and apply the oplog asynchronously (replication, oplog). If the primary is unavailable, an eligible secondary is elected; the manual says the median time to elect a new primary should not typically exceed 12 seconds under default settings. Reads use read preference primary by default, while secondary reads can be stale.
Writes go to the primary and are recorded in the oplog. Secondaries copy and apply that oplog asynchronously; if the primary is unavailable, an eligible secondary is elected.
4. Write concern, read concern and read preference are separate knobs. Since MongoDB 5.0, the implicit default write concern is usually majority, with the documented arbiter formula falling back to w:1 when data-bearing voting members are not greater than the voting majority (write concern). The default read concern for reads on primary and secondaries is local, which may read data that has not been majority committed and may roll back (local read concern).
Read concern majority returns data acknowledged by a majority and guaranteed not to roll back for non-transaction reads (majority read concern). Linearizable read concern applies to primary-only single-document reads that uniquely identify a document and may be slower (linearizable read concern). Snapshot read concern returns majority-committed data from one point in time, including transactions and, since MongoDB 5.0, supported reads outside transactions (snapshot read concern). Read preference decides where the read goes: primary is the usual route, while secondary modes can return stale data and should use staleness controls when needed (read preference).
5. Sharding is a shard-key design, not a switch. A sharded cluster has mongos routers, config servers and shard replica sets; MongoDB partitions a collection by the shard key into ranges or hashed ranges and routes queries with the shard key to targeted shards (sharding, shard keys). Hashed sharding spreads monotonic keys more evenly but hurts range locality; ranged sharding preserves range scans but can create a hot tail (hashed, ranged). This is the same access-pattern-first discipline as Sharding and partitioning.
The shard key partitions a collection into ranges or hashed ranges. Queries that include the shard key can target shards; missing it turns into scatter-gather.
Performance envelope
MongoDB performance envelope: sourced limits and defaults; throughput and latency depend on hardware, document shape, indexes, working set and deployment topology.
| Dimension | Number or behaviour | Why it matters |
|---|---|---|
| Document size | [16 MiB BSON document limit](https://www.mongodb.com/docs/manual/core/gridfs/) | Caps embedded arrays and documents; use references or GridFS-style patterns for larger logical objects |
| Compound index width | [Up to 32 fields](https://www.mongodb.com/docs/manual/core/indexes/index-types/index-compound/) | Enough for real query shapes, but each index slows writes and consumes memory |
| Replica failover | Median election time should not typically exceed [12 seconds](https://www.mongodb.com/docs/manual/core/replica-set-elections/) with defaults | Expect a write pause during failover and rely on driver retry behaviour |
| Default write concern | Since MongoDB 5.0, usually [w: majority](https://www.mongodb.com/docs/manual/reference/write-concern/#std-label-wc-default-behavior), with arbiter edge cases | Majority acknowledgements reduce rollback risk; w:1 writes can roll back on failover |
| Default read concern | [local](https://www.mongodb.com/docs/manual/reference/read-concern-local/) on primary and secondaries | Fast, but can read data that may roll back; use majority or snapshot where needed |
| Read preference | [primary is the normal routing choice](https://www.mongodb.com/docs/manual/core/read-preference/); secondary modes may be stale | Separate routing from read concern, and use max staleness or causal sessions when freshness matters |
| Chunk range size | Default range size is [128 MB](https://www.mongodb.com/docs/manual/core/sharding-data-partitioning/#std-label-sharding-chunk-size) | Large or unsplittable chunks affect balancing and can become jumbo chunks |
| Transaction runtime | Default transaction runtime is under [one minute](https://www.mongodb.com/docs/manual/core/transactions-production-consideration/) | Long transactions add cache pressure, lock waits and abort risk |
Capabilities in interviews
Aggregate documents and flexible schema
Store the data a screen needs together, while allowing different documents to carry different fields.
MongoDB's document model is useful when a product object is naturally nested: product details with bounded latest reviews, user profile settings, device state, content metadata or polymorphic event payloads. MongoDB's data modelling manual explicitly says to structure data from access patterns and keep data accessed together together (data modelling).
The trap is confusing flexible schema with no schema. You still choose required fields, indexes, validation rules and migration behaviour. If a child can grow without bound, reference it and paginate. If two pieces are updated at different rates, reference them or accept that the parent document becomes the write hotspot.
Choose this variant when
- One aggregate is the main read and write unit
- Fields vary by subtype or tenant
- Bounded embedded children avoid repeated joins
Field, compound and multikey indexes
Index document fields and arrays, then order compound keys with equality, sort and range in mind.
MongoDB's common performance move is a compound index that matches a query filter and sort. The ESR guideline says equality fields come first, then sort fields, then range fields; when the range predicate is very selective, ERS can be better even if it causes an in-memory sort (ESR guideline).
Arrays use multikey indexes. They are powerful for tags, genres and short membership lists, but they come with constraints: each document in a compound multikey index can have at most one indexed array field, hashed indexes cannot be multikey, and multikey indexes cannot be used as shard key indexes (multikey indexes).
Choose this variant when
- Queries filter by a few document fields and sort by one or two fields
- Arrays are bounded and queried by membership
- You can prove the index with explain rather than adding speculative indexes
Replica sets, concerns and preferences
Use one primary for writes, secondaries for redundancy or stale-tolerant reads, and concerns to state the guarantee.
Replica sets are the default production shape. The primary receives writes, records them in the oplog, and secondaries copy and apply those operations asynchronously (replication). Elections keep the set available, but writes pause until a primary exists.
State the three knobs separately. Write concern majority means the write is acknowledged after a majority condition is met and, with the default majority journaling behaviour, lowers rollback risk (write concern). Read concern local is the default and can return rollback-prone data; majority, linearizable and snapshot read concerns each add stronger but narrower guarantees (local, majority, linearizable, snapshot). Read preference primary is the default route; secondary reads trade freshness for read capacity or locality.
Choose this variant when
- You need high availability and automatic failover
- Some reads can tolerate lag and move to secondaries
- The answer needs an explicit durability and stale-read story
Multi-document ACID transactions
Use transactions when one invariant really spans documents, but treat them as the exception rather than the modelling default.
MongoDB supports ACID transactions across operations, collections, databases, documents and shards (transactions). MongoDB 4.0 added multi-document transactions scoped to replica sets, and MongoDB 4.2 extended them to sharded clusters (4.0 transaction announcement, 4.2 distributed transactions).
The cost is real. The manual says distributed transactions incur greater performance cost than single-document writes and should not replace good schema design. Production considerations include the default runtime limit under one minute, lock waits, cache pressure, chunk migration conflicts and multi-shard commit behaviour (production considerations, sharded transactions).
Choose this variant when
- A workflow updates several documents and partial success is a correctness bug
- You cannot remodel the invariant into one document
- You can keep the transaction short and avoid slow external calls inside it
Sharding by shard key
Scale a collection horizontally by choosing a shard key that distributes load and keeps hot queries targeted.
A shard key is one indexed field or multiple fields covered by a compound index that determines how documents are distributed (shard keys). The ideal key spreads documents evenly and supports common query patterns. Queries with the shard key or a compound shard-key prefix can target shards; missing it can broadcast to every shard (sharding).
Hashed sharding is often the right answer for monotonic keys and point lookups because it spreads inserts, while ranged sharding keeps close values together for range scans at the cost of possible hot tails (hashed, ranged). Jumbo chunks are chunks beyond the configured range size that cannot split, often because one shard key value is too frequent; MongoDB 5.0 added resharding so you can change the shard key, but it is still an operational migration (data partitioning, reshard a collection).
Choose this variant when
- One replica set no longer has storage or write headroom
- A shard key can keep the hottest reads targeted
- You have a plan for hashed vs ranged keys, jumbo chunks and resharding
Change streams
Watch inserts, updates, deletes and DDL events with resume tokens for projection pipelines.
Change streams let an application watch a collection, database or deployment for change event documents on replica sets and sharded clusters (change streams, change events). This is the MongoDB-native way to drive search indexing, cache invalidation, notifications or analytics projection.
They are not a full Kafka replacement. Each open stream holds a connection while waiting for events, sharded clusters open a stream on each shard through mongos, and resuming requires enough oplog history for the token. Use this with the discipline from CDC and eventing: consumers should be idempotent, checkpoint resume tokens and have a backfill plan.
Choose this variant when
- Search, cache or analytics should follow MongoDB writes
- You need a resumable feed of database changes
- The event volume and retention needs fit the oplog-backed stream model
Operating knobs
Embed vs reference
Embed when the child is bounded, owned by the parent, read with the parent and updated with the parent. Reference when the child is high-cardinality, independently queried, written at a different rate, or can grow without bounds. This is the single most important MongoDB modelling decision, and it should be defended with terminal cardinality, not launch-day data.
Concern and preference defaults
Since MongoDB 5.0, the default write concern is usually majority, except documented arbiter topologies can fall back to w:1. The default read concern is local, and the default read preference is primary. Linearizable and snapshot read concerns are stronger tools for narrower cases, not replacements for a routing decision. Read-your-writes or rollback-sensitive paths should explicitly combine majority write concern, majority or snapshot read concern, causal sessions where useful, and primary reads where stale secondary reads would surprise users.
Index order and multikey constraints
Use ESR for compound indexes: equality first, sort next, range last unless a highly selective range deserves ERS. Keep array fields short and intentional. A compound multikey index cannot index two array fields in the same document, hashed indexes cannot be multikey, and a multikey index cannot be the shard key index. That shapes both schema and shard-key choices.
Shard key and resharding
Choose the shard key from the dominant query and write distribution. Hashed keys smooth monotonic writes and point lookups; ranged keys preserve locality for range scans. MongoDB 5.0 added resharding for changing a bad key, but it needs resources, write blocking and application updates, so treat it as a migration plan, not permission to ignore the first key.
Transactions vs aggregate design
MongoDB transactions are correct and useful, but the cheap path is still one document. If a transaction appears in the hot path, ask whether the invariant can be remodelled into one aggregate or whether Postgres is a better source of truth. Keep necessary transactions short, use transaction-level read and write concerns, and expect retries on write conflicts or failover.
Versus the alternatives
MongoDB versus the closest datastore choices.
| Dimension | MongoDB | PostgreSQL + JSONB | DynamoDB |
|---|---|---|---|
| Best data shape | Aggregate documents with nested, bounded children | Relational core plus flexible JSON fields | Known key-value or partition-key access patterns |
| Query model | Rich document queries, secondary indexes, aggregations; no relational join default | SQL joins, constraints, ad-hoc queries and JSONB indexes | GetItem and Query by keys; GSIs for extra access patterns |
| Correctness default | Single-document atomic; multi-document transactions available with cost | ACID transactions, constraints and joins are the native model | Single-item conditional writes; transactions exist but key design comes first |
| Scaling model | Replica sets, then shard by chosen key through mongos | Vertical scale, read replicas, partitioning; sharding is manual or via extensions | Managed partitioning by primary key; hot keys and item limits still matter |
| Choose when | Product object is naturally a document and schema varies | Relationships, constraints and flexible querying matter most | Very large key-based workload with minimal operational burden |
Failure modes & gotchas
A room document that embeds every message or a post document that embeds every comment eventually collides with the 16 MiB BSON document limit and becomes a write hotspot. Reference unbounded children, paginate them, and embed only a bounded preview such as latest comments.
If most reads need data scattered across several collections, the app starts doing N queries and stitching results. That is the sign the workload is relational or needs a separate read model. MongoDB can aggregate and look up, but the happy path is still aggregate-local reads.
Secondaries apply the oplog asynchronously, so a user who writes and immediately reads from a secondary can see stale data. Route read-after-write flows to the primary, use causal sessions, or require the replica to reach the relevant cluster time. Do not treat read preference secondary as free capacity for every path.
An index ordered by range before equality or missing the sort prefix can force extra scans or in-memory sorts. Start from the hot query, apply ESR, confirm with explain, and remove redundant prefix indexes only when they are truly redundant.
A low-cardinality or very frequent shard key value can create jumbo chunks that cannot split smaller than one unique shard-key value. Refining or resharding may be needed, but that is an operational migration. Pick cardinality and distribution deliberately at design time.
Large, long or frequent transactions add locks, cache pressure, aborts and retry complexity. They are correct for real cross-document invariants, but if the hot path constantly needs them, either remodel the aggregate or consider Postgres for that invariant.
In production
Marketplace catalogue service (illustrative)
Flexible product attributes without one giant relational table
A marketplace catalogue can store product attributes as documents because shoes, phones and furniture do not share one clean set of columns. The product page reads one aggregate, so MongoDB fits. Reviews, seller history and stock adjustments still get their own collections or stores because they grow without bound or have separate correctness requirements.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“You are designing a product catalogue with flexible attributes, reviews and checkout inventory. Would you use MongoDB?”
Weak answer
"Yes, because MongoDB is schemaless and scales horizontally. I will store each product with all reviews in one document and shard it later if it gets big."
Strong answer
"I would use MongoDB for the catalogue document, not for the checkout invariant. Product attributes vary by category, and the product detail page reads one aggregate, so a product document fits well. I would embed bounded data like price display fields, category attributes and maybe the latest five reviews, but I would reference the full reviews collection because reviews are unbounded and have their own pagination.
I would index the catalogue around real queries: for example category equality, then sort by rank, then a range such as price if that filter matters, following the ESR guideline. If the catalogue outgrows one replica set, the shard key might be a compound key such as category plus hashed product id, depending on the browse pattern.
For checkout inventory, I would not hide behind MongoDB flexibility. If stock decrement and order creation must be atomic across records, I either keep that invariant in Postgres or use a short MongoDB transaction with majority commit and primary reads. Counts, search and recommendations can lag and follow through change streams."
Why it wins: It chooses MongoDB for the aggregate-shaped part, rejects unbounded arrays, names the index shape, separates the checkout invariant, and links stale-tolerant projections to change streams.
Interview playbook
When it comes up
- The prompt has product catalogues, profiles, metadata, content documents or polymorphic events
- Someone says document database, NoSQL, flexible schema or MongoDB Atlas
- A nested object has bounded children and the hot read is one aggregate
- The design needs change streams, replica-set failover or shard-key scale
Order of reveal
- 11. Justify by aggregate shape. I would use MongoDB if the main read and write is one aggregate document, not merely because the schema can vary.
- 22. Draw the embed/reference boundary. I embed bounded children read with the parent, and reference unbounded or independently queried children.
- 33. Name the indexes. For the hot query I order the compound index by equality, then sort, then range, and I check multikey limits if arrays are involved.
- 44. State replica-set guarantees. Writes go to the primary. The current default write concern is usually majority, reads default to local concern and primary preference, and secondary reads may lag.
- 55. Spend transactions and sharding carefully. Transactions are for real cross-document invariants; sharding depends on the shard key, not on MongoDB being NoSQL.
Signature phrases
- “Embed bounded, owned children; reference unbounded children.” — The cleanest rule for MongoDB schema design.
- “A flexible schema is not the same as no schema.” — Shows you still plan validation, indexes and migrations.
- “Read concern, write concern and read preference answer different questions.” — Prevents vague consistency claims.
- “The shard key is the scaling decision.” — Connects MongoDB sharding to the broader partitioning lesson.
Likely follow-ups
?“When would you pick Postgres with JSONB instead?”Reveal
When relationships and correctness dominate. If the product needs foreign keys, joins, uniqueness constraints across entities, reporting SQL or many transaction boundaries, Postgres should be the source of truth and JSONB can hold the flexible slice. MongoDB is stronger when the flexible document itself is the main object and cross-document relationships are secondary.
?“When would you pick DynamoDB instead?”Reveal
When the access pattern is strictly key-based and the reason for the choice is managed horizontal scale, not rich document querying. DynamoDB wants partition-key and sort-key design up front, and it has a smaller query surface. MongoDB gives richer secondary indexes and document queries, but you operate or buy a cluster whose shard key still matters.
?“How do you make read-your-writes work?”Reveal
Keep that read on the primary for a short window, or use a causally consistent session so the driver includes the session operation time. Pair majority write concern with majority read concern when you need rollback-resistant reads. Do not send a user directly to an arbitrary secondary after their write and expect freshness.
Worked example
Numbers in this section are illustrative.
Setup. Design storage for an e-commerce catalogue. Products have category-specific attributes, variants, seller-provided metadata, reviews, inventory and search.
The MongoDB-shaped part. The catalogue read is usually one product detail page, so I store one product document with stable fields, flexible category attributes and a bounded embedded preview: latest reviews, variant summaries and display-ready price metadata. This fits the document rule: accessed together, stored together. I add schema validation for required fields because flexible does not mean unvalidated.
The boundary. Full reviews are referenced in a reviews collection with product_id and created_at because reviews are unbounded and paginated. Inventory is not embedded in the product document if checkout decrements it under contention. That invariant either lives in Postgres or in a short MongoDB transaction with clear write concern and retry handling.
Indexes. Browse by category and rank uses an index shaped like category equality, rank sort, then optional price range. Attribute filters need carefully chosen indexes or Atlas Search if the query becomes search-like.
Replication and freshness. The replica set uses the current defaults as the baseline: write concern is usually majority, read concern local, read preference primary. Product detail reads can go to secondaries only if a little lag is acceptable. Seller edit pages read from the primary or use causal sessions so the seller sees their own changes.
Scale. If one replica set runs out of headroom, the shard key follows the dominant query. For browse-heavy catalogues, a compound key that starts with tenant or category and includes a hashed product id may keep reads targeted while spreading writes.
Derived views. Search, recommendations and analytics follow from change streams. Consumers checkpoint resume tokens and upsert idempotently. If the stream falls too far behind the oplog window, the recovery plan is a backfill plus resume from a safe point, not hope.
Result. MongoDB serves the flexible catalogue aggregate, references unbounded children, keeps checkout invariants out of the document if needed, and scales only after the shard key is tied to the access pattern.
Cheat sheet
- •Use MongoDB when the aggregate document is the natural read and write unit.
- •Embed bounded, owned children; reference unbounded or independently queried children.
- •A document is limited to [16 MiB](https://www.mongodb.com/docs/manual/core/gridfs/), so unbounded arrays are design bugs.
- •Indexes are access-pattern contracts: compound prefixes, ESR, and multikey limits matter.
- •Replica set: primary writes, async secondaries, elections on failure, oplog as the replication feed.
- •Current defaults: write concern usually majority, read concern local, read preference primary.
- •Transactions are ACID and useful, but costlier than one-document writes; keep them short.
- •Shard key choice decides scale: hashed for spread, ranged for locality, watch jumbo chunks.
- •Change streams are resumable CDC-like notifications; link them to /learn/cdc-eventing discipline.
Drills
Numbers in this section are illustrative.
A chat room document embeds all messages. What goes wrong?Reveal
Messages are unbounded. The document grows toward the 16 MiB BSON limit, every message append updates the same document, and pagination becomes awkward. Store the room as one document and messages as separate documents keyed by room_id and created_at, with maybe a bounded latest-message preview embedded in the room.
What do write concern majority, read concern majority and read preference primary each guarantee?Reveal
Write concern majority controls when a write is acknowledged and reduces rollback risk. Read concern majority controls what a read may return: majority-acknowledged data that will not roll back for non-transaction reads. Read preference primary controls where the read is routed. They are separate knobs, and you often need all three named clearly.
When would you use linearizable or snapshot read concern?Reveal
Use linearizable read concern for a primary-only read of one uniquely identified document when you need it to reflect successful majority-acknowledged writes before the read starts. Use snapshot read concern when you need a majority-committed point-in-time view, especially inside a transaction or supported snapshot read.
Why can a hashed shard key be good for writes but poor for range queries?Reveal
Hashing spreads adjacent key values across shards, which helps monotonic inserts avoid a hot tail. The cost is locality: close original values are no longer close after hashing, so range queries are more likely to broadcast or touch many chunks. Pick hashed for point lookups and spread; pick ranged when range locality is the product query.
What it is