Cost modelling
Pricing a design roughly: unit costs, egress, storage tiers, compute vs memory, cost per request and the per-user cost envelope.
Cost modelling turns a scale estimate into a price-shaped design. You will practise finding the bill that dominates, then choosing the simpler or cheaper architecture when it still meets the product promise.
Read this if your last attempt…
- Your last design sized capacity but did not ask what it costs
- You added a multi-region replica without pricing the traffic between regions
- You treated a cache tier as a QPS problem and forgot that RAM is the expensive part
- You could not answer "what dominates the bill?" with a ranked list
The concept
Cost modelling is a rough financial model for an architecture. Start with the numbers from capacity estimation, especially the load and storage units in the reference numbers card. Do not repeat that lesson's arithmetic. Use its QPS, storage growth, hot set and bandwidth as the inputs to a bill.
The method is the part to remember. AWS's Cost Optimization Pillar, publication date 2024-06-27 and opened 2026-09-28, frames cost optimisation as meeting requirements at the lowest practical price point. In an interview, that means you can justify a design by cost as well as by latency or reliability.
Take the load, storage and bandwidth estimates from capacity estimation, multiply by unit rates, then sort the monthly lines by size.
Cost levers: choose the one that attacks the largest line item.
| Lever | Use when | Risk to say out loud |
|---|---|---|
| Right-size compute | CPU or instance-hours dominate and load is predictable. | Too small hurts latency and leaves no failure headroom. |
| Right-size cache memory | The hot set is smaller than the full dataset. | Evictions can push traffic back to the database. |
| Tier storage | Old data is rarely read but must be retained. | Retrieval fees, restore delay and minimum duration can erase savings. |
| Reduce egress | Large payloads leave the region or origin repeatedly. | More caching or compression can add staleness and invalidation work. |
| Commit or reserve | A stable baseline runs most hours. | Unused commitment becomes waste if demand or instance shape changes. |
| Use spot or interruptible capacity | Work is queued, stateless, retryable or batch-like. | The design must handle interruption without user-visible loss. |
- Sort the ledger first. Optimising a small line item is theatre.
- Use illustrative prices in interviews unless the prompt supplies a source.
How interviewers grade this
- You turn QPS, storage and bandwidth from capacity estimation into billable units.
- You rank the bill before optimising: compute, memory, storage, requests and egress.
- You call out egress across internet, regions and availability zones before it surprises the design.
- You explain why cache cost usually tracks RAM and replicas more than raw request count.
- You choose on-demand, commitment or spot based on workload stability and interruption tolerance.
- You can say when a cheaper design is the right design because it still meets the product requirements.
Variants
Unit-rate ledger
A small spreadsheet that turns capacity inputs into monthly cost lines.
Use this for any system-design round once you have the capacity pass. Keep one row per billable dimension:
- Compute: instance count × hours × price per instance-hour.
- Memory: provisioned cache GB × hours × price per GB-hour or node-hour.
- Storage: GB stored × price per GB-month, split by tier.
- Requests: request count × price per request unit.
- Egress: GB transferred across a charged boundary × price per GB.
Then sort descending. The largest row decides where to spend design effort. If egress is the biggest row, changing instance families is a distraction. If cache RAM is the biggest row, compression and hot-set sizing matter more than app CPU.
Pros
- +Easy to explain at a whiteboard.
- +Makes hidden costs visible before choosing infrastructure.
- +Works across cloud vendors because the units are common.
Cons
- −Averages can hide peak commitments and minimum charges.
- −Needs periodic refresh because provider prices and workload shape change.
Choose this variant when
- You need to answer "what dominates the bill?" quickly.
- The design has several resource types and the expensive one is not obvious.
Egress-first model
Start from bytes crossing boundaries before you size more servers.
Use this for media, feeds, backups, analytics exports and multi-region designs. The math is the same as bandwidth sizing, but the boundary matters.
Ask:
- User egress: how many bytes leave for browsers or mobile clients?
- Origin egress: how often does a CDN miss and fetch from origin?
- Cross-AZ traffic: do chatty services or replicas sit on opposite sides of an availability-zone boundary?
- Cross-region traffic: are you replicating the full write stream or only compacted state?
The cheapest fix is often architectural: cache at the edge, compress responses, keep chatty dependencies in the same zone, precompute thumbnails, or replicate summaries instead of raw events.
Pros
- +Catches the line item candidates most often forget.
- +Connects cost to latency and data-placement decisions.
Cons
- −Can over-optimise locality if availability or disaster recovery requires another copy.
- −Needs product input on freshness and media quality.
Choose this variant when
- The system serves large media or public downloads.
- The architecture crosses regions or availability zones on a hot path.
Lifecycle tiering
Move data from hot to infrequent to archive when access probability falls.
Use this when storage grows with retention. The lifecycle design is not just "put old data in archive." It needs a rule and a retrieval story.
Example policy, all illustrative:
- 1Keep uploads hot while they are edited and shared often.
- 2Move originals to infrequent access after the active sharing window.
- 3Keep small thumbnails hot because they sit on the request path.
- 4Archive originals only when restore delay is acceptable.
- 5Delete derived files when they can be regenerated cheaper than stored.
AWS's S3 Lifecycle docs describe transition and expiration actions; the same mental model applies to other object stores. Name the lifecycle rule in your design, then state what user action would pay a retrieval cost.
Pros
- +Cuts storage cost without changing the user-facing hot path.
- +Forces a clear answer about retention and restore experience.
Cons
- −Retrieval and transition costs can erase savings if data is still active.
- −Archive restore delay can break product expectations.
Choose this variant when
- Logs, backups, media originals, audit records or exported reports dominate storage.
- Retention is required but most reads happen soon after write.
Pricing-model mix
Blend on-demand, committed and interruptible capacity by workload shape.
Use this when compute is a meaningful line item. Split the fleet into three buckets:
- Baseline: steady services that run most hours. Consider commitments once usage is stable.
- Burst: user traffic or launches that need capacity now. Keep on-demand or autoscaled capacity.
- Flexible: batch, backfill, video transcode, analytics and test jobs. Consider spot or interruptible capacity if the work checkpoints.
The interview signal is not "commit everything." It is matching the pricing model to risk. You can be cost-aware and still keep on-demand capacity for the parts that protect availability.
Pros
- +Saves money where utilisation is stable.
- +Keeps volatile or user-facing capacity flexible.
- +Makes batch work cheaper without risking the request path.
Cons
- −Commitments can become waste after migrations or demand shifts.
- −Spot needs retry, checkpoint and drain logic.
Choose this variant when
- The interviewer asks how you would reduce a compute bill.
- The workload mixes steady services with batch or backfill jobs.
Worked example
Numbers in this section are illustrative.
Scenario: a photo-sharing app has already done capacity estimation. We will cost the design without redoing that page.
Every number below is illustrative.
Capacity inputs from the estimate:
- API fleet: 20 app instances at peak, average 12 instances across the month.
- Cache hot payload: 250 GB; provisioned memory after overhead and replicas: about 1 TB.
- Object storage: 50 TB hot media and 120 TB older originals.
- Requests: 600M API requests and 900M object reads per month.
- Internet egress: 35 TB per month after CDN caching.
- Cross-region backup transfer: 10 TB per month.
Unit-rate ledger:
| Line | Illustrative calculation | Monthly cost |
|---|---|---|
| App compute | 12 instances × 730 hours × $0.08/hour | ~$700 |
| Cache memory | 1,000 GB × 730 hours × $0.010/GB-hour | ~$7,300 |
| Hot object storage | 50,000 GB × $0.020/GB-month | ~$1,000 |
| Older originals | 120,000 GB × $0.004/GB-month | ~$480 |
| Requests | 1.5B requests × $0.40/million | ~$600 |
| Internet egress | 35,000 GB × $0.080/GB | ~$2,800 |
| Cross-region transfer | 10,000 GB × $0.020/GB | ~$200 |
Sorted result: cache memory dominates, then egress, then storage. Compute is not the main cost, even though it is the most visible part of the architecture diagram.
Design response:
- 1Cache memory: cache small feed cards, compress if CPU allows, and cache the active slice. Dropping provisioned cache from an illustrative 1 TB to 500 GB beats changing app instance families.
- 2Egress: keep resized images at the CDN and avoid origin misses for public content.
- 3Storage: keep thumbnails hot, lifecycle originals, and archive only data whose restore delay is acceptable.
- 4Compute: keep app servers on-demand until traffic stabilises. Move reprocessing and backfills to checkpointed interruptible workers.
Per-user envelope:
The illustrative monthly total is about $13K. With 1M monthly active free users, that is roughly 1.3 cents each, but averages hide abuse. A public uploader can cost much more than a quiet reader, so free-tier quotas should cap storage, transformations and egress.
Interview conclusion: "The cheap design is not fewer app servers. It is smaller cached objects, stronger CDN behaviour, lifecycle rules for originals, and interruptible workers for reprocessing. That keeps the user experience while attacking the largest lines."
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“What dominates the bill, and would you change the design because of it?”
Weak answer
"Probably compute, because we have many servers. I would use autoscaling and maybe reserved instances. Storage is cheap, and the CDN should handle bandwidth."
Strong answer
"I would rank the lines before optimising. From the estimate, app compute is modest. The likely top lines are cache RAM, because the hot set needs replicas and headroom, and egress, because media leaves the system. I would first shrink the cached object and cache only the active slice. Then I would push resized media to the CDN and add lifecycle rules for originals. I would reserve or commit only the stable baseline; image backfills can run on interruptible workers. If the cheaper design still meets latency and retention, it is the better answer."
Why it wins: The strong answer does not guess from the diagram. It turns the design into cost lines, identifies the dominant ones, and chooses trade-offs that preserve the product promise.
When it comes up
- After capacity estimation, when the interviewer asks if the design is affordable.
- When you propose multi-region, CDN, a large cache, object storage, search, analytics or a managed database.
- When the interviewer asks "what dominates the bill?"
- When a simpler design might be better than a more scalable one.
Order of reveal
- 1Start from the capacity page. "I will reuse the QPS, storage, hot set and bandwidth from capacity estimation, then convert them into billable units."
- 2Name the ledger rows. "I will split the bill into compute hours, cache memory, storage GB-months, request count and egress GB."
- 3Sort before optimising. "I will sort the monthly lines first, because the design lever should attack the largest line item."
- 4Call out egress early. "I want to price internet, cross-region and cross-AZ traffic separately, because those lines hide between boxes."
- 5Pick pricing model by workload shape. "Baseline stays eligible for commitment after it stabilises; burst stays on-demand; retryable batch can use interruptible capacity."
- 6Close with the trade-off. "If a cheaper design meets the same latency, durability and product needs, I would choose it and state the limit where it changes."
Signature phrases
- ““Let’s rank the bill before we optimise it.”” — Prevents random cost-saving ideas from taking over.
- ““This cache is mostly a RAM bill.”” — Shows why hot-set sizing matters.
- ““Egress is a boundary cost, so I’ll price every boundary crossing.”” — Catches internet, region and zone transfer.
- ““The cheaper design wins if it still meets the requirement.”” — Makes cost a legitimate design constraint.
Likely follow-ups
?“The interviewer says storage is cheap. How do you respond?”Reveal
Agree with the direction, then split storage into hot bytes, cold bytes, requests, retrieval, lifecycle transitions and egress. Storage can be cheap per GB-month while old data still costs money to retrieve, replicate or send to users. If retention dominates, lifecycle rules and deletion policy are design decisions, not billing afterthoughts.
?“Should we reserve everything to reduce cost?”Reveal
No. Reserve or commit only the stable baseline after the workload shape is known. Keep launch spikes and uncertain growth on-demand. Use interruptible capacity only for work that can checkpoint and retry. That mix reduces waste without risking the request path.
?“A multi-region active-active design is expensive. Is cheaper single-region acceptable?”Reveal
It depends on the availability and data-loss requirement. If the product can tolerate regional failover time, a single active region with tested backup and warm standby may be the right trade-off. If the requirement is low recovery time and global latency, pay for multi-region and explain the cost. Cost does not override requirements; it tests whether the requirement is real.
?“How do you protect a free tier from one expensive user?”Reveal
Calculate a per-user envelope, then enforce quotas on the dimensions that can grow: stored bytes, transformed bytes, public egress and request rate. A free user should not be able to create unbounded storage or bandwidth cost. The limits belong in product design and abuse controls, not only in finance reporting.
Code examples
capacity_input -> billable_unit -> unit_rate -> monthly_line
compute -> instance_hours -> illustrative $/instance-hour -> total
cache -> provisioned_GB_hours -> illustrative $/GB-hour -> total
storage -> GB_months_by_tier -> illustrative $/GB-month -> total
requests -> request_units -> illustrative $/million -> total
egress -> GB_by_boundary -> illustrative $/GB -> total
Sort monthly_line descending, then choose the design lever.Common mistakes
Many diagrams make app servers look important because they are visible. A ranked ledger may show that CDN misses, internet egress or cache RAM costs more. Sort the bill before choosing the lever.
A service call can be cheap inside one host and expensive when it crosses a zone, region or public internet boundary. Price the path that bytes actually travel, including backup, replication and analytics exports.
Archive tiers reduce storage cost when access is rare. They can add retrieval fees, restore delay and minimum-duration constraints. Put only data with a matching product promise into archive.
A cache with enough throughput can still be the largest bill if the hot set is big and replicated. Model payload size, overhead, headroom and copies before choosing the cache tier.
Free does not mean costless. Attach quotas to stored bytes, generated derivatives, public downloads and expensive requests. Otherwise one free account can consume the margin of many paying accounts.
Practice drills
Numbers in this section are illustrative.
A design has high QPS but tiny payloads. What cost line do you check first?Reveal
Check request-driven costs and compute first, but do not stop there. If each request reads a cache object with many replicas, memory may still dominate. If responses are tiny and cache objects are small, egress is probably not the leading line.
A storage tier is half the illustrative GB-month price of hot storage. When might it still be more expensive?Reveal
When the data is still read often, retrieval and request costs can erase the storage saving. Minimum storage duration can also charge for data deleted too soon. Move data only when access probability and restore expectations match the tier.
Why is a cache tier often a memory bill rather than a CPU bill?Reveal
Simple cache operations are usually cheap compared with the amount of RAM kept online with overhead, headroom and replicas. If the cached objects are too large or the hot set estimate is too broad, you pay for memory even when QPS is manageable.
A free user stores many large files and shares them publicly. Which envelope limits matter?Reveal
Limit stored bytes, generated derivatives, public egress and request rate. If uploads trigger transformations, also cap transformation minutes or queued jobs. The point is to prevent one free account from creating unlimited storage and bandwidth cost.
What part of a fleet belongs on spot or interruptible capacity?Reveal
Use it for retryable, checkpointed or queued work such as batch jobs, reprocessing, analytics and test environments. Keep user-facing request capacity on more predictable models unless the service can tolerate workers disappearing without breaking the user promise.
Deep dives
Per-user cost envelope
A per-user cost envelope is the range of cost one account can create. It is more useful than a global average because averages hide the users who upload, transform, export or serve much more data than the median.
Start with the behaviours the product allows. A reader may create mostly request and egress cost. A creator may create storage, transcoding, thumbnails, search indexing, safety scanning and public downloads. A team account may create shared storage plus notification fan-out. Model each as usage × illustrative unit rate, then decide which dimensions need quotas or paid-plan gates.
The envelope is also a fairness tool. A free plan can be generous on cheap dimensions and strict on expensive ones. For example, allowing many text notes may be cheap, while allowing unlimited public video downloads is usually not. The design answer is not only "charge money." It can be compression, lifecycle expiry, per-day transformation budgets, signed URLs, rate limits and abuse detection.
In an interview, say the envelope out loud when the prompt has a free tier, user-generated content, AI inference, file hosting or public sharing. It shows that you understand cost as a product constraint, not just an infrastructure bill.
Questions to answer
- Which user action creates the largest marginal cost?
- Can one account amplify cost for others, such as public links or fan-out?
- Which dimensions should be free, quota-bound or paid?
- What data expires automatically?
- What expensive work can be queued, cached or reused?
Strong interview phrasing
"I would model free-tier users by cost envelope, not just by average. Quiet readers are cheap; public uploaders can create storage, transformation and egress cost. I will cap those dimensions and let cheap dimensions stay generous."
Cheat sheet
- •Start from capacity estimation; do not invent a second scale model.
- •Ledger rows: compute hours, cache memory, storage GB-months, requests and egress GB.
- •Sort the ledger before optimising.
- •Egress is a boundary cost: internet, region, zone, backup and export paths.
- •Cache cost is mostly RAM: payload × overhead × replicas.
- •Storage tiers need an access story: hot, infrequent, archive and delete.
- •Cost per request includes compute, datastore calls, storage reads and bytes out.
- •Free-tier envelope needs quotas on storage, transformations and egress.
- •Commit the stable baseline; keep burst flexible; use spot only for retryable work.
- •The cheaper design wins when it still meets the requirement.
Practice this skill
No problem is tagged directly to Cost modelling yet. These published problems still exercise the same interview category.
Read this if