System Design
Designing Systems That Scale, Adapt and Stay Reliable
Exploring system design through requirements, architecture decisions, failure modes and explicit trade-offs — from first principles to production-ready patterns.
How I Approach System Design
One consistent method, from the problem to the trade-offs.
What must it do, and how well?
What is the simplest structure that meets them?
What are the responsibilities and boundaries?
What is stored, how is it accessed, and what must be consistent?
What contract do clients and services depend on?
What breaks first as load grows?
How does it behave when parts fail?
Who can do what, and how is abuse limited?
What did we choose, and what did we give up?
URL Shortener
A worked walkthrough in which the architecture evolves step by step, with capacity assumptions clearly labelled as examples, and a possible Azure mapping kept separate from the cloud-agnostic design.
Open Design →Selected stages from the full walkthrough.
Design Library
Each design follows the same nine-step method.
URL Shortener
How do you turn long URLs into short, fast, durable links at scale?
Open Design →Rate Limiter
How do you limit request rates fairly across many servers?
Coming SoonNotification System
How do you deliver events across several channels reliably?
Coming SoonEvent-Driven Order Processing
How do you process orders asynchronously without losing or duplicating work?
Coming SoonDistributed Cache
How do you spread cached data across nodes and keep it useful?
Coming SoonDocument Processing / Knowledge Platform
How do you ingest, process and index documents for search and retrieval?
Coming SoonArchitecture Toolkit
Every building block is a trade: what it gives you, and what it costs.
Designing for Failure
In distributed systems, failure is normal. These patterns exist because of how failures spread, and each one has a price.
Timeouts
Why it existsA slow dependency can hold resources and stall callers indefinitely.
The trade-offToo short causes false failures; too long ties up capacity.
Retries
Why it existsMany failures are transient.
The trade-offRetries multiply load on a struggling dependency, so they need backoff, jitter and limits.
Idempotency
Why it existsRetries and duplicates are inevitable, so repeating a request must not repeat its effect.
The trade-offIt needs idempotency keys and stored outcomes.
Circuit breaking
Why it existsCalling a failing dependency makes things worse; stopping lets it recover and callers fail fast.
The trade-offIt can reject requests that would have succeeded, and thresholds need tuning.
Backpressure
Why it existsProducers faster than consumers exhaust queues and memory.
The trade-offSomething must be slowed, rejected or shed, and users notice.
Graceful degradation
Why it existsA partial service is better than a total outage.
The trade-offIt needs defined fallbacks, and reduced behaviour must be acceptable and visible.
Good architecture is a record of trade-offs.
Explore how I reason about systems, from requirements to the choices that shape them.