Design 01 · System Design

Designing a URL Shortener

A worked walkthrough from requirements to trade-offs, where the architecture evolves as the design does.

Cloud-agnostic first, with a possible Azure mapping at the end. The design is a worked example, and all capacity figures are labelled assumptions rather than measurements.

~9 min read

Start by deciding what the system must do, and how well it must do it. For a URL shortener the functional core is small. The non-functional requirements are what shape the architecture.

Functional

  • Create a short link for a long URL
  • Redirect a short link to its original URL
  • Optionally accept a custom alias
  • Optionally expire a link

Non-functional

  • Redirects are fast and highly available
  • The workload is read-heavy
  • Short codes are unique
  • Mappings are durable
  • Creation can tolerate slightly higher latency than redirects
Out of scope for this walkthrough: user accounts, analytics and link-editing interfaces.
The requirement that shapes the design: redirects far outnumber creations, so the read path is where most scaling effort will go.

Begin with the simplest structure that could work: a stateless application service behind a load balancer, backed by one database. It is easy to reason about, and it has a clear path to scale.

DESIGN v1 / BASELINEClientLoad balancerApp serviceDatabase
new in this version

This baseline cannot yet distinguish read traffic from write traffic, says nothing about how short codes are produced, and sends every redirect to a single database. Each of those gaps drives a later step.

Separating the read path from the write path lets each be scaled and protected independently. Creation is lower volume. Redirect is latency-critical.

DESIGN v2 / COMPONENTSClientLoad balancerCreate serviceRedirect serviceKey generatorDatabase
new in this version
ComponentResponsibilityDesign note
Load balancerDistribute requests, check healthStateless services allow horizontal scaling
Create serviceValidate the URL, obtain a key, store the mappingWrite path, lower volume
Redirect serviceLook up the code and return the redirectRead path, latency-critical
Key generatorProduce unique short codesOptions compared below
DatabaseDurable mapping storeUnique constraint on the code

Choosing how codes are generated

ApproachHow it worksStrengthWeakness
Hash the URLHash and truncateSame URL gives the same codeTruncation raises collision odds, which must be handled
Random codeGenerate a random base62 code, insert if unusedHard to guessCollision retries and an extra check on write
Counter encoded in base62Convert an increasing ID to base62No collisions, compactSequential codes are guessable; needs coordination

For this worked example, each service instance is allocated a range of IDs and encodes them in base62. That avoids collision checks and coordination on every write. Predictability is the cost, and the security step returns to it.

Access is dominated by one pattern: look up a long URL by its short code. That fits a table keyed by the code, or any key-value store.

short_links
  short_code   VARCHAR(10)   PRIMARY KEY
  long_url     TEXT          NOT NULL
  created_at   TIMESTAMP     NOT NULL
  expires_at   TIMESTAMP     NULL

Access patterns

  • Read by short code (dominant)
  • Insert a new mapping
  • Remove or ignore expired mappings

Consistency needs

  • Codes must be unique at creation
  • Reads can usually tolerate slight staleness
  • A just-created link should resolve promptly

A uniqueness constraint on the code also protects custom aliases: a clash surfaces as a conflict rather than silently overwriting an existing link.

The public contract is small: create a link and follow a link. Making creation safely repeatable matters, because clients and networks retry.

POST /api/v1/links
Idempotency-Key: <client-generated key>
{
  "longUrl": "https://example.com/some/long/path",
  "customAlias": "docs",            // optional
  "expiresAt": "2030-01-01T00:00:00Z"   // optional
}

201 Created
{ "code": "abc1234", "shortUrl": "https://short.example/abc1234" }
GET /abc1234

302 Found
Location: https://example.com/some/long/path
StatusMeaning
201Link created
302Redirect to the stored destination
400Invalid URL or request
404Unknown code
409Custom alias already taken
410Link expired
429Client is being rate limited
301 or 302? A permanent redirect (301) can be cached by browsers, which reduces traffic but makes it hard to expire or change a link. A temporary redirect (302) keeps control with the service at the cost of more requests. This design uses 302 and accepts that cost.

Capacity: a worked example

Worked example, not real data. The inputs below are assumptions chosen for illustration. The results are arithmetic from those assumptions, not measurements from any real system.
Assumption (input)Value
New links per month100 million
Read to write ratio100 : 1
Peak to average traffic5 : 1
Average record size500 bytes
Retention5 years
Derived (arithmetic)Result
Average writes per secondabout 39
Average redirects per secondabout 3,900
Peak redirects per secondabout 19,000
Records after 5 years6 billion
Storage for recordsabout 3 TB
Short codes available at 7 base62 charactersabout 3.5 trillion

The numbers indicate shape rather than precision. Writes are modest, redirects are thousands per second at peak, and storage is measured in terabytes. One database serving every redirect will be the first thing to strain.

Step up: add a cache

DESIGN v3 / + CACHEClientLoad balancerCreate serviceRedirect serviceKey generatorCacheDatabase
new in this version

Redirect traffic is likely to be concentrated on a small set of popular links. That is an assumption to validate, but it is what makes caching effective.

The redirect service uses a cache-aside pattern: check the cache, and on a miss read the database and populate the cache. Capping each entry's time to live by the link's expiry keeps expired links from being served.

GivesCosts
Lower latency and far less database load on popular linksStale results after a link is disabled or changed
A buffer in front of the databaseCold-start load after restarts; stampedes on a hot key unless requests are coalesced
Cheap horizontal read capacityMemory cost and another component to operate

Step up: partition the data

DESIGN v4 / + PARTITIONINGClientLoad balancerCreate serviceRedirect serviceKey generatorCacheShard 1+ replicasShard 2+ replicasShard N+ replicaspartition by hash(code)
new in this version

Replicas add read capacity. When data or write volume outgrows a single node, the table is partitioned. Hashing the short code spreads keys evenly and suits key lookups, and consistent hashing limits how much data moves when shards are added.

GivesCosts
Capacity beyond one nodeQueries across shards, such as all links for one owner, need another index or store
Even distribution of keysResharding is operational work; a very hot key can still overload one shard
Independent failure domainsMore moving parts to monitor and recover

Design for each component failing, and decide in advance which capability matters most. Here, redirects matter more than creation, so the design degrades by protecting redirects.

FailureEffectMitigation
Application instance failsCapacity dropsStateless services, load balancer health checks, instances in multiple zones
Cache unavailableMore database load and higher latencyFall back to the database, cap concurrent lookups, coalesce requests for hot keys
Shard or primary unavailableCreates and uncached redirects affectedReplication and failover; cached redirects continue to be served
Key generator unavailableCreation could failPre-allocated ID ranges let instances keep creating for a time
Duplicate create requestsDuplicate linksIdempotency key stored with the outcome
Slow dependencyRequests and threads pile upTimeouts, bounded retries with backoff and jitter, circuit breaking

Durability is handled separately: replicated storage and tested backups protect the mappings, which are the system's real asset.

A public redirect service is an attractive tool for abuse. Security here is mostly about limiting what the service can be used to do.

ThreatMitigation
Links to phishing or malwareAccept only http and https, optionally check destinations against reputation or block lists, and provide a way to report and disable links
Code enumerationSequential codes are easy to scan; scramble IDs or use random codes where guessing matters. Do not rely on an unguessable URL to protect private data
Resource abuseRate limit creation per client or API key, and apply protective limits at the edge
Open redirect misuseRedirect only to destinations stored at creation, never to a destination supplied in the redirect request
Transport and accessTLS everywhere, authenticated creation, least-privilege database credentials, secrets kept in a vault
AccountabilityAudit creation and disabling events

Good design is a record of choices and what each one cost. These are the main ones in this walkthrough.

DecisionChosen hereWhat it costs
Redirect status302 temporary redirectMore traffic reaches the service than with 301, in exchange for control over expiry and change
Key generationID ranges encoded in base62Predictable codes unless scrambled; a range allocator to operate
CachingCache-aside with expiry-aware TTLPossible stale results; stampede handling
PartitioningHash of the short codeNo range queries; resharding effort
ConsistencyReplicas may lag for redirectsA new link may briefly not resolve on a lagging replica; read from the primary or cache on create
Data modelKey-value style accessLimited ad-hoc querying

If requirements change, for example link analytics or per-user management, several of these choices would be revisited, starting with the data model and the partitioning scheme.

APPENDIX

Possible Azure Mapping

The design above is cloud-agnostic. This maps each component to Azure services that could implement it.

Architectural componentPossible Azure servicesNote
Edge and load balancingAzure Front Door, Azure Application GatewayTerminates TLS and routes traffic
Create and redirect servicesAzure Container Apps, App Service, Azure Kubernetes ServiceStateless, horizontally scaled
CacheAzure Cache for RedisCache-aside for redirects
Partitioned datastoreAzure Cosmos DB (partition key: short code), or Azure SQL DatabaseKey-based lookups
Key range allocatorA small service backed by durable storageDepends on the chosen store
Rate limiting and API managementAzure API ManagementPolicies per client or key
SecretsAzure Key VaultCredentials outside source control
ObservabilityAzure Monitor, Application InsightsMetrics, traces, logs
Design exercise only. This is a possible mapping, not a description of a system I have deployed.

Next in the design library

Rate Limiter, Notification System, Event-Driven Order Processing, Distributed Cache and a Document Processing / Knowledge Platform are coming soon.