Software Architecture

Seven Backend Patterns for Scalable Systems — and When Not to Use Them

12 min readRuhi GargaGuideIntermediate

Every backend pattern fixes one problem and creates another. Knowing which problem you actually have matters more than knowing the pattern’s name.

A pattern is a recorded trade-off

A pattern is useful because someone has already worked out what it fixes and what it costs. It becomes harmful when it is adopted for its name. For each of the seven below, this guide covers what it solves, what it adds to build and operate, and when to leave it out.

The sketches are in C#, but the ideas are not tied to .NET. Where a .NET library is named, it is one implementation of the pattern and not the pattern itself.

Three kinds of pattern

Strictly, none of these seven is a design pattern in the Gang of Four sense, which describes structure inside a program. Using the right word helps you ask the right question.

  • Design pattern. A solution to a recurring problem inside a program, such as Strategy or Decorator.
  • Architectural pattern. Shapes how a system is divided into parts and how they communicate.
  • Integration or resilience pattern. Governs how components exchange data, or how they survive each other’s failures.
  • Implementation technique. A specific mechanism that realises a pattern, such as an expiry on a cache entry, an idempotency key or a polling relay.
PatternKind
Cache-asideImplementation technique
Queue-based load levellingIntegration pattern
CQRSArchitectural pattern
Transactional outbox, idempotent consumerIntegration pattern, with an implementation technique
SagaIntegration pattern
Circuit breaker, with timeouts and retriesResilience pattern, with implementation techniques
Strangler figArchitectural (migration) pattern

The boundaries are fuzzy, and catalogues classify some of these differently.

1. Cache-aside

Implementation technique Caching strategy

The problem. The same data is read far more often than it changes, and every read pays for a trip to the data store.

How it works. The application checks the cache first. On a hit it returns the cached value. On a miss it reads from the data store, saves the result in the cache with an expiry, and returns it. Writes go to the data store, and the cached entry is removed or left to expire.

Cache hitRequest→Cache→Response
Cache missRequest→Cache→Data store→Fill cache→Response

Cache-aside: the application checks the cache, and fills it from the data store on a miss.

Use it when

  • Reads far outnumber writes.
  • Some staleness is acceptable.
  • The data is costly to fetch and keys are easy to define.

Avoid it when

  • The data must always be exact, such as a balance or stock at checkout.
  • Few reads repeat, so the hit rate would be low and the cache adds a hop for nothing.
  • Invalidation rules are unclear.
  • You haven’t measured and the data store may not be the bottleneck.

GivesLower latency and less load on the data store. It can be adopted one query at a time.

CostsStale reads, invalidation logic, cold-start load, stampedes when many callers miss the same key, and another system to run if the cache is distributed.

public sealed class ProductReader(HybridCache cache, ProductDb db)
{
    public ValueTask<ProductDto?> GetAsync(int id, CancellationToken ct) =>
        cache.GetOrCreateAsync(
            $"product:{id}",
            async token => await db.Products
                .Where(p => p.Id == id)
                .Select(p => new ProductDto(p.Id, p.Name, p.Price))
                .FirstOrDefaultAsync(token),
            new HybridCacheEntryOptions { Expiration = TimeSpan.FromMinutes(5) },
            cancellationToken: ct);

    public async Task UpdateAsync(Product product, CancellationToken ct)
    {
        db.Update(product);
        await db.SaveChangesAsync(ct);
        await cache.RemoveAsync($"product:{product.Id}", ct); // invalidate on write
    }
}

HybridCache (package Microsoft.Extensions.Caching.Hybrid, generally available from ASP.NET Core 9) performs the check-and-fill steps and ensures only one concurrent caller runs the factory for a given key, which guards against stampedes. Other caches follow the same steps.

2. Queue-based load levelling

Integration pattern With competing consumers

The problem. Traffic arrives in bursts, but the work behind it can only run at a steady rate. Without a buffer, a burst lands directly on the slower component, which then times out or fails.

How it works. A queue sits between the producer and the worker. The producer adds a message and returns at once. Consumers take messages at the rate they can sustain. More consumers reading the same queue (competing consumers) raise throughput without changing the producer.

Burst in, steady rate outBursty requests→Queue→Consumers at a steady rate→Downstream

The queue absorbs the burst. Consumers work at a rate the downstream can sustain.

Use it when

  • Spikes are short compared with average load.
  • The caller doesn’t need the result in the same request.
  • The downstream has a hard capacity limit.

Avoid it when

  • The user needs the result immediately.
  • Load is sustained, so the queue would only hide an overload while its delay grows without bound.
  • You can’t yet monitor queue depth or handle messages that keep failing.
  • Strict ordering across consumers is required.

GivesAbsorbs spikes, lets producer and consumer fail independently, and lets consumers scale on their own.

CostsAsynchronous behaviour that reaches the user interface, duplicates and ordering to handle, a queue to monitor, and a policy for when it fills.

// Producer: accept quickly, enqueue, answer "accepted"
app.MapPost("/reports", (ReportRequest request, Channel<ReportRequest> queue) =>
    queue.Writer.TryWrite(request)
        ? Results.Accepted()
        : Results.StatusCode(503));        // bounded queue is full: shed load

// Consumer: drain at a rate the downstream can sustain
public sealed class ReportWorker(Channel<ReportRequest> queue, ReportService reports) : BackgroundService
{
    protected override async Task ExecuteAsync(CancellationToken ct)
    {
        await foreach (var request in queue.Reader.ReadAllAsync(ct))
            await reports.GenerateAsync(request, ct);
    }
}

An in-process channel loses its contents if the process stops. For work that must survive a restart, use a durable broker. The pattern is the same.

3. CQRS

Architectural pattern Command Query Responsibility Segregation

The problem. The model that suits changing data, with validation and consistency rules, is often awkward for reading it, which needs joins and a different shape per screen. One model for both gets compromised.

How it works. Separate the operations that change state (commands) from those that read it (queries), and give each its own model. In its simplest form, both still use the same database. Separate read stores, kept in step by projections, are an option for when the read side needs different storage or scale. CQRS does not require separate databases, and it does not require event sourcing.

CQRS: separate command and query modelsCommands go through a command model to a write store. Queries go through a query model to a read store. In the simple form both stores are the same database.WRITE SIDEREAD SIDECommandsCommand modelvalidates, changes stateWrite storeQueriesQuery modelshaped for readingRead storeSimple form: one shared database.Optional: project changes intoa separate read store.

CQRS in its simple and optional forms.

Use it when

  • Read and write shapes differ noticeably.
  • Writes carry complex rules while reads are simple but varied.

Avoid it when

  • The application is mostly CRUD with matching shapes.
  • The team is small and the extra structure outweighs the benefit.
  • Users expect to see their own write immediately, and you plan separate read stores that lag.
  • The motivation is that it is fashionable.

GivesModels fitted to their purpose, reads that can be tuned independently, and simpler query code.

CostsMore types and code. With separate read stores, also sync lag, rebuild tooling and a second store to run. Start with the simple form.

public sealed class PlaceOrderHandler(OrdersDb db)
{
    public async Task HandleAsync(PlaceOrder command, CancellationToken ct)
    {
        var order = Order.Create(command.CustomerId, command.Lines);   // domain rules live here
        db.Orders.Add(order);
        await db.SaveChangesAsync(ct);
    }
}

public sealed class OrderSummaryHandler(OrdersDb db)
{
    public Task<OrderSummaryDto?> HandleAsync(GetOrderSummary query, CancellationToken ct) =>
        db.Orders.AsNoTracking()                                       // same database, read-shaped model
            .Where(o => o.Id == query.OrderId)
            .Select(o => new OrderSummaryDto(o.Id, o.Status, o.Lines.Count, o.Total))
            .FirstOrDefaultAsync(ct);
}

4. Transactional outbox and idempotent consumer

Integration pattern With an idempotent consumer as an implementation technique

The problem. A service saves a change and then publishes an event. If it saves and crashes before publishing, the event is lost. If it publishes first and the save fails, the event describes something that never happened. The database and the broker cannot share one transaction.

How it works. In the same transaction that saves the change, write the event to an outbox table. A separate relay reads unsent rows, publishes them and marks them as sent. The relay can publish a message and fail before marking it, so delivery is at least once and duplicates are possible.

That is why an idempotent consumer is the usual companion: it records the IDs of messages it has handled and ignores repeats. Neither needs the other. Naturally idempotent handlers can sit behind an outbox, and an idempotent consumer helps behind any at-least-once source.

Transactional outbox with an idempotent consumerA service writes its data change and an outbox row in one database transaction. A relay publishes unsent rows to a broker and marks them sent. The consumer checks whether it has seen the message ID before applying its effect.SERVICE DATABASE: ONE TRANSACTIONData tablethe changeOutbox tablethe eventServicesave change + eventRelaypublish, then mark sentBrokerCONSUMERSeen this ID?Apply effectThe relay can publish and fail before marking the row, so delivery is at least once.The consumer therefore ignores message IDs it has already handled.

Outbox and idempotent consumer. Each protects against a different failure.

Use it when

  • A state change must always be accompanied by an event.
  • You already have a transactional store.

Avoid it when

  • No event is published, or loss is tolerable.
  • The store can’t cover both writes in one transaction.
  • A change feed from the database already provides this. Don’t build both.
  • Polling the outbox would put real load on your database at your scale.

GivesNo lost or invented events, and a simple way to reason about publishing.

CostsAn extra table and relay to run, a delay before publishing, duplicates to handle, cleanup of old rows and care with ordering.

// Publisher side: one transaction writes the change and the event
db.Orders.Add(order);
db.Outbox.Add(new OutboxMessage
{
    Id = Guid.NewGuid(),
    Type = nameof(OrderPlaced),
    Payload = JsonSerializer.Serialize(new OrderPlaced(order.Id, order.Total)),
    CreatedAt = DateTimeOffset.UtcNow
});
await db.SaveChangesAsync(ct);

// Consumer side: ignore messages that were already handled
if (await db.Processed.AnyAsync(p => p.MessageId == messageId, ct)) return;
await ApplyEffectAsync(message, ct);
db.Processed.Add(new ProcessedMessage(messageId));  // unique index on MessageId guards races
await db.SaveChangesAsync(ct);

Some messaging libraries include outbox support. The sketch shows the mechanism only and is not a recommendation of any library.

5. Saga

Integration pattern Orchestration or choreography

The problem. A business operation spans several services, each with its own data. One database transaction cannot cover them, so when step three fails, steps one and two have already committed.

How it works. Split the operation into local transactions, one per service. If a step fails, run compensating actions for the steps that completed, such as releasing the stock or refunding the payment. A compensation is a new action and not a rollback. Some actions, such as an email already sent, cannot be undone, and other readers can see the intermediate state.

In orchestration, a coordinator tells each service what to do next, which is easier to follow. In choreography, each service reacts to the others’ events, which has fewer parts until the flow grows.

Saga with compensating actionsReserve stock, take payment and arrange shipping are local transactions in sequence. If shipping fails, the saga refunds the payment and releases the stock in reverse order.FORWARD: EACH STEP IS A LOCAL TRANSACTION IN ONE SERVICEReserve stockTake paymentArrange shippingfails in this exampleCOMPENSATION: UNDO IN REVERSE ORDERRelease stockRefund paymentstep fails: start compensating

A saga fails forward and unwinds by compensation. It does not roll back.

Use it when

  • The workflow crosses services that own separate data.
  • Each step can be undone, or the business accepts a compensation.
  • Temporary inconsistency is acceptable.

Avoid it when

  • Everything lives in one database. A local transaction is simpler and stronger.
  • A step can’t be compensated and strict atomicity is required. Rethink the boundaries instead.
  • Readers must never see intermediate states.

GivesWorkflows across services without holding locks across them, and failure handling that is explicit.

CostsCompensation logic, state to track, intermediate states visible to others, harder testing, and steps that must be idempotent.

Sagas and outboxes often appear together, because each step must update its state and send the next message reliably. A saga does not require an outbox.

public enum OrderSagaState { Started, StockReserved, PaymentTaken, Completed, Compensating, Failed }

// The orchestrator reacts to each event and decides what happens next
switch (@event)
{
    case StockReserved:
        saga.State = OrderSagaState.StockReserved;
        await bus.SendAsync(new TakePayment(saga.OrderId));
        break;
    case ShippingFailed when saga.State == OrderSagaState.PaymentTaken:
        saga.State = OrderSagaState.Compensating;
        await bus.SendAsync(new RefundPayment(saga.OrderId));   // compensating action
        break;
}

6. Circuit breaker, with timeouts and retries

Resilience pattern Timeouts and retries are implementation techniques

The problem. A slow or failing dependency holds its callers’ resources, and callers that keep retrying add load to something already struggling. One bad dependency can spread failure across the system.

Three tools, three jobs. They are often grouped, but none replaces another.

  • Timeout. Limits how long you wait for one call, so a slow dependency cannot hold your resources.
  • Retry. Repeats a call after a transient failure. It needs backoff, jitter and a limit, and it is only safe for operations that can be repeated.
  • Circuit breaker. Stops calling a dependency that keeps failing, so it can recover and callers fail fast.

Used together they form layers. Retries without a breaker keep pushing on a struggling service, and timeouts alone do not stop a pile-up.

Circuit breaker statesClosed passes calls and counts failures. When failures pass a threshold it opens and fails calls immediately. After a cool-down it goes half-open and allows a few trial calls. Success closes it, and a failure opens it again.Closedcalls pass, failures countedOpencalls fail immediatelyHalf-opena few trial callsfailure rate passes thresholdcool-down endstrial calls succeeda trial call fails

The three breaker states. Timeouts and retries sit around the breaker, not inside it.

Use it when

  • Calls cross a network to a dependency that can fail or slow down.
  • Callers can fail fast or have a fallback.

Avoid it when

  • The call is in-process.
  • Failing fast is worse than waiting and there is no fallback, as in some batch jobs.
  • You can’t tune thresholds from real traffic. A poorly tuned breaker rejects healthy calls or never opens.
  • You would retry operations that aren’t safe to repeat.

GivesProtects callers and the dependency, and failure that arrives quickly instead of slowly.

CostsThresholds to tune, calls rejected that would have succeeded, a fallback to design, and breaker state to monitor.

builder.Services
    .AddHttpClient<PaymentsClient>(c => c.BaseAddress = new Uri("https://payments.internal"))
    .AddStandardResilienceHandler(options =>
    {
        // The defaults retry every HTTP method. Don't retry calls that aren't safe to repeat.
        options.Retry.DisableForUnsafeHttpMethods();
    });

Microsoft.Extensions.Http.Resilience, which builds on Polly, provides a standard handler that layers a rate limiter, a total timeout, retries, a circuit breaker and a per-attempt timeout, each with documented defaults. Treat those defaults as a starting point and check them against your version. They retry every HTTP method unless you change that.

7. Strangler fig

Architectural (migration) pattern About modernising, not scaling

The problem. A legacy system is hard to change, but replacing it in one go is risky and means a long period with nothing shipped. This pattern does not make a running system faster or more scalable. It makes replacing one safer.

How it works. Put a façade, usually a proxy, in front of the legacy system. At first it routes everything to the legacy system. Then, one capability at a time, build the replacement and move the matching routes to it. The legacy system shrinks until it can be switched off, and then the façade can go too.

Phase 1Client→Façade→Legacy system
Phase 2Client→Façade→Some routes: new system→Rest: legacy
Phase 3Client→New system

Routes move one at a time. When the legacy system is retired, the façade can be removed.

Use it when

  • The system is too large or risky to replace at once.
  • You can intercept requests at a boundary such as HTTP.
  • The legacy system can stay for a long period.

Avoid it when

  • You can’t place a façade in front of the system.
  • The system is small enough to replace outright.
  • Both systems need to share data without a plan. Shared data is the hard part.
  • You need to decommission the original quickly.

GivesSmaller releases, risk spread over time, and rollback one route at a time.

CostsTwo systems running for longer, a façade that must stay reliable and not become a bottleneck, data to share or sync, and a risk that the migration never finishes.

{
  "ReverseProxy": {
    "Routes": [
      { "RouteId": "invoices", "ClusterId": "new-system",
        "Match": { "Path": "/invoices/{**catch-all}" } },
      { "RouteId": "everything-else", "ClusterId": "legacy",
        "Match": { "Path": "{**catch-all}" } }
    ],
    "Clusters": {
      "new-system": { "Destinations": { "d1": { "Address": "https://new.internal/" } } },
      "legacy":     { "Destinations": { "d1": { "Address": "https://legacy.internal/" } } }
    }
  }
}

YARP is a reverse proxy library for .NET. Any proxy or gateway that can route by path serves the same purpose. The configuration shape above follows YARP’s documentation, so check it against the version you use.

Choosing between them

SymptomConsiderCheck first
Reads are slow and repetitiveCache-asideIs the data store really the bottleneck?
Spikes overwhelm a downstream componentQueue-based load levellingCan the caller accept “accepted, not done”?
Reads and writes need different shapesCQRS, in its simple form firstDo the two models actually diverge?
Events are lost or invented after a saveTransactional outboxIs the store transactional, and is there already a change feed?
A workflow spans services and a step can failSagaCould the data live in one place instead?
One slow dependency drags its callers downTimeouts, retries, circuit breakerAre the calls safe to repeat?
A legacy system must be replaced without a big bangStrangler figCan you intercept the requests?

Patterns combine, but none requires another. A saga step may use an outbox. A queue consumer may call a dependency guarded by a circuit breaker. Add a combination only when each part has its own reason.

Adopting patterns gradually

I’d treat each pattern as a response to a symptom you can observe, not as a starting point. Measure first, adopt the smallest form that helps, such as CQRS over one database, and write down what you gave up. Each pattern also adds something to watch: hit rate, queue depth, outbox lag, saga states, breaker state.

Two related areas on this site go deeper. The System Design page covers designing for failure and the architecture toolkit, and the URL Shortener walkthrough shows caching and partitioning in a worked design.

Key takeaways

  1. A pattern is a trade-off. Adopt it for a problem you can observe, not for its name.
  2. Know the kind: architectural, integration, resilience or implementation technique.
  3. CQRS does not need separate databases, and a saga is not a general replacement for distributed transactions.
  4. Timeouts, retries and circuit breakers do different jobs and work as layers.
  5. Strangler fig is about safe modernisation, and it does not improve runtime scalability.

References

← Back to all Insights