Building Software

Engineering Fundamentals for the Agent Era

Contents Section 5, State

Distributed Systems and Partial Failure

Mistakes to catch in review

  1. Retries without idempotency that charge a customer twice when the first attempt had actually succeeded.

  2. Missing timeouts, so one slow dependency ties up every request that touches it.

  3. Immediate retries without backoff or jitter that turn a brief outage into a retry storm.

  4. Events ordered by wall-clock timestamps taken on different machines.

When work spans machines, any step can fail, stall or happen twice, and the design has to assume it will.

Topics

The Fallacies of Distributed Computing
The false assumptions, such as a reliable network and zero latency, that break distributed designs.
Timeouts, Retries and Backoff
Bounding every wait, and retrying safely with exponential backoff and jitter.
Idempotency
Designing operations that give the same result when applied twice, using idempotency keys and deduplication.
Delivery Guarantees
At-most-once, at-least-once and effectively-once processing with queues and the outbox pattern.
Time and Ordering
Clock skew, monotonic clocks and logical clocks, and why wall-clock time cannot order events across machines.
Consensus and Coordination
Leader election, quorums and the CAP tradeoff, and when to rely on a proven coordination service.

You understand it when you can

  • Design an idempotent payment endpoint and explain what happens when a client retries after a timeout.
  • Explain why exactly-once delivery is impossible over an unreliable network and what systems do instead.
  • Set timeouts and retry policies for a chain of three services so that each caller's timeout covers the retries of the service it calls.

Drill

An agent wrote a checkout endpoint that retries the payment call on timeout, up to five times, with no delay between attempts. Find the path that charges the customer twice and what the retries do to the payment provider during an outage, then write the test that proves the double charge.

Start here

Watch

Distributed Systems 2.1: The two generals problem

Martin Kleppmann, 2020. 12-minute lecture.

Shows why a sender can never be certain its message was acted on over an unreliable network, the root of the timeout-then-double-charge problem and of why exactly-once delivery is impossible.

Distributed Systems 3.1: Physical time

Martin Kleppmann, 2020. 21-minute lecture.

Explains clock skew, NTP corrections and monotonic versus wall-clock time, and why timestamps from different machines cannot order events.

Read

Understanding Distributed Systems

Roberto Vitillo, 2022, 2nd edition.

A practitioner's short guide to timeouts, retries with backoff, idempotency, the outbox pattern, leader election and replication, written for application developers.

Release It!: Design and Deploy Production-Ready Software

Michael T. Nygard, 2018, 2nd edition.

Catalogs production failure modes such as missing timeouts, cascading failures and self-inflicted retry storms, and the stability patterns that stop them: timeouts, circuit breakers and bulkheads.

Primary sources

  • Specification

    Stripe API Reference: Idempotent requests

    The payment API's own contract for idempotency keys: the server stores the first result for a key and returns it on a retry, so a retry after a timeout never charges twice.