Building Software

Engineering Fundamentals for the Agent Era

Contents Section 5, State

Transactions and Consistency

Mistakes to catch in review

  1. A read-modify-write on a balance or counter with no transaction or lock, losing updates under concurrent requests.

  2. A multi-step change where a failure halfway through leaves partial data behind.

  3. A read from a replica right after a write to the primary, so users don't see their own change.

  4. An assumption that the database's default isolation level prevents double-booking when it does not.

Making multi-step changes all-or-nothing, and understanding what concurrent readers and writers can see.

Topics

Atomicity
All-or-nothing changes, and what a transaction does and does not protect against.
Isolation Levels and Anomalies
Dirty reads, lost updates, write skew and phantoms, and which isolation levels prevent which.
Locking and Optimistic Concurrency
Row locks, version numbers and compare-and-set as ways to make concurrent writers safe.
Consistency Models
Strong, eventual and read-your-writes consistency, and what each one promises a user.
Replication and Replica Lag
How copies of data are kept in sync, and what stale reads look like to users.

You understand it when you can

  • Reproduce a lost update with two concurrent sessions and fix it in two different ways.
  • Name the isolation anomaly behind a given bug and the isolation level or technique that prevents it.
  • Explain what read-your-writes consistency means and how to provide it with a replicated database.

Drill

An agent wrote a gift-card redemption endpoint that reads the balance, checks that it covers the purchase, and writes the new balance in three separate statements. Find the concurrent sequence that spends the same balance twice, and write the test that reproduces it.

Start here

Watch

Distributed Systems 7.2: Linearizability

Martin Kleppmann, 2020. 19-minute lecture.

Defines strong consistency as linearizability and shows how a read from a lagging replica after a write breaks it, which is the stale-read bug behind replica reads.

Read

Database Internals: A Deep Dive into How Distributed Data Systems Work

Alex Petrov, 2019.

Explains transaction processing, recovery logs, locking and optimistic concurrency control, then replication and consistency models, from the storage engine up.

Primary sources

  • Manual

    PostgreSQL Documentation: Transaction Isolation

    The table of which anomalies each isolation level allows in PostgreSQL, and the serialization failures an application must retry under Repeatable Read and Serializable.

  • Reference

    Jepsen: Consistency Models

    A map of consistency models from serializable and linearizable down to read-your-writes, each with a short definition of what it promises.