Building Software

Engineering Fundamentals for the Agent Era

Contents Section 2, Verification

Tests as Executable Specifications

Mistakes to catch in review

  1. A failing test 'fixed' by changing its expected value to match the buggy output.

  2. Tests generated by reading the implementation, so they faithfully encode its bugs.

  3. Tests that assert implementation details, such as which private method ran or how many times a mock was called, and pass while the behavior is wrong.

  4. A test with no meaningful assertion, which passes no matter what the code does.

Test-first thinking without the ritual: every behavior has a test, each test comes from a requirement, and nothing is done until the tests say so.

Topics

Test-First, Reimagined
Every behavior gets a test written from the requirement, before or alongside the code, with red-green-refactor used as a tool when it helps.
Testing Observable Behavior
Asserting what callers and users can see, so tests survive refactoring and fail when something real breaks.
The Oracle Problem
Deciding who knows the right answer for a test, and why that answer must come from the specification instead of the code under test.
Regression Tests
Turning every bug into a test that fails before the fix and passes after it.
One Behavior per Test
Small, well-named tests whose failure message says exactly which expectation broke.

You understand it when you can

  • Write failing tests from a set of acceptance criteria before any implementation exists.
  • Review a test file and state, for each test, what broken behavior would make it fail.
  • Add a regression test that reproduces a reported bug before fixing the bug.

Drill

An agent fixed a failing test for a shipping-cost function by changing the expected value from 12.50 to 12.49, and its new tests check only that the function calls the rate lookup exactly once. Find what these tests would fail to catch, and write one test derived from the shipping requirement instead.

Start here

Watch

TDD, Where Did It All Go Wrong (Ian Cooper)

Ian Cooper, 2017. 64-minute talk.

Returns to Kent Beck's original rules to argue that tests should target behaviors behind a public API, not classes and methods, which is why implementation-coupled tests break on refactoring yet miss real bugs.

Watch

How to Write Acceptance Tests

Dave Farley, 2020. 15-minute explainer.

Explains how to write executable acceptance tests in the language of the requirement instead of the implementation, so the tests act as the specification an agent has to satisfy.

Read

Unit Testing Principles, Practices, and Patterns

Vladimir Khorikov, 2020.

Defines a good test as resistant to refactoring and protective against regressions, and shows why asserting on mock interactions and private details produces tests that pass while behavior is wrong.

Test Driven Development: By Example

Kent Beck, 2002.

The original red-green-refactor book: works through writing a failing test from a requirement before the code exists, in small steps you can reuse even when an agent writes the code.

Growing Object-Oriented Software, Guided by Tests

Steve Freeman and Nat Pryce, 2009.

Builds a real system outside-in from failing acceptance tests, showing how end-to-end tests derived from requirements drive and bound the unit-level work.

Primary sources

  • Reference

    Test Desiderata (Kent Beck)

    Kent Beck's list of properties tests should have, including behavioral, structure-insensitive, specific and deterministic, which works as a checklist for reviewing agent-written tests.

  • Reference

    Software Engineering at Google, Chapter 12: Unit Testing

    Google's guidance to test through public APIs, test behaviors instead of methods, keep logic out of tests and write failure messages that say exactly which expectation broke.