Building Software

Engineering Fundamentals for the Agent Era

Contents Section 7, Operations

Performance and Capacity

Mistakes to catch in review

  1. Code optimized where it is not the bottleneck, while the real cost sits in a query or a network call.

  2. A cache with no invalidation that serves stale data, or a cache key that serves one user's data to another.

  3. A service that works for ten users and falls over at a thousand because nobody considered the connection pool or a rate limit.

  4. Autoscaling or model usage with no ceiling, turning a traffic spike into a very large bill.

Measuring where time and money go, finding bottlenecks before users do, and planning for load.

Topics

Latency, Throughput and Tail Latency
The two basic performance measures, and why the slowest percentiles decide user experience.
Profiling
Measuring CPU, memory and I/O to find where time actually goes before optimizing anything.
Load Testing and Capacity Planning
Simulating realistic traffic to find limits, and planning resources ahead of growth.
Caching and Invalidation
Where caches help, how they go stale, and designing keys and expiry so they stay correct.
Cost as a Metric
Tracking compute, storage and model spend per feature, with budgets and alerts like any other signal.

You understand it when you can

  • Profile a slow endpoint and show where the time goes before changing any code.
  • Load test a service, find the point where it degrades, and name the bottleneck.
  • Explain what a cache key must include for a per-user cache to be safe.

Drill

An agent sped up a slow product page by memoizing a pricing function, and added a cache keyed only on the product ID for a response that includes the viewer's saved cart. Profile to find where the page actually spends its time, and find the cache bug that shows one user another user's cart.

Start here

Watch

"Performance Matters" by Emery Berger

Emery Berger, 2019. 42-minute talk.

Shows how layout effects and noise mislead naive benchmarks, and introduces causal profiling (Coz) to find the code whose speedup would actually make the program faster.

Watch

"How NOT to Measure Latency" by Gil Tene

Gil Tene, 2015. 43-minute talk.

Explains why averages and even p99 hide what users experience, and how coordinated omission makes most load-testing tools under-report tail latency.

Web Cache Deception Attack

Omer Gil, 2017. 23-minute talk.

Shows real sites whose caches stored one user's authenticated page and served it to others because the cache key and rules ignored who was asking, the per-user caching bug in this subsection's drill.

Read

Systems Performance: Enterprise and the Cloud

Brendan Gregg, 2020, 2nd edition.

The standard reference for performance methodology (USE method, workload characterization) and for profiling CPU, memory, disk and network before changing any code.

The Art of Capacity Planning: Scaling Web Resources in the Cloud

Arun Kejariwal and John Allspaw, 2017, 2nd edition.

Shows how to measure real capacity ceilings from production data, forecast growth, and plan resources and autoscaling limits before traffic arrives.

Primary sources

  • Paper

    The Tail at Scale (Dean and Barroso, Communications of the ACM, 2013)

    Shows mathematically why rare slow responses dominate user-facing latency once a request fans out to many servers, and describes hedged requests and other ways to reduce the tail.

  • RFC

    RFC 9111: HTTP Caching

    The HTTP caching standard, defining cache keys, Vary, Cache-Control private and no-store, and freshness, the rules that decide whether a shared cache may store and reuse a per-user response.