Watch
The medical test paradox, and redesigning Bayes' rule
Works through why a highly accurate test for a rare condition produces mostly false positives, which carries over directly to flaky-test reruns, alerts and classifiers.
Engineering Fundamentals for the Agent Era
Contents Section 4, Math and Algorithms
Average latency reported while a slow tail, which most users hit at least once per session, stays hidden.
A prompt or model change declared an improvement from a handful of runs, with no estimate of run-to-run variance.
Failures treated as independent when they share a cause, such as two replicas in the same data center.
A flaky test called fixed after a single green run.
Distributions, sampling, variance and significance: the tools for judging whether a measurement, an experiment or an eval result means anything.
An agent reports a flaky test as fixed after three consecutive green runs; before the fix it failed about one run in ten. Compute the chance of three greens with the bug still present, and say how many consecutive green runs it takes before a bug that is still there would produce that streak less than 1 percent of the time.
Watch
Works through why a highly accurate test for a rare condition produces mostly false positives, which carries over directly to flaky-test reruns, alerts and classifiers.
Shows why averages and naive percentiles hide the slow tail that most users hit, and how coordinated omission corrupts latency measurements.
Builds intuition for how sample means spread and shrink with sample size, the basis for confidence intervals and for deciding how many runs are enough.
Teaches distributions, percentiles, estimation and hypothesis testing through runnable notebooks, so the ideas turn into code a developer can apply to their own measurements.
Covers the ways experiments mislead people in practice, including underpowered tests, peeking and invalid randomization, drawn from running A/B tests at Microsoft, Google and LinkedIn.
The Harvard Stat 110 text on conditional probability, independence, expectation and distributions, for readers who want the foundations behind the rules of thumb.
Reference
The government reference for sampling, distributions, confidence intervals and hypothesis tests, with worked examples you can check an agent's statistical claim against.
Paper
Shows how rare slow responses become the common case once a request fans out across many servers, with the percentile arithmetic behind that.