9
Directing Agents: Delegation, Review and Accountability
How agents work and fail, how to hand them work, how to check what comes back, and what stays yours.
Why it matters when agents do the typing
Whoever approves a change answers for it, whether a person or an agent typed it, which makes the agents part of the system you are responsible for. Directing them well means knowing how they fail and how they are attacked, giving them clear intent, limiting what they can touch, and reviewing what comes back. This section comes last because it draws on everything before it: you can't review what you don't understand.
-
Core 19
Depth: do
How Agents Work and How They Fail
The mechanics of language models and agent loops, and the failure modes that follow from them: invented APIs, agreement with a bad spec, confident wrong fixes, injected instructions and different results on every run.
A mistake it teaches you to catch: A call to a client-library option that does not exist, such as a retries setting the library never had, written with full confidence.
-
Core 20
Depth: do
Prompting and Specifying Work for Agents
The acceptance criteria from Intent, plus the three things a prompt adds: the context the agent needs, the boundaries of what it may touch, and the check that proves it is done.
A mistake it teaches you to catch: An agent told to 'make the tests pass' that deletes or weakens the tests.
-
Core 21
Depth: do
Reading and Reviewing Code You Did Not Write
Reading diffs and unfamiliar code critically: what to check first, where agent changes typically go wrong, and how to review more code than you could ever write.
A mistake it teaches you to catch: A change that passes the tests but alters behavior outside the requested task.
-
Depth: explain
Evals: Testing AI Behavior
Measuring nondeterministic systems with datasets, graders and statistics, so a prompt or model change is judged on evidence.
A mistake it teaches you to catch: A prompt change shipped because it looked better on three examples.
-
Depth: explain
Orchestration, Permissions and Guardrails
Running agents in loops and pipelines with limited permissions, budgets and approval points, so their mistakes stay small.
A mistake it teaches you to catch: A CI agent given the organization-wide deploy token because scoping one to the repository was more setup, so any repository it touches can ship to production.
-
Core 22
Depth: do
Accountability: Owning Code You Did Not Type
What stays yours when agents do the work: the approval, the incident, the license and the dependencies, plus the skills to act on them yourself.
A mistake it teaches you to catch: A 600-line agent change approved within a minute, with approval treated as a formality.