The mechanics of language models and agent loops, and the failure modes that follow from them: invented APIs, agreement with a bad spec, confident wrong fixes, injected instructions and different results on every run.
Topics
- Next-Token Prediction and Hallucinated APIs
- Models generate the most plausible continuation, which is why they invent functions, flags and packages that look right and do not exist.
- Preference Tuning and Sycophancy
- Training on human approval rewards agreeable answers, so models tend to go along with a flawed spec or a wrong claim.
- Context Windows and Dropped Constraints
- What a model can see at once, and why instructions far back in a long session or buried in a large context get dropped.
- Sampling and Nondeterminism Between Runs
- Why the same request can produce different code on different runs, and what that means for reproducing and reviewing agent work.
- Agent Loops and Confident Wrong Fixes
- How an agent that iterates until a check passes can patch symptoms, weaken tests or report success it never verified.
- Prompt Injection Through Tool Results and Fetched Pages
- Instructions and data arrive in the same channel, so text in a web page, file, issue or tool output can steer the agent that reads it.