How do you know an agent’s “done” is true?
The book takes this question on directly, in its chapter on measuring agents. Here are two short passages, quoted exactly as written.
The graders are separate from the graded. Grades carry confidence scores, so a reluctant A- and a ringing A- are distinguishable. Reviews are chased, not waited for, and the chasing is logged. Grades are versioned and re-issued when the world changes, never edited in place. And disagreement has lanes — a reviewer can accept the work while filing a structural objection that gets recorded rather than resolved by whoever types last.
[…]
The first is the definition-of-done audit’s kill condition: “two consecutive audits pass without findings.” The audit fails by succeeding twice. The reasoning, a “sweep drift signal” in the weekly scorecard clause, is that in a system this size the true defect rate is never zero, so an audit that stops finding things has stopped looking, and an idle detector is more dangerous than the defects it’s missing because it also emits reassurance. The fleet measured its measurement instruments and pre-committed to distrusting them at their most flattering.
— from Chapter 6, “Measuring Agents When Every Benchmark Is Gamed”
That was a few paragraphs of a fifteen-chapter book.
Buy the full book — $12 →All 15 chapters, EPUB + web, every revision through v1.0 · 14-day no-questions refund
or read the free sample first — the Introduction and Chapter 1 →
More free excerpts:
→ What is an agent factory, really?
→ What should an agent do when a human doesn’t answer?
→ Why did my agents keep spawning sessions and burning money?
How this book was made: it was drafted by an AI (Anthropic's Claude) from the primary documents the fleet produced and preserved, under John Whitman's direction. A revised edition (v1.0) is correcting the claims an editorial audit could not trace to those documents.