Field worksheet · fill in before funding
Chapter 14 decision worksheet · companion to Human Out of the Loop by John Whitman · v1.0, 2026-09-29 · printable, standalone, no external assets
Fill every field before funding or dispatching. A field you cannot fill is itself a finding — write “cannot fill” and the reason, and count it against GO.
Why fill it in honestly: the framework’s first full run returned a verdict its own authors hated — it was designed to return a verdict (ch14:11). That is the point of the instrument.
Built for the technical leader deciding whether to fund one of these things (ch14:11) — not for choosing a model or a vendor. The decision this sheet forces is whether to run agents at all, how many, and under what stop rules.
| Name and role | |
| Team / organization | |
| Date filled | |
| The problem, in one sentence (who feels it, how often) |
Worksheet prompt: list what you can show today — tickets per month, hours spent, measured cost of the status quo, named people asking. If the list is empty, write “zero receipts” and treat Gate 1 accordingly.
| Receipt 1 | |
| Receipt 2 | |
| Receipt 3 |
The decision decomposes into four gates, in order (ch14:17–20): (1) agent at all, or deterministic software? (2) more than one agent? (3) factory or hand-built? (4) own or rent?
The dividing line is not “AI code” versus “human code”; it is deterministic execution paths versus stochastic decision points (ch14:32). Rule: confine the stochastic part to the smallest zone that pays for it, and let everything else stay software (ch14:32).
Agents earn their keep where specification is the bottleneck — underspecified requirements, cross-domain integration, one-off tasks, prototyping, runtime adaptation (ch14:28). Traditional architecture wins where execution is the bottleneck — high-volume low-latency paths, regulated and auditable domains, cost-sensitive scale, long-lived systems (ch14:30).
Answer before any model question: who can tell my agent it’s wrong, how fast, and at what cost? (ch14:34) If the answer is “a professional, slowly, with liability attached,” the bar is much higher — or the real project is building the verifier first (ch14:34).
| The bounded task, in one sentence (single-purpose) | |
| Why specification — not execution — is the bottleneck here | |
| Who/what can verify outputs: machine test / named expert / nobody | |
| If expert or nobody: cost and date to build the verifier first |
One wrong answer costs us $ because
If you cannot write that sentence with a defensible number, you are not ready to choose an architecture (ch14:50).
The second gate of the four-gate decomposition (ch14:17–20). The chapter frames it with reported volume, team, budget and accuracy thresholds from a fleet-authored synthesis; this worksheet deliberately reproduces none of them. A threshold you did not measure and did not choose is not a gate — it is somebody else’s number. The gate opens on five of your own inputs: a measured baseline, priced alternatives, the incremental value of a fleet over the best alternative, the support capacity you can actually sustain, and pass/fail criteria you declare before scoring.
| Monthly task volume, measured from your ticket/request system (or: not measured — measure first) | |
| Measured baseline: today’s single-agent or scripted accuracy and cost per task (or: not measured — measure first) | |
| Alternatives you will compare against (scripted workflow / single agent / vendor tooling / do nothing) | |
| Incremental value of a fleet over the best alternative — measured, or explicit estimate with basis | |
| Support capacity: engineers and hours/week able to sustain orchestration, review and evals | |
| Budget actually committed to the program ($/mo) | |
| Pass/fail criteria for this gate, written before scoring — your metric, threshold and window | |
| Gate 2 verdict by those criteria: pass / fail / cannot decide — and why |
Practice fields; the survey statistics the book attached to this list are provenance-flagged and omitted — the practices stand (ch14:48).
| Named executive sponsor | |
| Success metric, one sentence, set at kickoff | |
| Integration with the system of record | |
| Evaluation + review share of program budget (fixed share, named now) | % |
The third gate (ch14:17–20). The chapter’s rule of thumb is qualitative: if your realistic roster is a handful of agents, you do not need a factory — you need a few well-built agents and a good eval harness (ch14:46). The break-even roster count the chapter prints alongside that rule is a reported synthesis figure and is not reproduced here; price factory versus hand-built with your own numbers instead.
| Realistic roster size, 12 months out — your projection and its basis | |
| Hand-built basis: who builds and maintains each agent, by name; measured or estimated cost per agent | |
| Factory basis: who builds and maintains the generation and eval machinery; measured or estimated cost | |
| Owner of the evaluation harness | |
| Pass/fail criterion for this gate, written before scoring — factory beats hand-built on what measured number, by when | |
| Gate 3 verdict by that criterion: factory / hand-built / cannot decide — and why |
Own the identity layer, rent everything else (ch14:88). Own: the identity specification, the provenance/falsification machinery, governance, dispatch logic — the things that encode your judgment about your agents. Rent: inference, embeddings, vector storage, runtime/hosting, messaging, observability (ch14:88).
Operational test: if you can swap the vendor without your agents changing who they are, rent it (ch14:88). Build a protocol, not a platform (ch14:90). Sequence: substrate before factory — no new capabilities until existing ones are reliable (ch14:92).
| Layers you will own | |
| Layers you will rent | |
| Swap test: can you change model vendor without changing agent identity? |
Fund the fleet and the audit function as one budget item or you have funded neither (ch14:60). Budget the ops, not the generation: generation is the cheap part — generation scales cheaply with complexity while verification does not (ch14:58, ch14:112).
| Hard spend ceiling, enforced at infrastructure layer ($/mo) | |
| Hard token ceiling AND rate limit (not the same control) | |
| Evaluation + review share of budget (from §3) | % |
| Pilot window (weeks), then this sheet is re-filled | |
| Decision date for GO/HOLD/STOP (§9) | |
| Year-two review capacity: who reviews generated output, hours/week |
The dated maintenance-cliff figure the book carries is provenance-flagged and omitted here (ch14:161); the practice — fund review capacity up front — stands.
Attach a falsifier with a date. The pattern is the fleet memo’s Recommendation 3: “Strategy without revenue is a hobby.” and “If yes, that’s the pitch deck. If no, we learn why before investing further.” (ch14:7). Name the program’s own falsifier — the receipt that would prove the whole thing is working, and the date by which its absence means it isn’t (ch14:143).
| Program falsifier: “By [date], [artifact/receipt] will show [measured change].” | |
| Success metric (same words as §3 — no drifting) | |
| Stop rule: “If by [date] [metric] is not [threshold], we stop / downgrade to single agent / re-scope.” | |
| Weekly ledger to keep: fired / done / failed / revenue — or your ground-truth equivalent (ch14:132) | |
| Where the ledger lives (wired to ground truth, not self-reported) |
Track trends over snapshots — a single green number is noise (ch14:129).
The honest default, stated before the decision block: most organizations reading this should not run an autonomous fleet (ch14:98). A STOP on this sheet is a legitimate outcome, not a failure of the exercise.
| Boundary (money / send / delete / identity / other) | Enforced by (code path, not prompt) | Named approver | Verified in live-fire path (date, by whom) |
|---|---|---|---|
| This program’s never-automate list (minimum the three above) | |
| Who ratifies self-modification (named human) | |
| Who audits composition as built (named human, outside the program) | |
| Kill-switch drill: date run and result |
Check only with a receipt (document, date, path — “done” is not evidence).
| Done | Item (source) | Evidence / receipt |
|---|---|---|
| Success metric and error price written at kickoff (ch14:50) | ||
| Gate 2 pass/fail criteria written before scoring and applied to your measured numbers (§3) | ||
| Verifier named; its cost and speed known (ch14:34) | ||
| Stochastic zone confined; generated artifacts run as conventional, testable code (ch14:32) | ||
| Ops and audit funded as one budget item (ch14:60) | ||
| Year-two review capacity funded and named (ch14:58) | ||
| Least privilege enforced on tools and filesystem (ch14:117) | ||
| Money/send/delete/identity enforced by external code, not prompts (ch14:118) | ||
| External enforcement layers in place; ceilings AND rate limits (ch14:119) | ||
| Reversibility-gated approvals wired (ch14:120) | ||
| Never-automate list excluded by construction (ch14:121) | ||
| Kill-switch drill passed and recorded (ch14:122) | ||
| As-built audit completed by a named outsider, on a schedule (ch14:123) | ||
| Tool layer scanned (ch14:124) | ||
| Substrate stable: no new capabilities until existing ones are reliable (ch14:92) |
GO requires the gates to actually open (ch14:100) — opened by the criteria you wrote down in §3–§4 before you scored them, not by any default of this sheet: work in the stochastic-decision zone and machine-verifiable output; your Gate-2 criteria met, with the baseline measured and the error price written down; factory machinery justified by your own cost basis (or a deliberate single-agent scope); verification and governance funded as a first-class budget with a named owner; the identity layer in place before the first unattended dispatch; both permanent human posts staffed (ch14:100).
| Decision — mark one: GO / HOLD / STOP | |
| Date of decision | |
| Owner (the named executive sponsor from §3) | |
| Second reader (human outside the program) | |
| If GO: conditions attached (e.g., “after items 7–14 of §9 are evidenced”) | |
| Review date — this sheet is re-filled as a governance event, not by drift (ch14:141) |
A platform team fills: bounded task = “triage the weekly infrastructure-audit findings”; verifier = linter + test suite + on-call acceptance; error price = “$150 per false positive paged to on-call” (their own on-call math, written as the §2 sentence). Volume they can state from the ticket system — but the single-agent baseline has never been run and no alternative’s cost is measured, so the pass/fail criterion they wrote for Gate 2 (“beat the scripted baseline on false-positive pages over a four-week window, at lower cost per finding”) cannot be scored. The sheet returns HOLD — measure the baseline and re-score, with a falsifier (“by [date], the agent closes low-severity findings with zero false-positive pages, measured against the ticket ledger”) and a stop rule attached. Every number in this example is invented to illustrate evidence-based selection; the sheet supplies no threshold — the team’s own unmet criterion, not a book figure, is what returns HOLD.