Field worksheet · fill in before funding

Should You Run an Autonomous Agent Fleet?

Chapter 14 decision worksheet · companion to Human Out of the Loop by John Whitman · v1.0, 2026-09-29 · printable, standalone, no external assets

§0How to use this sheet

Fill every field before funding or dispatching. A field you cannot fill is itself a finding — write “cannot fill” and the reason, and count it against GO.

Why fill it in honestly: the framework’s first full run returned a verdict its own authors hated — it was designed to return a verdict (ch14:11). That is the point of the instrument.

Deliberate omissions, stated up front. The book’s ROI-distribution figures, the named “18-Month Wall” with its 16–18-month dating, and the vendor latency case study (12% abandonment / ~$27K) are excluded from this worksheet. The chapter itself flags all three: the latency case is “one vendor case study this edition could not independently verify” (ch14:43); the ROI stack “circulates through secondary aggregators with no recoverable primary” (ch14:48); the wall’s provenance is marked thin in the chapter’s own TK list (ch14:161). The practices those numbers decorated are kept below as fill-in prompts, without the figures. The Gate-2 volume, team-size, budget and accuracy thresholds and the Gate-3 factory break-even count (ch14:40, ch14:41, ch14:42, ch14:44, ch14:46) are likewise not reproduced: they are reported synthesis figures, and no number from the book is made an operative criterion here — every pass/fail line in this worksheet is one the reader writes.

§1Intended user, problem, and evidence for demand

Built for the technical leader deciding whether to fund one of these things (ch14:11) — not for choosing a model or a vendor. The decision this sheet forces is whether to run agents at all, how many, and under what stop rules.

Identity and problem
Name and role
Team / organization
Date filled
The problem, in one sentence (who feels it, how often)

Evidence for demand — receipts, not enthusiasm

Worksheet prompt: list what you can show today — tickets per month, hours spent, measured cost of the status quo, named people asking. If the list is empty, write “zero receipts” and treat Gate 1 accordingly.

Receipts
Receipt 1
Receipt 2
Receipt 3

The decision decomposes into four gates, in order (ch14:17–20): (1) agent at all, or deterministic software? (2) more than one agent? (3) factory or hand-built? (4) own or rent?

§2Gate 1 — Does this problem want an agent at all?

The dividing line is not “AI code” versus “human code”; it is deterministic execution paths versus stochastic decision points (ch14:32). Rule: confine the stochastic part to the smallest zone that pays for it, and let everything else stay software (ch14:32).

Agents earn their keep where specification is the bottleneck — underspecified requirements, cross-domain integration, one-off tasks, prototyping, runtime adaptation (ch14:28). Traditional architecture wins where execution is the bottleneck — high-volume low-latency paths, regulated and auditable domains, cost-sensitive scale, long-lived systems (ch14:30).

Answer before any model question: who can tell my agent it’s wrong, how fast, and at what cost? (ch14:34) If the answer is “a professional, slowly, with liability attached,” the bar is much higher — or the real project is building the verifier first (ch14:34).

Gate 1 fields
The bounded task, in one sentence (single-purpose)
Why specification — not execution — is the bottleneck here
Who/what can verify outputs: machine test / named expert / nobody
If expert or nobody: cost and date to build the verifier first

Write the error price sentence (ch14:50)

One wrong answer costs us $ because

If you cannot write that sentence with a defensible number, you are not ready to choose an architecture (ch14:50).

§3Gate 2 — Does it want more than one agent?

The second gate of the four-gate decomposition (ch14:17–20). The chapter frames it with reported volume, team, budget and accuracy thresholds from a fleet-authored synthesis; this worksheet deliberately reproduces none of them. A threshold you did not measure and did not choose is not a gate — it is somebody else’s number. The gate opens on five of your own inputs: a measured baseline, priced alternatives, the incremental value of a fleet over the best alternative, the support capacity you can actually sustain, and pass/fail criteria you declare before scoring.

Gate 2 fields
Monthly task volume, measured from your ticket/request system (or: not measured — measure first)
Measured baseline: today’s single-agent or scripted accuracy and cost per task (or: not measured — measure first)
Alternatives you will compare against (scripted workflow / single agent / vendor tooling / do nothing)
Incremental value of a fleet over the best alternative — measured, or explicit estimate with basis
Support capacity: engineers and hours/week able to sustain orchestration, review and evals
Budget actually committed to the program ($/mo)
Pass/fail criteria for this gate, written before scoring — your metric, threshold and window
Gate 2 verdict by those criteria: pass / fail / cannot decide — and why

Governance artifacts to name at kickoff

Practice fields; the survey statistics the book attached to this list are provenance-flagged and omitted — the practices stand (ch14:48).

Kickoff governance
Named executive sponsor
Success metric, one sentence, set at kickoff
Integration with the system of record
Evaluation + review share of program budget (fixed share, named now) %

§4Gate 3 — Does it want a factory?

The third gate (ch14:17–20). The chapter’s rule of thumb is qualitative: if your realistic roster is a handful of agents, you do not need a factory — you need a few well-built agents and a good eval harness (ch14:46). The break-even roster count the chapter prints alongside that rule is a reported synthesis figure and is not reproduced here; price factory versus hand-built with your own numbers instead.

Gate 3 fields
Realistic roster size, 12 months out — your projection and its basis
Hand-built basis: who builds and maintains each agent, by name; measured or estimated cost per agent
Factory basis: who builds and maintains the generation and eval machinery; measured or estimated cost
Owner of the evaluation harness
Pass/fail criterion for this gate, written before scoring — factory beats hand-built on what measured number, by when
Gate 3 verdict by that criterion: factory / hand-built / cannot decide — and why

§5Gate 4 — If you run it: own vs rent

Own the identity layer, rent everything else (ch14:88). Own: the identity specification, the provenance/falsification machinery, governance, dispatch logic — the things that encode your judgment about your agents. Rent: inference, embeddings, vector storage, runtime/hosting, messaging, observability (ch14:88).

Operational test: if you can swap the vendor without your agents changing who they are, rent it (ch14:88). Build a protocol, not a platform (ch14:90). Sequence: substrate before factory — no new capabilities until existing ones are reliable (ch14:92).

Own vs rent
Layers you will own
Layers you will rent
Swap test: can you change model vendor without changing agent identity?

§6Budget and time limits

Fund the fleet and the audit function as one budget item or you have funded neither (ch14:60). Budget the ops, not the generation: generation is the cheap part — generation scales cheaply with complexity while verification does not (ch14:58, ch14:112).

Limits
Hard spend ceiling, enforced at infrastructure layer ($/mo)
Hard token ceiling AND rate limit (not the same control)
Evaluation + review share of budget (from §3) %
Pilot window (weeks), then this sheet is re-filled
Decision date for GO/HOLD/STOP (§9)
Year-two review capacity: who reviews generated output, hours/week

The dated maintenance-cliff figure the book carries is provenance-flagged and omitted here (ch14:161); the practice — fund review capacity up front — stands.

§7Success, failure, and stop criteria

Attach a falsifier with a date. The pattern is the fleet memo’s Recommendation 3: “Strategy without revenue is a hobby.” and “If yes, that’s the pitch deck. If no, we learn why before investing further.” (ch14:7). Name the program’s own falsifier — the receipt that would prove the whole thing is working, and the date by which its absence means it isn’t (ch14:143).

Criteria
Program falsifier: “By [date], [artifact/receipt] will show [measured change].”
Success metric (same words as §3 — no drifting)
Stop rule: “If by [date] [metric] is not [threshold], we stop / downgrade to single agent / re-scope.”
Weekly ledger to keep: fired / done / failed / revenue — or your ground-truth equivalent (ch14:132)
Where the ledger lives (wired to ground truth, not self-reported)

Track trends over snapshots — a single green number is noise (ch14:129).

The honest default, stated before the decision block: most organizations reading this should not run an autonomous fleet (ch14:98). A STOP on this sheet is a legitimate outcome, not a failure of the exercise.

§8Approval boundaries — before any unattended run

Boundary register — enforced by code, not prompts
Boundary (money / send / delete / identity / other)Enforced by (code path, not prompt)Named approverVerified in live-fire path (date, by whom)
Approval fields
This program’s never-automate list (minimum the three above)
Who ratifies self-modification (named human)
Who audits composition as built (named human, outside the program)
Kill-switch drill: date run and result

§9Evidence and release checklist — gate to first unattended dispatch

Check only with a receipt (document, date, path — “done” is not evidence).

Release checklist
DoneItem (source)Evidence / receipt
Success metric and error price written at kickoff (ch14:50)
Gate 2 pass/fail criteria written before scoring and applied to your measured numbers (§3)
Verifier named; its cost and speed known (ch14:34)
Stochastic zone confined; generated artifacts run as conventional, testable code (ch14:32)
Ops and audit funded as one budget item (ch14:60)
Year-two review capacity funded and named (ch14:58)
Least privilege enforced on tools and filesystem (ch14:117)
Money/send/delete/identity enforced by external code, not prompts (ch14:118)
External enforcement layers in place; ceilings AND rate limits (ch14:119)
Reversibility-gated approvals wired (ch14:120)
Never-automate list excluded by construction (ch14:121)
Kill-switch drill passed and recorded (ch14:122)
As-built audit completed by a named outsider, on a schedule (ch14:123)
Tool layer scanned (ch14:124)
Substrate stable: no new capabilities until existing ones are reliable (ch14:92)

§10The decision

GO requires the gates to actually open (ch14:100) — opened by the criteria you wrote down in §3–§4 before you scored them, not by any default of this sheet: work in the stochastic-decision zone and machine-verifiable output; your Gate-2 criteria met, with the baseline measured and the error price written down; factory machinery justified by your own cost basis (or a deliberate single-agent scope); verification and governance funded as a first-class budget with a named owner; the identity layer in place before the first unattended dispatch; both permanent human posts staffed (ch14:100).

GO HOLD STOP
Decision record
Decision — mark one: GO / HOLD / STOP
Date of decision
Owner (the named executive sponsor from §3)
Second reader (human outside the program)
If GO: conditions attached (e.g., “after items 7–14 of §9 are evidenced”)
Review date — this sheet is re-filled as a governance event, not by drift (ch14:141)

§11Example use (synthetic — invented to show the form; not a customer, not evidence)

A platform team fills: bounded task = “triage the weekly infrastructure-audit findings”; verifier = linter + test suite + on-call acceptance; error price = “$150 per false positive paged to on-call” (their own on-call math, written as the §2 sentence). Volume they can state from the ticket system — but the single-agent baseline has never been run and no alternative’s cost is measured, so the pass/fail criterion they wrote for Gate 2 (“beat the scripted baseline on false-positive pages over a four-week window, at lower cost per finding”) cannot be scored. The sheet returns HOLD — measure the baseline and re-score, with a falsifier (“by [date], the agent closes low-severity findings with zero false-positive pages, measured against the ticket ledger”) and a stop rule attached. Every number in this example is invented to illustrate evidence-based selection; the sheet supplies no threshold — the team’s own unmet criterion, not a book figure, is what returns HOLD.