Delegating to AI agents safely

Hand off the work, not the accountability — bound what it can touch, and check the trace where a wrong turn would compound.

The idea

An agent that runs many steps on its own is only as safe as the around it. A wrong assumption early on doesn't announce itself — it quietly feeds every step after it, so the tidy-looking result at the end can be built on sand.

Three moves keep you in control: scope what it may touch (read, draft, or act), place a checkpoint right after the step where a mistake would start compounding, and review the trace — the reasoning — not just the final output. The human stays accountable throughout.

Run the agent

permission scope
checkpoints — review after step…
A ten-step agent run with permission scope wall and review checkpoints.
review overhead0
cleanup cost—
total cost—
Set a scope, drop a checkpoint or two, then run the agent. A wrong assumption is waiting at step 5.

How it works

You're weighing two costs against each other: the review overhead of checking the agent, and the cleanup cost when a mistake slips through. The art is spending a little review exactly where it saves a lot of cleanup.

total cost = review overhead + cleanup

The wrong assumption enters at step 5 and compounds each step.

no checkpoint, "act" scope
  it reaches the irreversible send -> cleanup 48   total 48
checkpoint after step 3 (too early to catch anything)
  pays review, catches nothing, still sends -> total 51   (worse!)
checkpoint after step 5 (right where it compounds)
  caught immediately -> review 3 + cleanup 4   total 7
"read-only" scope, no checkpoint
  walled before any change -> nothing ships, but the wrong
  total still sits in your drafts -> cleanup 12   total 12

Scope limits the ; a checkpoint is what actually catches the mistake; and set a time/cost stop-rule so a looping or runaway agent halts on its own.

When to use it

Reach forWhen
Tight scope (read-only / draft)Actions are irreversible or hard to undo — sending, paying, deleting, publishing.
A checkpointA step makes an assumption everything downstream depends on (a total, a match, a query).
Trace reviewThe output looks plausible — plausibility is exactly when a silent wrong assumption hides.
A stop-ruleThe agent could loop, retry, or spend unbounded time or tokens.

The trade-off: every checkpoint and every narrowing of scope costs you speed and autonomy. Too much oversight and you've just done the task yourself; too little and one bad assumption ships.

Watch out for

Worked example

You ask an agent to reconcile last month's invoices and email vendors any corrections. At step 5 it sums the totals assuming every invoice is in USD — but three are in EUR, so the figures are quietly wrong. With "act" scope and no checkpoint, that error rides all the way to step 9 and lands in twelve vendors' inboxes: expensive to walk back. Instead you give it "draft" scope so it can't send, and drop one checkpoint right after step 5. Reviewing the trace there, you spot the currency assumption, correct it, and let it finish — a few minutes of review instead of a day of apologies. In an interview, name the three moves out loud: scope, checkpoint at the compounding step, review the trace — and add that you'd keep a stop-rule on cost and own the result either way.

Check yourself

You'll let an agent draft and send follow-up emails to leads overnight, unattended. What's the safest single move?