It is tempting to make every agent action wait for a person. It feels safe, but it often is not. The more useful question is where human judgement adds the most protection.
Approving everything does not scale
If people are asked to approve every step, approvals quickly become routine. Reviewers click through to keep work moving, and the process produces a record of approval without the substance of review. That can be worse than no gate, because it creates false assurance.
Place humans at authority boundaries
Risk-based autonomy lets the agent act on its own where consequences are small and reversible, and brings a person in where they are not. Factors worth weighing include reversibility, the breadth of impact, the target environment, data sensitivity, the quality of the supporting evidence and how unusual the situation is.
Where an action sits on this spectrum is decided by policy and risk, not by how confident the model sounds.
What a good approval looks like
An approver needs the proposed action, the reasoning and evidence behind it, the assessed risk, what will change, and how to undo it. The approval should apply to that action only, and should expire if conditions change. A request that lacks context invites a rubber stamp.
Autonomy should be earned and adjustable
Start with narrow autonomy. Use observability, review of past actions and evaluation to decide where it can be safely extended, and keep the ability to reduce it or switch it off. Oversight does not end at approval: monitoring and post-action review remain part of the loop.
Key takeaways
- Blanket approval leads to approval fatigue and false assurance.
- Use risk factors such as reversibility, scope and evidence quality to decide where people intervene.
- Make approvals informed, specific and recorded.
- Treat autonomy as something to be extended deliberately, with the ability to pull it back.