The usual advice is correct as far as it goes. Start an agent under close review. Widen what it may do as it proves itself. Keep a person accountable for the outcome. We have written a version of it ourselves, and said that evidence should demote as well as promote without setting out what that means in practice.
This is the piece that sets it out. Not the exact numbers, which belong in a policy signed with a particular client, but the shape: what the levels are, what each one costs in review time, what evidence moves an agent between them, and what takes autonomy back without waiting for anybody to notice.
An AI agent supervision level specifies what the agent may do, how much of its work a person must review, and which evidence can increase or reduce its autonomy.
What does everyone already agree on?
That autonomy should be graded rather than switched on. That is not a recent insight. Thomas Sheridan and William Verplank laid out a ten-point scale of human supervisory control in 1978, running from a person doing everything unaided, through a machine that narrows the options and suggests one, to a machine that acts and tells nobody. Later work by Parasuraman, Sheridan and Wickens separated automation into four functions: gathering information, interpreting it, choosing an action and carrying that action out.
So a graded ladder is close to fifty years old. Current guidance on agents and their operating loops has arrived at the same broad principle: autonomy should expand in steps. The disagreement is not about the ladder.
It is about the rungs. Read the promotion criteria on offer and they are qualitative: the correction rate becomes acceptable, the exception list stabilises, reviewers stop changing things. Those describe a feeling about an agent. None of them can be checked by somebody who was not in the room, which means none of them can be audited later, and it means the decision to widen an agent’s scope is made by whoever is most confident in the meeting.
Why can’t human review carry the control?
Because reliable automation changes how people allocate attention. The issue is not a lack of diligence.
Raja Parasuraman and Dietrich Manzey reviewed the empirical work on this in Human Factors in 2010. Their findings are uncomfortable for anyone whose only control is a person checking outputs. Automation complacency appears in experts as well as novices and is not removed by simple practice. Automation bias is not reliably prevented by training or instructions. It produces both kinds of error: accepting what the machine got wrong, and missing what the machine never raised.
An earlier experiment gives the underlying result. In a 1999 simulated-flight task, participants without an automated aid outperformed participants using a highly but imperfectly reliable aid on monitoring. The aided group made omission errors when the automation failed to flag an event, and commission errors when they followed a bad recommendation despite other valid information being available.
This research predates production AI agents and comes from automated decision support, much of it in aviation. It does not establish a failure rate for agent review. It does establish why human review should not carry the control by itself: attention is bounded, and dependable automation changes where people spend it.
Human review remains necessary. The review rate, the sampling and the escalation rules need to be structural. How much of the agent’s output a person must see should be set in policy and enforced by the system, rather than left to the reviewer’s judgement on the day.
What are the three levels?
Every agent we run sits at one of three, and the level is a setting in Aitonomy Control. Anyone with access can see which level an agent is at today and the evidence that put it there.
| What the agent does | What a person does | |
|---|---|---|
| Suggests | Produces a recommendation and changes nothing | Substantively reviews every result and performs or approves the action |
| Drafts | Completes the work and stages the action | Approves each action, with targeted review of exceptions and selected detail |
| Auto | Completes agreed routine cases and acts within scope | Reviews a sampled share and every escalation, and can pause the agent |
Every agent starts at Suggests. There is no fast track for a process that somebody is confident about, and no agent begins at Auto because a similar one elsewhere earned it.
The important column is the third. A level is a commitment about how much human attention the work will consume, which puts it in the operating cost as well as in the controls. Moving an agent from Suggests to Drafts can start giving time back, provided approval takes less effort than doing the original work and the released capacity can be used. That is also why a level is never widened casually.
What moves an agent up a level?
Evidence of five kinds, in a fixed combination. Control measures all five continuously, against thresholds set before the agent ran.
Time at the current level. Long enough for the process to show its variation, including the parts of the month or the quarter that behave differently.
Volume. Enough completed cases that the intervention rate means something. A handful of clean results is not evidence.
Intervention rate. The share of the agent’s output a person changed before it took effect, falling and staying below the level’s figure.
Durability. The share still standing after the agreed window, which is a different number from the share accepted on the day. The gap between the two is work somebody reopened, corrected or reversed, and an agent promoted on day-one acceptance alone has been promoted on the wrong signal.
No critical errors. A critical error is an outcome that crosses a boundary defined for that agent before it runs. The policy states which outcomes qualify, what evidence is required and who may record one. None may occur during the qualifying period.
We begin from a default policy and agree the thresholds agent by agent, then load them into Control so the number being enforced is always the number that was agreed. A process-specific deviation is configured there too, with its reason attached. The figures stay in the signed policy, because the right threshold for an invoice match is not the right threshold for something whose mistakes reach a customer. What does not vary is that the thresholds are fixed before the agent runs and remain the numbers used to judge it later.
Why fix them in advance
A threshold agreed after the results are in is not a threshold. If the number can move to fit the evidence, the evidence has stopped doing any work, and the promotion decision has quietly returned to whoever is most confident in the meeting.
What moves it back down?
The same numbers. A ladder that only goes up describes a launch. The downward half is what makes it supervision, and it is often the part left unspecified.
A critical error demotes immediately. Once one is recorded under the agreed definition, Control drops the level. One is enough. It does not wait for the monthly review and is not weighed against how well the agent has otherwise been running.
A rising intervention rate freezes scope. This is not a demotion by itself, but the agent stops widening while somebody establishes the cause. Its scope does not expand while the rate of people correcting it is going up.
A model change sends it back to the test set. Changing the model invalidates the evidence attached to the previous configuration until the held-out cases have been replayed. The agent keeps its level only if the new configuration clears the same standard.
Control applies the configured response when the evidence reaches one of these conditions. It does not rely on a reviewer noticing drift during routine approval, which would inherit the weakness the levels exist to reduce.
Where do the levels live?
In Aitonomy Control. It holds the current level for each agent, the evidence behind it, the thresholds being enforced and the decision history. A person approves promotion. When recorded evidence reaches a configured demotion condition, Control freezes or reduces autonomy without waiting for the monthly review. This produces a traceable operating record that an auditor or regulator can inspect, rather than a reconstruction after the event.
Once a month, every agent gets a decision: widen, hold, step back or retire. The evidence for it is already gathered, so the meeting is about the verdict rather than about assembling the case for one. Hold should be a normal result. A portfolio where every agent is being widened should prompt a question about whether the intervention rate and durability evidence are carrying enough weight.
There is a short test for whether any of this is real in an organisation, and it does not require access to the system. Ask when an agent was last moved down a level. Ask who has the authority to do it, and whether that person needs anyone’s permission. If nobody can explain how demotion would happen, or nobody owns the authority to do it, the levels are labels.
An agent does not become trustworthy because it has been running for a while. It becomes trustworthy because it produced evidence against a standard set before anyone knew whether it would clear it, and because the same standard is still able to take the trust back.
Common questions
What are AI agent supervision levels?+
How does an AI agent earn more autonomy?+
Can human review catch every mistake an AI agent makes?+
When should an AI agent lose autonomy?+
Sources
- Thomas Sheridan and William Verplank, Human and Computer Control of Undersea Teleoperators, MIT Man-Machine Systems Laboratory (1978), the ten-point scale of human supervisory control that later levels-of-automation work is built on
- Raja Parasuraman, Thomas Sheridan and Christopher Wickens, A Model for Types and Levels of Human Interaction with Automation, IEEE Transactions on Systems, Man, and Cybernetics 30:3 (2000), separating automation into information acquisition, analysis, decision selection and action implementation
- Linda Skitka, Kathleen Mosier and Mark Burdick, Does Automation Bias Decision-Making?, International Journal of Human-Computer Studies 51:5 (1999), comparing monitoring with and without a highly but imperfectly reliable automated aid
- Raja Parasuraman and Dietrich Manzey, Complacency and Bias in Human Use of Automation: An Attentional Integration, Human Factors 52:3 (2010), reviewing automation complacency, automation bias and the limits of practice, training and instructions
- The three levels, the promotion evidence and the demotion rules are Aitonomy’s own, from the autonomy policy we sign per agent and enforce in Aitonomy Control. The thresholds themselves are written in that policy and are not published here.