Field notes · Supervision

Autonomy is earned: supervision levels for agents in production

An agent earns autonomy by clearing thresholds agreed before it starts: enough time and volume at its current level, a falling intervention rate, durable outcomes and no critical errors. The same policy must also be able to take autonomy back, and reviewer confidence should not be what decides.

Published by Aitonomy We run agent workforces for companies. This is where we write down what that actually takes. What we do

The usual advice is correct as far as it goes. Start an agent under close review. Widen what it may do as it proves itself. Keep a person accountable for the outcome. We have written a version of it ourselves, and said that evidence should demote as well as promote without setting out what that means in practice.

This is the piece that sets it out. Not the exact numbers, which belong in a policy signed with a particular client, but the shape: what the levels are, what each one costs in review time, what evidence moves an agent between them, and what takes autonomy back without waiting for anybody to notice.

An AI agent supervision level specifies what the agent may do, how much of its work a person must review, and which evidence can increase or reduce its autonomy.

What does everyone already agree on?

That autonomy should be graded rather than switched on. That is not a recent insight. Thomas Sheridan and William Verplank laid out a ten-point scale of human supervisory control in 1978, running from a person doing everything unaided, through a machine that narrows the options and suggests one, to a machine that acts and tells nobody. Later work by Parasuraman, Sheridan and Wickens separated automation into four functions: gathering information, interpreting it, choosing an action and carrying that action out.

So a graded ladder is close to fifty years old. Current guidance on agents and their operating loops has arrived at the same broad principle: autonomy should expand in steps. The disagreement is not about the ladder.

It is about the rungs. Read the promotion criteria on offer and they are qualitative: the correction rate becomes acceptable, the exception list stabilises, reviewers stop changing things. Those describe a feeling about an agent. None of them can be checked by somebody who was not in the room, which means none of them can be audited later, and it means the decision to widen an agent’s scope is made by whoever is most confident in the meeting.

Why can’t human review carry the control?

Because reliable automation changes how people allocate attention. The issue is not a lack of diligence.

Raja Parasuraman and Dietrich Manzey reviewed the empirical work on this in Human Factors in 2010. Their findings are uncomfortable for anyone whose only control is a person checking outputs. Automation complacency appears in experts as well as novices and is not removed by simple practice. Automation bias is not reliably prevented by training or instructions. It produces both kinds of error: accepting what the machine got wrong, and missing what the machine never raised.

An earlier experiment gives the underlying result. In a 1999 simulated-flight task, participants without an automated aid outperformed participants using a highly but imperfectly reliable aid on monitoring. The aided group made omission errors when the automation failed to flag an event, and commission errors when they followed a bad recommendation despite other valid information being available.

This research predates production AI agents and comes from automated decision support, much of it in aviation. It does not establish a failure rate for agent review. It does establish why human review should not carry the control by itself: attention is bounded, and dependable automation changes where people spend it.

Human review remains necessary. The review rate, the sampling and the escalation rules need to be structural. How much of the agent’s output a person must see should be set in policy and enforced by the system, rather than left to the reviewer’s judgement on the day.

What are the three levels?

Every agent we run sits at one of three, and the level is a setting in Aitonomy Control. Anyone with access can see which level an agent is at today and the evidence that put it there.

What the agent doesWhat a person does
SuggestsProduces a recommendation and changes nothingSubstantively reviews every result and performs or approves the action
DraftsCompletes the work and stages the actionApproves each action, with targeted review of exceptions and selected detail
AutoCompletes agreed routine cases and acts within scopeReviews a sampled share and every escalation, and can pause the agent

Every agent starts at Suggests. There is no fast track for a process that somebody is confident about, and no agent begins at Auto because a similar one elsewhere earned it.

The important column is the third. A level is a commitment about how much human attention the work will consume, which puts it in the operating cost as well as in the controls. Moving an agent from Suggests to Drafts can start giving time back, provided approval takes less effort than doing the original work and the released capacity can be used. That is also why a level is never widened casually.

What moves an agent up a level?

Evidence of five kinds, in a fixed combination. Control measures all five continuously, against thresholds set before the agent ran.

Time at the current level. Long enough for the process to show its variation, including the parts of the month or the quarter that behave differently.

Volume. Enough completed cases that the intervention rate means something. A handful of clean results is not evidence.

Intervention rate. The share of the agent’s output a person changed before it took effect, falling and staying below the level’s figure.

Durability. The share still standing after the agreed window, which is a different number from the share accepted on the day. The gap between the two is work somebody reopened, corrected or reversed, and an agent promoted on day-one acceptance alone has been promoted on the wrong signal.

No critical errors. A critical error is an outcome that crosses a boundary defined for that agent before it runs. The policy states which outcomes qualify, what evidence is required and who may record one. None may occur during the qualifying period.

We begin from a default policy and agree the thresholds agent by agent, then load them into Control so the number being enforced is always the number that was agreed. A process-specific deviation is configured there too, with its reason attached. The figures stay in the signed policy, because the right threshold for an invoice match is not the right threshold for something whose mistakes reach a customer. What does not vary is that the thresholds are fixed before the agent runs and remain the numbers used to judge it later.

Why fix them in advance

A threshold agreed after the results are in is not a threshold. If the number can move to fit the evidence, the evidence has stopped doing any work, and the promotion decision has quietly returned to whoever is most confident in the meeting.

What moves it back down?

The same numbers. A ladder that only goes up describes a launch. The downward half is what makes it supervision, and it is often the part left unspecified.

A critical error demotes immediately. Once one is recorded under the agreed definition, Control drops the level. One is enough. It does not wait for the monthly review and is not weighed against how well the agent has otherwise been running.

A rising intervention rate freezes scope. This is not a demotion by itself, but the agent stops widening while somebody establishes the cause. Its scope does not expand while the rate of people correcting it is going up.

A model change sends it back to the test set. Changing the model invalidates the evidence attached to the previous configuration until the held-out cases have been replayed. The agent keeps its level only if the new configuration clears the same standard.

Control applies the configured response when the evidence reaches one of these conditions. It does not rely on a reviewer noticing drift during routine approval, which would inherit the weakness the levels exist to reduce.

Where do the levels live?

In Aitonomy Control. It holds the current level for each agent, the evidence behind it, the thresholds being enforced and the decision history. A person approves promotion. When recorded evidence reaches a configured demotion condition, Control freezes or reduces autonomy without waiting for the monthly review. This produces a traceable operating record that an auditor or regulator can inspect, rather than a reconstruction after the event.

Once a month, every agent gets a decision: widen, hold, step back or retire. The evidence for it is already gathered, so the meeting is about the verdict rather than about assembling the case for one. Hold should be a normal result. A portfolio where every agent is being widened should prompt a question about whether the intervention rate and durability evidence are carrying enough weight.

There is a short test for whether any of this is real in an organisation, and it does not require access to the system. Ask when an agent was last moved down a level. Ask who has the authority to do it, and whether that person needs anyone’s permission. If nobody can explain how demotion would happen, or nobody owns the authority to do it, the levels are labels.

An agent does not become trustworthy because it has been running for a while. It becomes trustworthy because it produced evidence against a standard set before anyone knew whether it would clear it, and because the same standard is still able to take the trust back.

Common questions

What are AI agent supervision levels?+
They define how much of an agent’s work a person must review. At Suggests, the agent recommends an action but changes nothing. At Drafts, it completes the work and waits for approval. At Auto, it acts within its approved scope, while escalations and a sample of its work are still reviewed.
How does an AI agent earn more autonomy?+
It earns autonomy by meeting thresholds set before it starts. The evidence includes time and volume at the current level, the rate of human intervention, whether results hold up after an agreed period and whether any critical errors occurred. Every agent starts at the most supervised level.
Can human review catch every mistake an AI agent makes?+
No. People become less alert when an automated system is usually right, a well-established effect known as automation complacency. Human review remains necessary, but it should not be the only control. Review rates, escalation rules, automatic checks and a readable record of what the agent did need to be built into the operating model.
When should an AI agent lose autonomy?+
A critical error should reduce its autonomy and increase its required oversight immediately. A rising intervention rate should stop any expansion until the cause is understood. A model change should trigger re-testing before the agent keeps its existing level. These rules are part of running an agent in production and should be enforced rather than left to someone’s memory.

Sources

  • Thomas Sheridan and William Verplank, Human and Computer Control of Undersea Teleoperators, MIT Man-Machine Systems Laboratory (1978), the ten-point scale of human supervisory control that later levels-of-automation work is built on
  • Raja Parasuraman, Thomas Sheridan and Christopher Wickens, A Model for Types and Levels of Human Interaction with Automation, IEEE Transactions on Systems, Man, and Cybernetics 30:3 (2000), separating automation into information acquisition, analysis, decision selection and action implementation
  • Linda Skitka, Kathleen Mosier and Mark Burdick, Does Automation Bias Decision-Making?, International Journal of Human-Computer Studies 51:5 (1999), comparing monitoring with and without a highly but imperfectly reliable automated aid
  • Raja Parasuraman and Dietrich Manzey, Complacency and Bias in Human Use of Automation: An Attentional Integration, Human Factors 52:3 (2010), reviewing automation complacency, automation bias and the limits of practice, training and instructions
  • The three levels, the promotion evidence and the demotion rules are Aitonomy’s own, from the autonomy policy we sign per agent and enforce in Aitonomy Control. The thresholds themselves are written in that policy and are not published here.

Mik Nijhuis

Co-founder and Chief AI Officer

Mik has architected enterprise systems for ABN AMRO, Aegon and Shell. He builds agent systems that add throughput without removing the controls complex organisations need.

Aitonomy builds and runs agent roles for recurring business work, supervised in production and measured against a baseline agreed before the first live case.What we do

All insights

Fieldwork

How an agent workforce is actually run. The roles we put into production, and the evidence behind the decisions.

Ready when you are

Bring one workflow.
Leave with a business case.

Bring us the work that slows your teams down. We will map it, count what it costs today, and select the first role only when the numbers support it.