Guides · Making the case

How to build the business case for an AI agent

Start the business case with the work as it runs today. Record its current volume, effort, quality and cost, then test how much of that work an agent could realistically take. The calculation has to include the work of reviewing and running the agent, and what the organisation intends to do with any capacity released. Otherwise the case may win approval and still be impossible to prove.

Published by Aitonomy We run agent workforces for companies. This is where we write down what that actually takes. What we do

Most guidance on AI business cases focuses on how to present the result: translate the benefits into money, show a range, calculate the payback period. That matters, and it assumes the underlying figures are sound.

An AI agent business case compares an after with a before. To build one, start with a single process and record its current volume, effort, quality, waiting time and cost. Then estimate what share of that effort suits an agent, subtract review and operating costs, state whether the result is cash, avoided cost or released capacity, and make a named owner responsible for the outcome.

Keep every estimate visible, because the same case should be used after go-live to check whether the agent was worth it.

Why does the baseline come first?

Because different measures of AI value can all be correct while describing different things.

McKinsey’s State of AI survey, published in August 2026, gathered responses from 1,719 participants. Eighty percent said AI had improved their individual productivity. Thirty-seven percent attributed at least some effect on earnings before interest and taxes to AI. Six percent met McKinsey’s much higher bar for an AI high performer: at least 5 percent of EBIT attributed to AI, and a self-described significant impact.

Those figures do not contradict one another. They measure individual experience, some enterprise financial impact, and a demanding high-performance threshold. Collapsing them into a single success rate removes the distinction a business case needs most: what exactly moved, for whom, and against which starting point.

Gartner predicts that more than 40 percent of agentic AI projects will be cancelled by the end of 2027. The reasons it gives are escalating costs, unclear business value and inadequate risk controls. Each belongs in the case before the build starts.

A 2026 California Management Review study of 24 practitioners found a similar pattern. Ten named the absence of a success metric before purchase as the main barrier to measuring financial impact. Only two used a control group or an A/B test. Sixteen used self-reported time savings, while nine said they had real confidence in that evidence.

You cannot measure improvement against a target that does not exist. Mahboubeh Cheraghian · California Management Review, August 2026

The sample was small and purposive, and it studied workplace AI rather than production agents, so it cannot establish a population rate. It does illustrate the measurement problem a business case has to solve.

What is a baseline?

A baseline is a recorded account of how the work performs before the operating model changes. It should cover a period long enough to represent normal variation. Where demand is seasonal, that may mean a full cycle. Where the work is stable and high volume, a shorter period may be enough.

The useful measures depend on the value being claimed. They usually include volume, handling effort, cycle time, queue age, rework, rejection, exceptions, escalations, and the cost of the people and systems involved. Averages alone are rarely enough. If the difficult fifth of cases carries half the effort, average handling time hides the part that will remain.

The hourly cost also determines what the number means. A fully loaded rate can describe capacity-equivalent benefit, but it does not prove a cash saving, and the difference between cash, avoided cost and released capacity decides what the case may actually claim.

Most organisations will be missing some of these measures. The case can still be built, provided it is clear about which figures come from evidence and which remain assumptions.

Measured baselineModelled baseline
Starts fromThe organisation’s own recorded historyAvailable records plus a benchmark or explicit assumption
AnswersWhat the work demonstrably didWhat the work probably did, within a stated range
CarriesData-quality limits and normal variationAssumptions, ranges and a confidence level
Fails whenMissing records are treated as completeAn estimate is presented as a measurement

A historical baseline may still be recoverable after an agent goes live, but it may no longer be comparable. The process, volumes, team and behaviour may already have changed. Recording the baseline first reduces hindsight bias and preserves the conditions the decision was made against.

The rule we hold to

A modelled baseline honestly labelled beats a fabricated one silently assumed. Keep the label attached when the number is reused, and if an estimate is later replaced by a measurement, the record should show what changed and why.

What belongs on the one page?

We keep the business case to six fields and agree them with the process owner before anything is built. Limiting it to one page keeps the focus on the decision rather than the proposed technology.

SIX FIELDS, AND WHERE THE BASELINE COMES FROM THE SIX FIELDS, ON ONE PAGE The work and business outcome the problem, the result, the boundary Quality and consequence where approval stays mandatory The baseline volume, effort, quality and cost today Investment and operating cost the build, and what continues after it The target and measurement plan what improves, and how it is attributed The named owner and conversion plan one person, and what the capacity becomes WHERE THE BASELINE COMES FROM MEASURED Your own recorded history volume, effort, cycle time, rework and exceptions, over a representative period MODELLED, ONE OF THREE ROUTES A comparable process elsewhere An external benchmark An assumption the board signs
The other five fields depend on the baseline. It is also the only field you can arrive at two ways, so the page records which one was used.

The work and business outcome. The operational problem, the desired result, and the boundary of the agent’s proposed responsibility. Start from the outcome, then decide what belongs with people, an agent, and existing automation.

Quality and consequence. The standard the work must meet, what happens when it is wrong, and where human approval or escalation remains necessary.

The baseline. Current volume, effort, quality, waiting time and cost. Each important number is labelled measured or modelled, with the period and source attached.

Investment and operating cost. Build and integration, along with the costs that continue afterwards: model and tool use, human review, exception handling, monitoring, governance, maintenance, process change, and the cost of stopping or correcting failures.

The target and measurement plan. The primary outcome that must improve, the range expected, how it will be attributed, and the evidence required to proceed, narrow the scope or stop. A staggered rollout, matched comparison or holdout is stronger than a simple before-and-after where the process permits one.

The named owner and conversion plan. One person on the customer side is accountable for the result, and for deciding what released capacity will become. A department cannot redeploy time or answer for a missed target.

The costs, expected value and decision gate should be readable together. The target must say how it will be measured, and any capacity claim must explain how the released time will be used.

Why does theoretical capacity overstate value?

Theoretical capacity is usually calculated as volume multiplied by average handling time multiplied by the share of cases an agent might take. That gives a useful ceiling, not the expected benefit. Turning the ceiling into a usable estimate takes three further adjustments.

Coverage must be weighted by effort. Agents often start with the work that is easiest to recognise and verify, which can remove a large share of cases while leaving a much larger share of effort.

Accounts payable provides a useful illustration. Ardent Partners’ 2025 benchmark, published by Medius, reports an average invoice exception rate of 22 percent. For the calculation below, assume an exception takes four times the effort of a straight-through invoice.

Straight throughExceptions
Share of cases78 percent22 percent
Assumed relative effort per case14
Share of total effort47 percent53 percent

In this simplified example, taking every straight-through invoice would remove 78 percent of case volume but only 47 percent of effort. The exception multiplier is illustrative rather than an Ardent Partners finding, and an organisation should replace it with its own handling data before using the calculation in a case.

Review and residual work remain. A person may still need to verify results, handle exceptions, and keep enough context to intervene. Bainbridge described this operating problem in Ironies of Automation: automating routine work can leave people with the abnormal conditions and a new monitoring duty. She wrote about industrial automation rather than AI agents, and the same workload question belongs in this case. Human oversight takes time, and that time needs to be measured.

Released capacity has to be usable. Thirty minutes removed from one queue may improve service immediately. Two minutes scattered across fifteen people may never appear as cash or additional output. The business case must say which one it expects.

Customer support offers another example. A deflected case ended without reaching a person. A resolved case ended because the customer’s problem was solved. Counting both as the same outcome turns a system event into a value claim without checking what happened.

What number should you plan on?

Use the available evidence to plan a range, rather than starting from a universal discount.

Potential released capacity

Current case volume × current effort × effort-weighted coverage, minus review, rework and residual exception effort.

Then classify the result, because the four lines are not interchangeable.

ClassificationWhat the case may count
Cashable savingSpend that will actually leave the cost base
Avoided costA specific hire, contractor or service expense no longer required
Released capacityCapacity-equivalent hours, kept separate from cash until they produce a measured outcome
Operational valueQuality, availability, throughput or risk improvement, with its own baseline

Finally subtract the full cost of ownership: build, integration, usage, review, control, maintenance and change. If a material input is unknown, show it as a range and make closing that range a decision gate.

Workday’s January 2026 research shows why review belongs in the calculation. In a survey of 3,200 active workplace AI users, nearly 40 percent of reported time savings went back into rework, and only 14 percent consistently reported a clear positive net result. The evidence is self-reported and concerns workplace AI rather than production agents, so it cannot supply a discount for a specific process. It does show the risk of treating gross time saved as value without checking the rework behind it.

The low, expected and high cases should vary the assumptions that matter: effort-weighted coverage, verification cost, exception load, operating cost, and the share of capacity that can be converted. Use the scenarios to expose uncertainty, not to give the same unsupported estimate three different labels.

If the range is too wide to support a decision, measure what is missing before the build. Depending on the result, the proposed role may need to be narrowed or stopped.

What if the value is not hours?

Hours are commonly used because they are easy to count, although they are often not the strongest measure of value.

Quality. Where a process has a measurable error or rework rate, value may come from fewer corrections, credit notes, complaints or reopened cases. An agent may apply a rule more consistently, and the case should treat that as a target to prove rather than a guaranteed property.

Availability. Work that waits overnight or behind a queue carries a service cost. The baseline is queue age, time to first response, and the share of cases breaching a service level. The result is measured in cycle time and service performance, not automatically in headcount.

Capacity without hiring. Where demand is growing, holding the service level while volume rises may avoid a planned hire or external service cost. The case needs the volume trend, the staffing trigger, and the point at which the expense would otherwise occur.

The work that remains. Removing routine administration can give people more capacity for exceptions and judgement. That is valuable only when the organisation identifies the work that will receive the time and measures what improves.

Value lineBaselineEvidence after go-live
QualityError, rework and rejection ratesCorrections, complaints and credit notes
AvailabilityQueue age and response timeCycle time and service-level performance
Avoided hiringVolume trend and staffing triggerService maintained without the planned expense
Residual workBacklog, overtime and exception ageDifficult cases handled faster or better

Faster handling, released capacity and avoided hiring may all describe the same underlying change, so counting each one separately would overstate the case. Choose one primary financial claim, keep the operational measures that prove it, and leave money off other expected effects unless they can be measured independently.

What does a modelled baseline look like?

In one anonymised case, no measured handling-time baseline was available for the work in question, and a decision was needed before one could be produced.

We combined the recorded volumes with assumptions about handling time across the candidate processes. The result was an estimate of roughly 5,800 capacity-equivalent hours a year, which we treated as a starting point rather than a saving.

Redeployability was the one factor with firmer evidence behind it, because the staffing pattern showed how much released time could realistically be taken up. We put that at about two thirds, which produced an expected case of roughly 3,750 capacity-equivalent hours.

Coverage and verification had not yet been measured on real cases. Instead of assigning them preset discounts, the case carried a range of 2,000 to 6,000 hours. The upper end was slightly above the starting estimate because handling time itself remained uncertain. Every figure was labelled a hypothesis.

The range identified what the assessment still had to measure: handling effort by case type, achievable coverage on held-out work, and the time required to verify a result. As those findings came in, the range would narrow and the build decision could change.

Used this way, a modelled baseline makes uncertainty testable without presenting it as fact.

How does the case survive after approval?

The baseline, target and measurement plan stay attached to the work after go-live. Each reporting period compares live evidence with the conditions recorded before the decision, and feeds the operating record.

A simple before-and-after comparison may be enough for a stable, high-volume process. Where seasonality, staffing or demand changes could explain the result, use stronger attribution if practical: a staggered rollout, a matched team, a holdout queue, or a documented trend comparison.

The numbers can change when reality changes, provided the change stays traceable. If the baseline, scope or target is revised, the record shows the old value, the new one, the reason, and the person who approved it.

A business case must leave room for a no. If the evidence cannot support an attainable range, the full operating cost removes the value, the consequence of failure is too high, or nobody owns the outcome, the right decision may be to prepare first, narrow the role, or leave the work with people. Unlike a proposal, the case has to support a decision not to proceed.

Where do the Scan and Aitonomy Control fit?

The Scan and our team establish the baseline before the build decision. With approved access to the systems involved, we reconstruct the recorded process, test the quality of the evidence, and investigate the gaps with the people who own the work. Together we label what is measured, what is modelled, and what must still be proven.

The result is a one-page case that shows the open assumptions and supports one of four decisions.

DecisionWhat it means
BuildThe evidence supports proceeding with the proposed role and controls
Prepare firstPromising, but the data, process, ownership or controls need work before building
Narrow the scopeOnly part of the proposed role is sufficiently valuable, feasible and safe
StopThe expected value, feasibility or risk does not justify proceeding

If the work proceeds, Aitonomy Control provides the runtime, monitoring, governance and traceability in production. Its operating record is what compares live performance against the agreed baseline and target.

You do not need a Scan to test the first version. Take the process already under consideration and try to complete the six fields. If you cannot fill in the baseline, the total operating cost or the named owner, you know where to start. Producing the evidence behind those fields is where we start.

A business case should be built before the agent and kept in use after go-live, so the expected value can be checked against what actually happened.

Common questions

How do you calculate the ROI of an AI agent?+
Start with work that is suitable for an AI agent, then record the current volume, effort, quality, waiting time and cost. Estimate how much of the effort the agent can take on and subtract review and operating costs. Finally, name the value honestly as cash saved, cost avoided or capacity released. They are not the same thing.
What if you have no baseline data?+
Build a modelled baseline from a comparable process, an external benchmark or a stated assumption, and label it clearly. Keep the assumptions visible so they can be challenged and replaced as real data arrives. An estimate is useful when everyone can see that it is an estimate.
Why can theoretical time savings overstate the value?+
Case volume and human effort are not distributed evenly. An agent may clear many straightforward cases while the team keeps the slower exceptions. That can remove a large share of the volume without removing the same share of the work. Value should be based on the effort that actually disappears or becomes usable elsewhere.
What belongs in an AI agent business case?+
Put six things on one page: the work and intended outcome, the required quality and consequence of error, the baseline, the investment and operating cost, the target and measurement plan, and a named owner. If the case depends on released capacity, say how that capacity will be used. The evidence should support four possible decisions: build, prepare first, narrow the scope or stop. If the decision is build, keep the case in use while the agent runs in production.
Is released capacity the same as a cost saving?+
No. Released capacity means people have time available for other work. It becomes a cash saving only when spend leaves the cost base, and an avoided cost when a planned hire or expense is no longer needed. Keep those three forms of value separate in the case.

Sources

Joep

Co-founder and CEO

Joep spent two decades at DEPT®, latterly as SVP Clients, working with enterprise leadership on how their organisations run without losing governance or cost control.

Aitonomy builds and runs agent roles for recurring business work, supervised in production and measured against a baseline agreed before the first live case.What we do

All insights

Fieldwork

How an agent workforce is actually run. The roles we put into production, and the evidence behind the decisions.

Ready when you are

Bring one workflow.
Leave with a business case.

Bring us the work that slows your teams down. We will map it, count what it costs today, and select the first role only when the numbers support it.