Most guidance on AI business cases focuses on how to present the result: translate the benefits into money, show a range, calculate the payback period. That matters, and it assumes the underlying figures are sound.
An AI agent business case compares an after with a before. To build one, start with a single process and record its current volume, effort, quality, waiting time and cost. Then estimate what share of that effort suits an agent, subtract review and operating costs, state whether the result is cash, avoided cost or released capacity, and make a named owner responsible for the outcome.
Keep every estimate visible, because the same case should be used after go-live to check whether the agent was worth it.
Why does the baseline come first?
Because different measures of AI value can all be correct while describing different things.
McKinsey’s State of AI survey, published in August 2026, gathered responses from 1,719 participants. Eighty percent said AI had improved their individual productivity. Thirty-seven percent attributed at least some effect on earnings before interest and taxes to AI. Six percent met McKinsey’s much higher bar for an AI high performer: at least 5 percent of EBIT attributed to AI, and a self-described significant impact.
Those figures do not contradict one another. They measure individual experience, some enterprise financial impact, and a demanding high-performance threshold. Collapsing them into a single success rate removes the distinction a business case needs most: what exactly moved, for whom, and against which starting point.
Gartner predicts that more than 40 percent of agentic AI projects will be cancelled by the end of 2027. The reasons it gives are escalating costs, unclear business value and inadequate risk controls. Each belongs in the case before the build starts.
A 2026 California Management Review study of 24 practitioners found a similar pattern. Ten named the absence of a success metric before purchase as the main barrier to measuring financial impact. Only two used a control group or an A/B test. Sixteen used self-reported time savings, while nine said they had real confidence in that evidence.
You cannot measure improvement against a target that does not exist. Mahboubeh Cheraghian · California Management Review, August 2026
The sample was small and purposive, and it studied workplace AI rather than production agents, so it cannot establish a population rate. It does illustrate the measurement problem a business case has to solve.
What is a baseline?
A baseline is a recorded account of how the work performs before the operating model changes. It should cover a period long enough to represent normal variation. Where demand is seasonal, that may mean a full cycle. Where the work is stable and high volume, a shorter period may be enough.
The useful measures depend on the value being claimed. They usually include volume, handling effort, cycle time, queue age, rework, rejection, exceptions, escalations, and the cost of the people and systems involved. Averages alone are rarely enough. If the difficult fifth of cases carries half the effort, average handling time hides the part that will remain.
The hourly cost also determines what the number means. A fully loaded rate can describe capacity-equivalent benefit, but it does not prove a cash saving, and the difference between cash, avoided cost and released capacity decides what the case may actually claim.
Most organisations will be missing some of these measures. The case can still be built, provided it is clear about which figures come from evidence and which remain assumptions.
| Measured baseline | Modelled baseline | |
|---|---|---|
| Starts from | The organisation’s own recorded history | Available records plus a benchmark or explicit assumption |
| Answers | What the work demonstrably did | What the work probably did, within a stated range |
| Carries | Data-quality limits and normal variation | Assumptions, ranges and a confidence level |
| Fails when | Missing records are treated as complete | An estimate is presented as a measurement |
A historical baseline may still be recoverable after an agent goes live, but it may no longer be comparable. The process, volumes, team and behaviour may already have changed. Recording the baseline first reduces hindsight bias and preserves the conditions the decision was made against.
The rule we hold to
A modelled baseline honestly labelled beats a fabricated one silently assumed. Keep the label attached when the number is reused, and if an estimate is later replaced by a measurement, the record should show what changed and why.
What belongs on the one page?
We keep the business case to six fields and agree them with the process owner before anything is built. Limiting it to one page keeps the focus on the decision rather than the proposed technology.
The work and business outcome. The operational problem, the desired result, and the boundary of the agent’s proposed responsibility. Start from the outcome, then decide what belongs with people, an agent, and existing automation.
Quality and consequence. The standard the work must meet, what happens when it is wrong, and where human approval or escalation remains necessary.
The baseline. Current volume, effort, quality, waiting time and cost. Each important number is labelled measured or modelled, with the period and source attached.
Investment and operating cost. Build and integration, along with the costs that continue afterwards: model and tool use, human review, exception handling, monitoring, governance, maintenance, process change, and the cost of stopping or correcting failures.
The target and measurement plan. The primary outcome that must improve, the range expected, how it will be attributed, and the evidence required to proceed, narrow the scope or stop. A staggered rollout, matched comparison or holdout is stronger than a simple before-and-after where the process permits one.
The named owner and conversion plan. One person on the customer side is accountable for the result, and for deciding what released capacity will become. A department cannot redeploy time or answer for a missed target.
The costs, expected value and decision gate should be readable together. The target must say how it will be measured, and any capacity claim must explain how the released time will be used.
Why does theoretical capacity overstate value?
Theoretical capacity is usually calculated as volume multiplied by average handling time multiplied by the share of cases an agent might take. That gives a useful ceiling, not the expected benefit. Turning the ceiling into a usable estimate takes three further adjustments.
Coverage must be weighted by effort. Agents often start with the work that is easiest to recognise and verify, which can remove a large share of cases while leaving a much larger share of effort.
Accounts payable provides a useful illustration. Ardent Partners’ 2025 benchmark, published by Medius, reports an average invoice exception rate of 22 percent. For the calculation below, assume an exception takes four times the effort of a straight-through invoice.
| Straight through | Exceptions | |
|---|---|---|
| Share of cases | 78 percent | 22 percent |
| Assumed relative effort per case | 1 | 4 |
| Share of total effort | 47 percent | 53 percent |
In this simplified example, taking every straight-through invoice would remove 78 percent of case volume but only 47 percent of effort. The exception multiplier is illustrative rather than an Ardent Partners finding, and an organisation should replace it with its own handling data before using the calculation in a case.
Review and residual work remain. A person may still need to verify results, handle exceptions, and keep enough context to intervene. Bainbridge described this operating problem in Ironies of Automation: automating routine work can leave people with the abnormal conditions and a new monitoring duty. She wrote about industrial automation rather than AI agents, and the same workload question belongs in this case. Human oversight takes time, and that time needs to be measured.
Released capacity has to be usable. Thirty minutes removed from one queue may improve service immediately. Two minutes scattered across fifteen people may never appear as cash or additional output. The business case must say which one it expects.
Customer support offers another example. A deflected case ended without reaching a person. A resolved case ended because the customer’s problem was solved. Counting both as the same outcome turns a system event into a value claim without checking what happened.
What number should you plan on?
Use the available evidence to plan a range, rather than starting from a universal discount.
Potential released capacity
Current case volume × current effort × effort-weighted coverage, minus review, rework and residual exception effort.
Then classify the result, because the four lines are not interchangeable.
| Classification | What the case may count |
|---|---|
| Cashable saving | Spend that will actually leave the cost base |
| Avoided cost | A specific hire, contractor or service expense no longer required |
| Released capacity | Capacity-equivalent hours, kept separate from cash until they produce a measured outcome |
| Operational value | Quality, availability, throughput or risk improvement, with its own baseline |
Finally subtract the full cost of ownership: build, integration, usage, review, control, maintenance and change. If a material input is unknown, show it as a range and make closing that range a decision gate.
Workday’s January 2026 research shows why review belongs in the calculation. In a survey of 3,200 active workplace AI users, nearly 40 percent of reported time savings went back into rework, and only 14 percent consistently reported a clear positive net result. The evidence is self-reported and concerns workplace AI rather than production agents, so it cannot supply a discount for a specific process. It does show the risk of treating gross time saved as value without checking the rework behind it.
The low, expected and high cases should vary the assumptions that matter: effort-weighted coverage, verification cost, exception load, operating cost, and the share of capacity that can be converted. Use the scenarios to expose uncertainty, not to give the same unsupported estimate three different labels.
If the range is too wide to support a decision, measure what is missing before the build. Depending on the result, the proposed role may need to be narrowed or stopped.
What if the value is not hours?
Hours are commonly used because they are easy to count, although they are often not the strongest measure of value.
Quality. Where a process has a measurable error or rework rate, value may come from fewer corrections, credit notes, complaints or reopened cases. An agent may apply a rule more consistently, and the case should treat that as a target to prove rather than a guaranteed property.
Availability. Work that waits overnight or behind a queue carries a service cost. The baseline is queue age, time to first response, and the share of cases breaching a service level. The result is measured in cycle time and service performance, not automatically in headcount.
Capacity without hiring. Where demand is growing, holding the service level while volume rises may avoid a planned hire or external service cost. The case needs the volume trend, the staffing trigger, and the point at which the expense would otherwise occur.
The work that remains. Removing routine administration can give people more capacity for exceptions and judgement. That is valuable only when the organisation identifies the work that will receive the time and measures what improves.
| Value line | Baseline | Evidence after go-live |
|---|---|---|
| Quality | Error, rework and rejection rates | Corrections, complaints and credit notes |
| Availability | Queue age and response time | Cycle time and service-level performance |
| Avoided hiring | Volume trend and staffing trigger | Service maintained without the planned expense |
| Residual work | Backlog, overtime and exception age | Difficult cases handled faster or better |
Faster handling, released capacity and avoided hiring may all describe the same underlying change, so counting each one separately would overstate the case. Choose one primary financial claim, keep the operational measures that prove it, and leave money off other expected effects unless they can be measured independently.
What does a modelled baseline look like?
In one anonymised case, no measured handling-time baseline was available for the work in question, and a decision was needed before one could be produced.
We combined the recorded volumes with assumptions about handling time across the candidate processes. The result was an estimate of roughly 5,800 capacity-equivalent hours a year, which we treated as a starting point rather than a saving.
Redeployability was the one factor with firmer evidence behind it, because the staffing pattern showed how much released time could realistically be taken up. We put that at about two thirds, which produced an expected case of roughly 3,750 capacity-equivalent hours.
Coverage and verification had not yet been measured on real cases. Instead of assigning them preset discounts, the case carried a range of 2,000 to 6,000 hours. The upper end was slightly above the starting estimate because handling time itself remained uncertain. Every figure was labelled a hypothesis.
The range identified what the assessment still had to measure: handling effort by case type, achievable coverage on held-out work, and the time required to verify a result. As those findings came in, the range would narrow and the build decision could change.
Used this way, a modelled baseline makes uncertainty testable without presenting it as fact.
How does the case survive after approval?
The baseline, target and measurement plan stay attached to the work after go-live. Each reporting period compares live evidence with the conditions recorded before the decision, and feeds the operating record.
A simple before-and-after comparison may be enough for a stable, high-volume process. Where seasonality, staffing or demand changes could explain the result, use stronger attribution if practical: a staggered rollout, a matched team, a holdout queue, or a documented trend comparison.
The numbers can change when reality changes, provided the change stays traceable. If the baseline, scope or target is revised, the record shows the old value, the new one, the reason, and the person who approved it.
A business case must leave room for a no. If the evidence cannot support an attainable range, the full operating cost removes the value, the consequence of failure is too high, or nobody owns the outcome, the right decision may be to prepare first, narrow the role, or leave the work with people. Unlike a proposal, the case has to support a decision not to proceed.
Where do the Scan and Aitonomy Control fit?
The Scan and our team establish the baseline before the build decision. With approved access to the systems involved, we reconstruct the recorded process, test the quality of the evidence, and investigate the gaps with the people who own the work. Together we label what is measured, what is modelled, and what must still be proven.
The result is a one-page case that shows the open assumptions and supports one of four decisions.
| Decision | What it means |
|---|---|
| Build | The evidence supports proceeding with the proposed role and controls |
| Prepare first | Promising, but the data, process, ownership or controls need work before building |
| Narrow the scope | Only part of the proposed role is sufficiently valuable, feasible and safe |
| Stop | The expected value, feasibility or risk does not justify proceeding |
If the work proceeds, Aitonomy Control provides the runtime, monitoring, governance and traceability in production. Its operating record is what compares live performance against the agreed baseline and target.
You do not need a Scan to test the first version. Take the process already under consideration and try to complete the six fields. If you cannot fill in the baseline, the total operating cost or the named owner, you know where to start. Producing the evidence behind those fields is where we start.
A business case should be built before the agent and kept in use after go-live, so the expected value can be checked against what actually happened.
Common questions
How do you calculate the ROI of an AI agent?+
What if you have no baseline data?+
Why can theoretical time savings overstate the value?+
What belongs in an AI agent business case?+
Is released capacity the same as a cost saving?+
Sources
- McKinsey & Company, The state of AI in 2026: On the road to ROI (August 2026), a survey of 1,719 participants, on individual productivity, EBIT attribution and the high-performer threshold
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 2025), naming escalating costs, unclear business value and inadequate risk controls
- Mahboubeh Cheraghian, Beyond AI Licenses: How to Measure Whether Your AI Is Actually Working, California Management Review (August 2026), a structured study of 24 practitioners on measurement barriers and attribution
- Workday, Beyond Productivity: Measuring the Real Value of AI (January 2026), a survey of 3,200 active workplace AI users, fielded in November 2025
- Ardent Partners, Accounts Payable Metrics that Matter in 2025, on an average invoice exception rate of 22 percent. The four-times effort multiplier used above is illustrative and is not an Ardent Partners finding
- Lisanne Bainbridge, Ironies of Automation, Automatica 19:6 (1983), on the residual work and monitoring duty left to people after automation
- The six-field business case, the measured and modelled baseline labels, and the anonymised example are part of Aitonomy’s Scan method. The example has its organisation, sector and commercial terms removed.