Skip to content
AZGARD
strategy

How to measure AI ROI when hours saved are not money

Measure AI ROI by baselining the process first: count, time per unit, a named owner, a date, and the one downstream change that turns hours into money.

Angus McDonald · 18 Aug 2026 · 18 min read

How to measure AI ROI when hours saved are not money

Measure AI ROI on one process before you build, using an Azgard AI Baseline Card: how often it runs in a week, how long one run takes, who measured it and when. Re-take at 30 and 90 days. Then check the one downstream thing that had to change for the saving to become money - a role you did not backfill, a deadline you stopped missing. If nothing downstream moved, the hours are real and the money is imaginary.

This page assumes you have already mapped the work, which is where a defensible baseline comes from. A defensible baseline still is not money.

Why AI ROI is hard to measure in a small business

Measuring AI ROI is hard in a small business for three reasons, none of them about AI: attribution, missing baselines, and no cost centre.

Attribution first. Deloitte surveyed 1,854 executives across Europe and the Middle East between 15 August and 5 September 2025, and quoted one saying they "only managed to get a ballpark estimate of the benefits because it was hard to separate the gains from AI initiatives from those of other initiatives". Enterprise, two regions, not a small-business sample - and the problem is worse in a ten-person business, because you changed three things that quarter and the same two people did all of them.

Then the missing baseline. Most small businesses cannot say what a process cost before the AI arrived, so there is nothing to compare against. IBM's study with Oxford Economics, covering 2,000 CEOs in 33 countries between February and April 2025, found only 25% of AI initiatives have delivered expected ROI over the last few years: enterprise again, and as much a verdict on measurement discipline as on technology. MIT Sloan Management Review, on interviews with more than 30 leaders, names the mechanism: "few companies apply the same financial discipline to artificial intelligence as they would to a new factory or piece of machinery".

And the cost centre. A van has a purchase price, a depreciation schedule and a job it does. An AI subscription has a monthly charge spread across five people who use it differently and none of whom bills time to it. The spend is visible; the unit it attaches to is not.

Hours saved are not money saved until something changes downstream

Hours saved by an AI system are not money saved until something downstream of those hours changes: a role you did not backfill, work you took on that you would have turned away, overtime you stopped paying, a deadline you stopped missing. If none of those move, the hours are real and the money is imaginary.

Not a new observation, and this page will not pretend otherwise. CIO ran the argument in March 2026 under the headline "why tools don't pay off until work changes". The consensus is right. What it lacks is an instrument: the downstream test is nearly always applied afterwards, as a post-mortem, rather than written down in advance as a commitment. Meanwhile every failing ROI method fails at the same joint: it multiplies hours saved by an hourly rate and calls the product money. That multiplication is an assumption, not a measurement.

What you measuredWhat has to change downstream for it to be moneyHow you check that it did
Quoting time fell from 3.5 hours to just over 2More quotes go out, or the second estimator is never hiredQuotes issued per week; next year's headcount line
Two hours a week off bookkeepingContracted hours are cut, or that person moves onto collectionsInvoiced bookkeeping hours; debtor days
Proposal first drafts come back in minutesProposals go out same day, and deal count or win rate movesEnquiry-to-proposal days; proposals per month
An assistant drafts 40% of support repliesThe support backfill is never hired, or churn movesOpen roles not filled; first-response time; cancellations
Fewer hours on the monthly reportThe report lands early enough to change a decisionThe date it is issued; one decision that moved

The hours themselves are usually self-reported, and self-reports about AI are measurably unreliable. METR ran a randomised controlled trial: 16 experienced open-source developers, 246 real tasks in repositories they already maintained, Cursor Pro with Claude 3.5 and 3.7 Sonnet. "Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%". METR later quantified the gap: people "overestimated AI's effect on their time spent on tasks by 40 percentage points on average". Its 2026 survey of 349 technical workers recorded a median self-reported speed change of 3x, with the write-up conceding that "survey results are not necessarily grounded in reality".

METR's limits deserve as much weight as its headline: 16 developers, all experienced, all on codebases they knew well, and the paper states plainly that it does not claim its "developers or repositories represent a majority or plurality of software development work". What it measures is self-report, not AI.

Workday's "Beyond Productivity" research, run by Hanover Research across 3,200 employees at large enterprises and announced on 14 January 2026, puts the gap closer to the P&L: 85% of respondents saved one to seven hours a week, and only 14% consistently got clear, positive net outcomes from AI. The hours were real; the money mostly was not. Two caveats: Workday sells AI-powered HR software, so this is vendor-sponsored research, and the sample is large-enterprise employees, not small businesses. Neither caveat means the gains are fake. They are real, and the best-measured ones are counted per unit of work.

Azgard's own number sits at the same joint between hours and money. At Apex Signage, a Sydney signage manufacturer, quoting went from about 3.5 hours per quote to a little over two, against a baseline of roughly 12 quotes a week. The arithmetic is sitting right there and we are not going to do it: 1.5 hours times 12 quotes times somebody's notion of an estimator's hourly cost is a number nobody measured. The hours are real and checkable in Apex's own job records. What they become in money depends on the downstream field, and that figure belongs on Apex's books rather than the marketing page of an AI consultancy in Sydney.

That number also fails two fields of our own Card. Mapping the work established the baseline, which is where 3.5 hours came from, but nobody wrote down who took the timings or on what date. The measurer and the date, fields four and five, are blank. A page arguing that an unattributed baseline is not evidence has to say that about its own best number: the Apex figure is a real measurement with no measurer and no date attached to it.

The Azgard Three-Question Vendor Filter: what to ask before you sign

The Azgard Three-Question Vendor Filter is three questions to put to any AI vendor, including this one, before you sign anything: ask what number they will measure before they build, who takes the measurement, and the date it gets taken again. If they can only tell you the number afterwards, they are selling a story. It is the same discipline Azgard applies to its own builds, the six-field Azgard AI Baseline Card, turned around and pointed at the person selling to you, which is why a vendor who cannot answer it is telling you their own work will not be measurable.

The Azgard AI Baseline Card: six fields to fill in before anyone builds

The Azgard AI Baseline Card is a six-field record, filled in for one process, before anything gets built. One card per process. It exists because "we will see if it works" is not a measurement plan, and because the field you cannot fill in is the one that predicts the project.

  1. The process. One named workflow with a start and an end, not a department and not a tool. "Producing a quote from an enquiry", not "estimating".
  2. The count. How many times it ran over one week, from records you already keep.
  3. The time per unit. How long one run takes, measured on real runs, recorded as a range, not a tidy average.
  4. The measurer. The name of the person who took the count and the timings. Not a team, not a system. A person.
  5. The date. The day it was taken, so the later comparison is between two dated numbers rather than a number and a memory.
  6. The downstream change. The one thing that has to change for the saving to become money, written as a sentence somebody can be held to. "We do not replace the estimator who leaves in March." "We stop declining Friday jobs."

The same person re-takes the same six fields at 30 days and again at 90 days.

Field six is the whole argument. Its retrospective version is already common advice: after the project, ask whether the time got redeployed. Asking afterwards is a post-mortem. Writing it down beforehand, with a name against it and a date to check it, is a commitment, and commitments change what gets built. Field six kills projects at the whiteboard, the cheapest place to kill them. If nobody in the room can say what would change downstream, the build is a preference rather than an investment.

Fields four and five are the same discipline as an owner and an expiry date on a company brain: a measurement with no name against it decays exactly like a document with no owner.

How to baseline an AI project in one week without new tooling

Baselining an AI project takes about a week and needs no new tooling, because the counts already sit in systems you pay for. Instrument nothing; count what is already there.

  • Volume, from records you already keep. Export the last eight weeks of jobs, tickets, quotes or invoices and count the rows. Eight weeks absorbs one quiet week without becoming a research project.
  • Time per unit, from a stopwatch, five times. Ask the person who does the work to time five real runs on their phone. The range matters more than the average, because the slow runs are the ones AI changes most.
  • Queue time, from timestamps you already have. The gap between an enquiry arriving and someone logging it is two date fields in an inbox or a CRM. In most small businesses the queue is longer than the task.
  • Rework, from the revision trail. Count how many outputs came back for a second pass over those same eight weeks. Without the before number you cannot tell later whether AI reduced rework or manufactured it.

Raid your own systems before buying anything that promises measurement. Job management tools, invoicing systems and shared inboxes already log every transaction they process, and in most small businesses the count you need is one export away.

If the AI is already running, none of this is closed to you. The same records give a retrospective baseline: count the eight weeks before you switched it on, take the timings now, and record on the card that the before-number was reconstructed rather than observed. That is weaker evidence than a measurement taken in advance, and it is the honest version of what most businesses actually have.

Rough but honest beats precise and absent. A count taken by hand from eight weeks of job records, with a name and a date on it, survives a room full of sceptics. A dashboard figure nobody can trace does not.

How to measure AI time savings you can actually defend

AI time savings are defensible when measured per unit of work and fragile when reported as hours per week. Hours per week moves with volume, staffing and season, so a busy quarter buries a real improvement and a quiet one invents a fake improvement. Time per unit holds still while everything around it moves.

Per unit is what the best-run studies measure. Brynjolfsson, Li and Raymond studied 5,179 customer support agents and found AI assistance "increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers". Note the unit, and note the distribution: the gain went to the least experienced people, which changes who you should be measuring.

The preregistered Harvard Business School study of 758 Boston Consulting Group consultants by Dell'Acqua and colleagues found the same shape with a sharper edge. Inside the AI's capability frontier, consultants completed 12.2% more tasks, 25.1% faster, at higher quality. On a task outside it, "subjects using AI were 19% less likely to produce correct solutions". Same tool, same people, and the sign flips with task selection, which is why the Baseline Card covers one named process rather than a tool. ROI is not a property of the tool.

The practical version is unglamorous. Same person, same kind of work, five timed runs before and five at each re-take, range recorded alongside the average. If the measurer changes between the baseline and the 30-day check, you have two numbers rather than a comparison.

The AI KPIs worth tracking, and the soft benefits worth counting but not pricing

Four numbers are enough to track an AI project in a small business: volume, time per unit, rework rate, and the one named downstream metric from the Baseline Card.

  • Volume. How many times the process ran. If time per unit falls and volume falls with it, you found a slow month, not a productivity gain.
  • Time per unit. The defensible measure of the saving, taken the same way each time.
  • Rework rate. The share of outputs needing a second pass. Workday's research found nearly 40% of AI time savings are lost to rework, which is vendor-sponsored and drawn from large enterprises, so read it as a warning rather than a benchmark for your own business. Rework is the one KPI that can invert a project's result, and the one nobody baselines.
  • The named downstream metric. Whatever field six says: quotes issued, roles not filled, overtime hours, debtor days, jobs declined.

Two numbers people ask for that this method deliberately will not produce. Cost per task is time per unit multiplied by an hourly rate, which is the same multiplication that turns real hours into imaginary money: fine for comparing two ways of doing one job, worthless as evidence of return, because it inherits the redeployment assumption whole. And the ROI of an AI automation is not a separate calculation from the ROI of anything else. An automation moves time per unit and volume, so it goes on a Baseline Card like any other change, and it still has to name the one thing downstream that turns the saving into money.

IBM splits AI KPIs into hard ROI (labour cost, operational efficiency, conversion, revenue) and soft ROI (satisfaction, retention, decision making), and says the soft ones "are often measured with surveys and qualitative research initiatives". Deloitte says why that resists a fix: AI "frequently delivers outcomes that matter" but they "are hard to monetize".

So count the soft event and never price it. "The owner stopped doing quotes on Sunday" is a countable event with a date on it. Attaching an hourly rate to those Sundays and adding the total to a business case is where a defensible measurement turns into a sales deck. Soft benefits are often the real reason a business keeps a system running. They are still not evidence of return.

Why an AI ROI calculator returns the number you fed it

An AI ROI calculator returns the number you fed it, because every input except the licence fee is an assumption. Hours saved is self-reported, the hourly rate is a choice and adoption is a guess. The redeployment of the saved hours, the step that makes the arithmetic true, is not an input at all.

Think Technology Australia, an Australian IT provider, states the consensus method clearly: "Before deploying any AI tool, record how long the target task currently takes per week and the hourly cost of the person doing it. Review those same numbers at 90 days." The instruction is fine. It is the worked example that gives the game away. Their published example has a tool saving five hours a week for one staff member at AU$45 an hour, for an annual saving of AU$11,700 against AU$1,800 of tool cost. They publish the formula as well: "ROI: ($11,700 minus $1,800) divided by $1,800, times 100 equals 483%."

Run that division and it returns 550%. Nor is it a rounding slip, because 483% would need a saving of AU$10,500 or a tool cost of about AU$2,006, different numbers entirely. This is not a gotcha: the author did the harder thing by publishing the formula, and without it there would be nothing to check at all. It is simply a worked ROI example that does not survive its own arithmetic. These figures were checked on 18 August 2026; if the page is corrected upstream, this paragraph is the stale one.

Elsewhere on the same page, 7.5 hours a week recovered is "roughly $15,000 per year in capacity that can be redeployed or avoided as a future hire", in Australian dollars. "Can be redeployed" is the joint: the whole return sits inside a modal verb, stated honestly and never tested. Field six of the Baseline Card exists to turn that into "will be redeployed, here is who says so, and here is the date we check".

CIO named the same failure and called out calculators "that treat 'hours saved' as 'money saved' without tying into staffing or throughput plan". Its diagnosis lists the downstream changes that do exist: AI "can shave minutes off tasks, but P&L only moves when that time converts into throughput, reduced overtime, fewer contractors, higher conversion or shorter cycle time with the same headcount. Without an operating plan, saved minutes can simply dissolve into the calendar."

Calculator defaults are old, too. Gartner's July 2024 release is the source of three figures still seeded into 2026 business cases - 15.8% revenue increase, 15.2% cost savings, 22.6% productivity improvement - from a survey of 822 leaders conducted between September and November 2023. The same release predicted "at least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025". A 2023 survey should not be the default value in a 2026 business case.

Payback runs longer than most calculators assume. Deloitte found "most respondents reported achieving satisfactory ROI on a typical AI use case within two to four years", against the seven to 12 months typically expected of a technology investment, with only 6% reporting payback in under a year and 13% of even the most successful projects seeing returns within 12 months. That is enterprise scale in Europe and the Middle East, and one fixed-price build on one small-business process is a different animal. But if your model shows five months, you are more optimistic than the median large enterprise.

What the "95% of AI projects fail" statistic actually measures

The most-quoted number in AI ROI, the 95% figure, has come loose from the thing that was measured. Here is the sentence it comes from, in the MIT NANDA report "The GenAI Divide: State of AI in Business 2025": "Despite $30-40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return." Organisations. The same report separately finds that 5% of integrated pilots are extracting millions, a different population pointing the other way. Elsewhere it calls the figure a "95% failure rate for enterprise AI solutions". IBM's own AI ROI page renders it as "95% of generative AI pilots are failing", footnoted to that same PDF. One statistic, at least three denominators: organisations, solutions, pilots. The measured one is the first. Open the PDF and you do not catch anybody out - you watch the drift happen inside a single document.

The NANDA report also states its own limits: 52 interviews, 153 surveyed leaders, directionally accurate on individual interviews rather than official company reporting, and it "may not represent broader market patterns". It is an industry report, not a peer-reviewed paper.

Two findings in that report beat its own headline. Top performers moved from pilot to implementation in around 90 days, against nine months or more at large enterprises. And where returns showed up, they came from reduced external spend: eliminating BPO contracts, cutting agency fees. Cancelled invoices, in other words. That is the downstream change, and it sits inside the most-cited AI ROI report there is.

FAQ

tags: roimeasurementadoptionprocess

Angus McDonald

Angus McDonald

Founder, Azgard

Builds and operates production AI systems for organisations that need results, not slide decks.