How do you actually measure AI ROI?

    Only 29% of executives can confidently measure AI ROI. The other 71% are running programs they can't defend at the board. Here's the framework that works for mid-market businesses, and the CFO test that filters real value from theatre.

    By Will Turner · Founder & Head of Data & AI26 April 20267 min read

    Only about 29% of executives say they can confidently measure the ROI of their AI investments. The other 71% are running programs they can't explain at the board table, which is one of the main reasons 42% of AI projects were abandoned in 2025. The technology mostly works. The measurement is what's missing.

    A measurement framework that works for mid-market businesses needs to do four things. Connect every AI initiative to a specific business outcome the CFO already tracks. Distinguish productivity from value. Run on a horizon long enough to give the work room to compound, but short enough that the board doesn't lose patience. And produce a number that survives an honest interrogation.

    The three-horizon framework

    Most ROI failures come from measuring everything as if it's the same thing. It isn't. AI initiatives produce returns on three different time horizons, and conflating them makes the whole portfolio look broken.

    Horizon 1: efficiency (3-9 months payback). Hours saved, processing errors avoided, time-to-output improvements. This is the easiest horizon to measure and the easiest to overcount. Saving an analyst four hours a week is worth zero dollars unless those four hours get redeployed to revenue-producing work, or unless you actually reduce headcount. Most "efficiency wins" never translate to a P&L line because nothing changes downstream of the freed-up time.

    Horizon 2: revenue impact (9-24 months payback). Conversion lift, retention improvement, expansion revenue, new product capability. Measurable, attributable, but slower. Requires controlled rollouts (A/B testing, holdout groups, before-and-after with cohort matching) to defend the number against the inevitable "it would have happened anyway" challenge.

    Horizon 3: risk and capability (24+ months payback). Compliance posture, security incidents avoided, talent attraction, strategic optionality. Genuinely real, hardest to put a dollar number on, easiest to dismiss as hand-waving. Worth measuring qualitatively even when you can't measure it quantitatively.

    A healthy AI portfolio has wins on all three horizons. A portfolio that's all efficiency stalls at the dashboard. A portfolio that's all revenue takes too long to fund. A portfolio that's all risk is a CDO defending a budget nobody else understands.

    The CFO test

    The single most useful question to ask of any proposed AI ROI metric: would the CFO sign for it?

    If the metric is "30% fewer manual touches on the invoice process," the CFO asks: does that mean we hire fewer accounts payable clerks next year, or do they have time to do other things they weren't doing before? If the answer is "they have more time," the CFO signs for hours saved as a productivity number, not a dollar number. If the answer is "we hire one fewer clerk," the CFO signs for the dollar value of the avoided hire.

    The number that goes on the ROI line of an AI initiative should be the one the CFO would put on next year's plan with a real budget impact behind it. Everything else is interesting but not ROI.

    If a metric needs three slides of context to defend, it's not ROI. It's narrative. Both are useful. Don't confuse them.

    What to measure in the first 12 months

    For a typical mid-market AI initiative, the credible measurement set looks like this:

    • One leading indicator, measured weekly. The thing the model output is supposed to change. Calls made on top-ranked leads. Documents auto-processed. Forecast MAPE.
    • One operational metric, measured monthly. The downstream business operation the leading indicator should improve. Conversion rate. Headcount per claim. Stockout frequency.
    • One financial metric, measured quarterly. The dollar number the operational metric should ladder into. Pipeline-weighted revenue. Operating cost per claim. Gross margin.
    • One risk metric, measured ad hoc. The thing that would invalidate the whole project. Model drift. False positive rate. Audit finding.

    Three of those are leading-edge measures the team controls. The fourth is the one the CFO actually cares about. Track all four, report all four, and do not let the team cherry-pick the favourable number on report day.

    Common measurement mistakes

    Counting time saved as if it were money saved. It isn't, until something changes downstream. An hour of analyst time freed up is an hour of analyst time freed up, not a dollar.

    Attributing all change to the AI when other things changed too. A new sales prioritisation engine launched the same quarter as a new SDR team is two interventions, not one. Build a holdout group or accept that your number is noisy.

    Measuring against the wrong baseline. If the previous quarter was unusually bad, year-over-year improvement looks great and is fake. Compare against a sensible same-period baseline or against a controlled holdout.

    Letting model accuracy stand in for business value. A 92% AUC churn model that doesn't change retention rate has zero ROI, however accurate.

    Stopping the measurement after launch. The first three months of any AI project are honeymoon results. Real ROI is the steady-state number after month six, not the launch number.

    The honest payback expectation

    Surveys put average AI initiative payback at 28 months for enterprise programs. Mid-market initiatives, narrower in scope, often pay back faster, in the 9-18 month range for well-scoped efficiency or revenue projects. The gap between what executives expect (most surveys put expected payback at under 12 months) and what actually happens is the single biggest driver of project cancellation in 2026.

    The fix is not to accelerate the project. The fix is to set the right expectation at sponsor level before kickoff, and to stage the project so the first defensible win lands in month six even if full payback is in month eighteen. A board sees the month-six win, the project survives, and month-eighteen value compounds. A board that expected payback in month twelve and didn't see it pulls the plug, and the month-eighteen value never arrives.

    The one-page ROI brief that ships projects

    Every AI initiative we run for clients starts with a one-page brief that answers four questions before any code is written:

    1. What specific business decision does this change, and who makes it?
    2. What financial metric does that decision feed into, and what's the current baseline?
    3. What's the realistic value range and the realistic time-to-value range?
    4. What's the kill criteria, and who decides?

    If the sponsor can't sign that page, the project isn't ready. If the sponsor signs it and the CFO refuses to acknowledge the value range, the project isn't ready either. The one-pager is cheap. The post-mortem on a year of unmeasured work is expensive.

    Common Questions

    Frequently asked

    Got one of these problems in front of you?

    Beyond Data runs engagements that put the ideas in this insight into practice.

    Cookie preferences

    We use essential cookies to run the site. With your permission, we also use analytics cookies to understand what content helps visitors make better data and AI decisions.