How We Measure AI ROI

IX Value Engine

Most AI business cases are written after the money is spent.

By then nobody remembers what the process actually cost before. We baseline the workflow before we write code, measure what the AI does in production, and report the change against the number we started with.

Baselined first Before anything is built
Measured in production Not sampled or estimated
Attributed, not assumed Separated from other changes
Reported to finance In language a CFO accepts
Why This Matters

Four reasons AI programmes lose their funding.

Almost never because the technology failed. Usually because nobody could prove what changed.

01

No baseline was taken

Nobody measured what the process cost before. Every later claim becomes an argument between people with different recollections.

02

Hours saved were counted as cash

Ten people saving an hour a day is capacity, not savings. Finance notices this immediately, and the credibility of the whole case goes with it.

03

Nothing was attributed

Cycle time improved. So did staffing, seasonality and a process change. Without attribution the AI gets no credit, or gets credit it did not earn.

04

Run cost was never counted

Model, infrastructure and operations cost continues every month. A business case that ignores it is not a business case.

The Method

Five steps. The first one happens before we build anything.

The IX Value Engine is how measurement is built into the architecture rather than reconstructed from memory two quarters later.

Step 01

Baseline Measure what it costs you today

Two weeks of actual measurement, not a workshop estimate. We instrument the current process before a single agent is designed, because this number is the one every later claim is judged against.

  • Volume and cycle time
  • Manual touch count per item
  • Error and rework rate
  • Fully loaded cost per transaction
  • Exception and escalation volume
Step 02

Measure Instrument the agent, not the demo

Every run is traced. We capture what the agent handled, what it escalated, how long it took and what it cost, at the level of individual transactions rather than monthly averages.

  • Transactions handled end to end
  • Human interventions and overrides
  • Output quality against a golden set
  • Latency and throughput
  • Model and infrastructure cost per run
Step 03

Attribute Answer the question a CFO will ask

Was it the AI, or was it everything else that changed at the same time? Attribution is the difference between a measurement and a claim, and it is the step most programmes skip.

  • Per-transaction traceability
  • Held-back control segment where volume allows
  • Interventions logged, so assisted work is not counted as automated
  • Run cost allocated per workflow, not pooled
  • Same metric definition as the baseline
Step 04

Optimize Improve the number, then re-measure

Once the system is live the cheapest wins are usually in routing, caching and prompt design rather than in more build. We tune against the same baseline so improvement is visible.

  • Model routing and right-sizing
  • Caching and retrieval tuning
  • Automation rate improvement
  • Escalation reduction
  • Cost per transaction reduction
Step 05

Prove ROI Report it in the language of finance

Not dashboards of token counts. A statement of what the process cost before, what it costs now, what the AI costs to run, and what the net position is against the original investment.

  • Net value against baseline
  • Payback position
  • Cost per transaction, before and after
  • Realisation assumptions stated openly
  • Reported at the cadence finance already uses
What We Measure

Hours saved is the easy part. The rest is where the argument is won or lost.

Four categories, because a CFO, a COO and a risk committee will each ask about a different one.

Speed

  • Cycle time
  • Time to decision
  • Turnaround
  • Queue depth
  • Time to first response

Automation

  • Manual touches removed
  • Transactions processed
  • Straight-through rate
  • Escalation volume
  • Hours released

Cost & margin

  • Cost per transaction
  • Cost avoided
  • Margin protected
  • AI operating cost
  • Return on investment

Quality & risk

  • Error and rework rate
  • First-contact resolution
  • Accuracy against golden set
  • Audit findings
  • Policy exceptions
The Honest Part

Freed-up hours are not savings.

If ten people each save an hour a day and nothing else changes, you have more capacity and exactly the same cost base. That is a real benefit, but it is not a number you can put in a budget.

Value is realised when that capacity is redeployed to work that generates revenue, or when cost genuinely comes out. Most AI business cases quietly assume one hundred percent conversion. We do not, and we say so before you fund anything.

Discuss your workflow
Hours released each month1,080
Loaded hourly cost$48
Gross value of released time$51,840
Realisation rate applied60%
Realised value$31,104
AI operating cost−$7,500
Net monthly value$23,604

Illustrative example. The realisation rate is agreed with your finance team at the start, not negotiated after the fact.

Attribution

The hardest question is “was it the AI?”

Cycle time improves for many reasons. A process change, a staffing change, a quiet quarter. If you cannot separate the AI’s contribution from everything else, the business case does not survive its first serious review. These six mechanisms are configured on every engagement, which is why the answer is evidenced rather than argued.

Baseline captured before build
Two weeks of measurement, not a workshop estimate.
Per-transaction traceability
Every item the agent touched is identifiable and separable.
Held-back control segment
Where volume allows, part of the flow keeps running the old way.
Interventions logged
Escalations and overrides counted, so assisted work is never reported as automated.
Run cost attributed per workflow
Model, infrastructure and operations cost allocated by use case, not pooled.
Same metric, same definition
Reported against the number we started with, not a new one.
The Baseline

Two weeks that decide whether the rest is worth doing.

Before any design work begins, we measure what the process actually costs today. It is the cheapest insurance available on an AI programme.

Days 1–3

Map the real process

Not the documented one. Where work enters, who touches it, where it waits and where it goes wrong.

Days 4–7

Instrument and sample

Pull system data where it exists, time the work where it does not, and count the exceptions everyone forgets.

Days 8–11

Cost it properly

Volume, effort, rework, escalation and fully loaded cost per transaction, agreed with finance.

Days 12–14

Agree the target

The metric, its definition, the realisation rate and the number the programme will be judged against.

If the baseline shows the process is not worth automating, we tell you. That outcome is a good result for both of us.

Reporting

Three audiences. Three different reports.

The same measurements, presented for the person who has to act on them.

For finance

A position against the original investment, with the assumptions visible rather than buried.

  • Net value against baseline
  • Payback position
  • Cost per transaction
  • Realisation rate applied
  • Run cost forecast

For operations

What actually changed in the workflow, at the level someone can act on this week.

  • Cycle time and throughput
  • Straight-through rate
  • Exception and escalation volume
  • Rework and error rate
  • Queue and backlog position

For risk and audit

Evidence that the system behaved as approved, retained and retrievable.

  • Agent action audit trail
  • Evaluation results over time
  • Human oversight and overrides
  • Policy exceptions raised
  • Model and version history
Where It Lives

Measurement is part of the architecture, not a reporting project.

The IX Value Engine is the MEASURE layer of the IX AI Foundry. It ships with the solution, which is why the numbers exist without anyone having to go and find them.

The Straight Talk

What we will tell you that others will not.

Some workflows are genuinely hard to attribute value to. We flag those before you fund them, not after.

Straight-through rates below forty percent are normal in year one for anything involving exceptions, judgement or a regulator.

Build cost is driven by integration, not by AI. Anything writing back to a system of record costs materially more than read-only work.

Run cost grows with volume. Budget for it every year, not just in the first.

If the baseline says the process is not worth automating, that is the finding. We would rather lose the build than deliver something you cannot defend.

FAQ

Measuring AI ROI questions

Baseline the workflow before building: volume, cycle time, manual touches, error rate and fully loaded cost per transaction. Measure agent activity in production at transaction level. Attribute the change to the AI rather than to everything else happening at the same time. Apply an agreed realisation rate, subtract the AI operating cost, and report the net position against the original investment.

It is the share of released time that becomes real money. Capacity only converts to cash when it is redeployed to revenue-generating work or when cost genuinely comes out. Most business cases assume one hundred percent without stating it. We agree the rate with your finance team at the start.

About two weeks for a single workflow. Days one to three map the real process, days four to seven instrument and sample it, days eight to eleven cost it with finance, and days twelve to fourteen agree the target metric and its definition.

Per-transaction traceability, a held-back control segment where volume allows, logged human interventions so assisted work is not counted as automated, run cost attributed per workflow, and reporting against the same metric definition used in the baseline.

Yes. Retrospective baselining is harder because the original state was never measured, but system logs and operational data are usually enough to reconstruct a defensible starting point.

Then that is the finding and we say so. A two-week baseline that prevents a poorly targeted build is a far better outcome than a project that quietly loses its funding in month nine.

No. It is the measurement layer of the IX AI Foundry, our delivery foundation. It is deployed as part of the solution we build, and the data belongs to you.

Talk to an AI Expert

Bring us the workflow. We will tell you what it costs you today.

A thirty-minute session is usually enough to tell whether the value is there. If it is not, you will hear that too.

What Happens Next

A 30-minute working session with an AI architect and a delivery lead. No pitch deck.

A 30-Minute Working Session

With an AI architect and a delivery lead. No pitch deck.

Bring One Workflow

The volume, the systems it touches and what it costs you today.

You Leave With

A view on feasibility, the metric worth targeting and what a baseline would involve.