Four reasons AI programmes lose their funding.
Almost never because the technology failed. Usually because nobody could prove what changed.
No baseline was taken
Nobody measured what the process cost before. Every later claim becomes an argument between people with different recollections.
Hours saved were counted as cash
Ten people saving an hour a day is capacity, not savings. Finance notices this immediately, and the credibility of the whole case goes with it.
Nothing was attributed
Cycle time improved. So did staffing, seasonality and a process change. Without attribution the AI gets no credit, or gets credit it did not earn.
Run cost was never counted
Model, infrastructure and operations cost continues every month. A business case that ignores it is not a business case.
Five steps. The first one happens before we build anything.
The IX Value Engine is how measurement is built into the architecture rather than reconstructed from memory two quarters later.
Baseline Measure what it costs you today
Two weeks of actual measurement, not a workshop estimate. We instrument the current process before a single agent is designed, because this number is the one every later claim is judged against.
- Volume and cycle time
- Manual touch count per item
- Error and rework rate
- Fully loaded cost per transaction
- Exception and escalation volume
Measure Instrument the agent, not the demo
Every run is traced. We capture what the agent handled, what it escalated, how long it took and what it cost, at the level of individual transactions rather than monthly averages.
- Transactions handled end to end
- Human interventions and overrides
- Output quality against a golden set
- Latency and throughput
- Model and infrastructure cost per run
Attribute Answer the question a CFO will ask
Was it the AI, or was it everything else that changed at the same time? Attribution is the difference between a measurement and a claim, and it is the step most programmes skip.
- Per-transaction traceability
- Held-back control segment where volume allows
- Interventions logged, so assisted work is not counted as automated
- Run cost allocated per workflow, not pooled
- Same metric definition as the baseline
Optimize Improve the number, then re-measure
Once the system is live the cheapest wins are usually in routing, caching and prompt design rather than in more build. We tune against the same baseline so improvement is visible.
- Model routing and right-sizing
- Caching and retrieval tuning
- Automation rate improvement
- Escalation reduction
- Cost per transaction reduction
Prove ROI Report it in the language of finance
Not dashboards of token counts. A statement of what the process cost before, what it costs now, what the AI costs to run, and what the net position is against the original investment.
- Net value against baseline
- Payback position
- Cost per transaction, before and after
- Realisation assumptions stated openly
- Reported at the cadence finance already uses
Hours saved is the easy part. The rest is where the argument is won or lost.
Four categories, because a CFO, a COO and a risk committee will each ask about a different one.
Speed
- Cycle time
- Time to decision
- Turnaround
- Queue depth
- Time to first response
Automation
- Manual touches removed
- Transactions processed
- Straight-through rate
- Escalation volume
- Hours released
Cost & margin
- Cost per transaction
- Cost avoided
- Margin protected
- AI operating cost
- Return on investment
Quality & risk
- Error and rework rate
- First-contact resolution
- Accuracy against golden set
- Audit findings
- Policy exceptions
Freed-up hours are not savings.
If ten people each save an hour a day and nothing else changes, you have more capacity and exactly the same cost base. That is a real benefit, but it is not a number you can put in a budget.
Value is realised when that capacity is redeployed to work that generates revenue, or when cost genuinely comes out. Most AI business cases quietly assume one hundred percent conversion. We do not, and we say so before you fund anything.
Discuss your workflowIllustrative example. The realisation rate is agreed with your finance team at the start, not negotiated after the fact.
The hardest question is “was it the AI?”
Cycle time improves for many reasons. A process change, a staffing change, a quiet quarter. If you cannot separate the AI’s contribution from everything else, the business case does not survive its first serious review. These six mechanisms are configured on every engagement, which is why the answer is evidenced rather than argued.
Two weeks that decide whether the rest is worth doing.
Before any design work begins, we measure what the process actually costs today. It is the cheapest insurance available on an AI programme.
Map the real process
Not the documented one. Where work enters, who touches it, where it waits and where it goes wrong.
Instrument and sample
Pull system data where it exists, time the work where it does not, and count the exceptions everyone forgets.
Cost it properly
Volume, effort, rework, escalation and fully loaded cost per transaction, agreed with finance.
Agree the target
The metric, its definition, the realisation rate and the number the programme will be judged against.
If the baseline shows the process is not worth automating, we tell you. That outcome is a good result for both of us.
Three audiences. Three different reports.
The same measurements, presented for the person who has to act on them.
For finance
A position against the original investment, with the assumptions visible rather than buried.
- Net value against baseline
- Payback position
- Cost per transaction
- Realisation rate applied
- Run cost forecast
For operations
What actually changed in the workflow, at the level someone can act on this week.
- Cycle time and throughput
- Straight-through rate
- Exception and escalation volume
- Rework and error rate
- Queue and backlog position
For risk and audit
Evidence that the system behaved as approved, retained and retrievable.
- Agent action audit trail
- Evaluation results over time
- Human oversight and overrides
- Policy exceptions raised
- Model and version history
Measurement is part of the architecture, not a reporting project.
The IX Value Engine is the MEASURE layer of the IX AI Foundry. It ships with the solution, which is why the numbers exist without anyone having to go and find them.
What we will tell you that others will not.
Some workflows are genuinely hard to attribute value to. We flag those before you fund them, not after.
Straight-through rates below forty percent are normal in year one for anything involving exceptions, judgement or a regulator.
Build cost is driven by integration, not by AI. Anything writing back to a system of record costs materially more than read-only work.
Run cost grows with volume. Budget for it every year, not just in the first.
If the baseline says the process is not worth automating, that is the finding. We would rather lose the build than deliver something you cannot defend.
Measuring AI ROI questions
Baseline the workflow before building: volume, cycle time, manual touches, error rate and fully loaded cost per transaction. Measure agent activity in production at transaction level. Attribute the change to the AI rather than to everything else happening at the same time. Apply an agreed realisation rate, subtract the AI operating cost, and report the net position against the original investment.
It is the share of released time that becomes real money. Capacity only converts to cash when it is redeployed to revenue-generating work or when cost genuinely comes out. Most business cases assume one hundred percent without stating it. We agree the rate with your finance team at the start.
About two weeks for a single workflow. Days one to three map the real process, days four to seven instrument and sample it, days eight to eleven cost it with finance, and days twelve to fourteen agree the target metric and its definition.
Per-transaction traceability, a held-back control segment where volume allows, logged human interventions so assisted work is not counted as automated, run cost attributed per workflow, and reporting against the same metric definition used in the baseline.
Yes. Retrospective baselining is harder because the original state was never measured, but system logs and operational data are usually enough to reconstruct a defensible starting point.
Then that is the finding and we say so. A two-week baseline that prevents a poorly targeted build is a far better outcome than a project that quietly loses its funding in month nine.
No. It is the measurement layer of the IX AI Foundry, our delivery foundation. It is deployed as part of the solution we build, and the data belongs to you.
Bring us the workflow. We will tell you what it costs you today.
A thirty-minute session is usually enough to tell whether the value is there. If it is not, you will hear that too.