- Why Enterprise AI ROI Looks Broken in 2026
- What AI ROI Actually Means at the Enterprise Scale
- The Five Categories of AI Value You Can Actually Defend
- The Honest Total Cost of Ownership Model
Two numbers should anchor every enterprise AI conversation in 2026. MIT Sloan’s review of 300 public deployments and 153 executive surveys found that 95% of generative AI pilots produced no measurable profit-and-loss impact. In parallel, McKinsey’s State of AI work shows that more than 80% of organisations report no tangible enterprise-level EBIT impact from generative AI, even though 88% are actively experimenting with it.
That gap is the entire story of enterprise AI right now. The technology works. The funding is there. The execution is not converting into measurable value at the EBIT line. And the failure mode is rarely the model. It is the absence of a defensible AI ROI measurement framework. Your CFO does not need a louder demo. Your CFO needs a number that survives audit, a TCO that includes the costs your vendor would prefer you forget, and a benefits model that can be re-measured after launch.
This is the practitioner guide we use with Sthambh clients in Singapore, Hong Kong, London, and the US to build that framework. It is built for CTOs, VPs of Engineering, Heads of AI, and finance partners who have to defend the spend, not just request it.
Why Enterprise AI ROI Looks Broken in 2026
The headline failure rate is a symptom. Three structural issues sit underneath it.
The first is that most enterprises approved their first wave of GenAI projects against projected ROI that was never measured again after launch. A 2025 MIT Sloan study found that 61% of enterprise AI projects were approved on projected ROI that nobody went back to validate. The same study found that 73% of failed AI projects had no agreed definition of success before the work began. Defensible measurement is not a reporting problem. It is a procurement problem that compounds into a reporting problem.
The second is that production economics diverge sharply from pilot economics. MIT Sloan’s data shows production cost overruns averaging 380% compared to pilot projections. A surface like 85% of organisations misestimate AI project costs by more than 10%, and nearly a quarter underestimate by 50% or more. The costs that surface in production are the ones a typical pilot does not capture. Persistent inference at production volume. Cross-region data transfer. Continuous retraining as content drifts. Storage sprawl. Compliance review cycles. Vendor management overhead. In production environments, hidden costs make up 60 to 80% of total spend, not the model API line your finance team was watching.
The third is that the benefits side is poorly modelled. Productivity gains, the historical headline benefit, have been falling as the primary ROI metric tracked by enterprises. Futurum’s research shows direct financial impact, top-line revenue and bottom-line profit, doubling to 21.7% of primary ROI responses in 2026. Productivity has fallen from 23.8% to 18.0% as the lead metric. Boards no longer accept “saved hours” without a translation into either revenue, headcount neutrality, or a cost line. If you cannot draw that translation, the savings do not get counted.
The takeaway: the AI ROI problem is not that GenAI does not produce value. It is that the value is being captured by 5% of organisations with disciplined measurement, and lost by the 95% who cannot tell their CFO what changed at the P&L line.
What AI ROI Actually Means at the Enterprise Scale
The classic formula still works. ROI equals net benefits minus total costs, divided by total costs, expressed as a percentage. The problem is not the formula. The problem is the inputs.
At enterprise scale, the inputs you need to defend are these. Total costs must include the full lifecycle, not just the build. Benefits must distinguish hard benefits, which can be journal-entered against a P&L line, from soft benefits, which need a translation layer before they appear in any financial model. Time horizon matters more than at any other software investment because most enterprise AI does not deliver satisfactory ROI inside 7 to 12 months. The realistic payback window for transformational AI is 2 to 4 years, which means your single-year ROI calculation will lie to you in both directions.
The second concept your model needs is a difference between attributable value and correlational value. DBS Bank’s SGD 1 billion AI economic value claim for 2025 is widely cited because the bank uses a control-group benchmarking approach. Customer outcomes from AI-powered solutions are compared against a matched control group. The SGD 1 billion is the lift, not the gross revenue passing through the AI surface. Most enterprises skip this step. The result is an ROI number that the AI team believes, the operations team disputes, and the CFO eventually discounts to zero.
Defensible AI ROI measurement at enterprise scale requires three commitments before a single line of code is written. A pre-registered hypothesis with quantified success criteria. A baseline measurement of the workflow being changed. A control mechanism, whether a holdout group, a pre-and-post window, or a comparison against a matched cohort, that lets you separate AI lift from background change.
The Five Categories of AI Value You Can Actually Defend
The AI ROI frameworks that survive audit treat value as five categories, not one. Each requires a different measurement approach. Each lands on a different line of the P&L. Mixing them is the single most common mistake we see in enterprise business cases.
Cost Reduction Value
The most defensible category. AI absorbs work that would otherwise require headcount, vendor spend, or third-party processing. The right metric is unit cost reduction, measured per transaction, per ticket, per document, per claim, or per contract. Multiply by volume and you get an annualised cost saving.
The trap: AI rarely eliminates a full role in a single workflow. It compresses time per unit. Unless you take the compressed time and either reduce headcount, redirect to higher-value work that itself generates measurable revenue, or cap hiring against a growing workload, the saving stays theoretical. Most enterprises do not enforce that translation. The CFO is right to ignore the savings number until they do.
Revenue Lift Value
AI surfaces drive incremental revenue through better recommendations, faster sales cycles, improved conversion, recovered abandoned baskets, expanded cross-sell, or new product capability that justifies a price increase. This is the value category that has grown fastest in 2026 measurement frameworks because it lands directly on the top line.
The trap: revenue lift requires a counterfactual. Without a holdout group or an A/B framework, the lift number is contaminated by everything else changing in the market. Singapore retail banks running personalisation engines on customer advisor copilots have learned to randomise at the relationship-manager level, not the customer level, so they can defend the lift number to internal audit.
Risk Avoidance Value
The category where AI prevents losses that would otherwise show up as fraud, fines, errors, customer attrition, or regulatory action. Risk avoidance is real, sometimes the largest single value pool, and the hardest to defend without statistical rigour. The CFO is trained to discount loss-avoidance claims because they cannot be journal-entered until the loss happens.
The right approach: model the expected loss using historical incident data, multiply by the probability change the AI introduces (measured against a baseline cohort or period), and report the avoided expected value with confidence intervals. UK insurers running claims-triage AI use this approach to defend the spend against PRA examiners.
Productivity and Capacity Value
The category that has lost the most ground in 2026 measurement frameworks. Productivity gains, expressed as hours saved per employee per week, were the dominant ROI metric three years ago. They no longer carry the CFO conversation alone.
Productivity value becomes defensible when it converts into one of four downstream effects. Headcount neutrality against rising volume, where you absorb growth without hiring. Service-level improvements that protect renewal revenue or reduce churn. Capacity redirection to revenue-generating work, with the revenue tracked back. Reduced overtime, contractor spend, or outsourced processing. Without one of these conversions, the saved hours do not appear at the P&L line, and the productivity number stops counting.
Strategic and Capability Value
The category that funds the platform, not the use case. This is the value of having a working data foundation, an evaluation harness, a governance framework, and an internal team that can ship the next AI use case faster than the last one. Strategic value rarely appears in a single-use-case ROI calculation. It appears in the velocity at which the second, third, and tenth use case clear pilot.
The right approach: track time-from-idea-to-production for the first use case versus the fifth versus the tenth. Track the percentage of pilots that reach production. Track reuse of platform components across teams. Strategic value is not soft. It is a measured reduction in marginal cost per use case shipped. Boards understand this argument when it is quantified.
The Honest Total Cost of Ownership Model
The cost side of the equation is where most enterprise AI business cases collapse. The vendor will hand you a price per million tokens and a one-time integration estimate. The actual TCO is roughly 3 to 5 times that line.
A defensible AI TCO model has seven cost categories. None of them is optional. The first is direct inference and model cost, the line your vendor focuses on. The second is data foundation cost, which covers ingestion, cleaning, vectorisation, lineage, and retention. The third is integration and engineering cost, including the work to wire AI into existing systems of record, identity, and workflow. The fourth is evaluation and observability cost, the eval harness, the tracing, the regression suite, the on-call coverage. The fifth is governance and compliance cost, which is rising fast as MAS AIRG, HKMA SA-2, the EU AI Act, the UK ICO guidance, and US sectoral rules add review cycles to every production change. The sixth is change management and training cost. The seventh is ongoing maintenance, drift management, and version migrations as your model vendor releases new generations.
The table below shows the typical cost split for a mid-size enterprise GenAI initiative reaching production at scale, calibrated from Sthambh client engagements in 2025 and 2026. Numbers are USD ranges per year for an initiative supporting roughly 500 daily users at production volume.
| Cost Category | Typical Annual Range (USD) | % of Total TCO | Where it gets missed |
|---|---|---|---|
| Direct inference & model | $180k - $480k | 18 - 22% | The only line vendors quote |
| Data foundation & retrieval | $200k - $550k | 20 - 24% | Vector DB, ingestion, lineage |
| Integration & engineering | $220k - $620k | 22 - 26% | Systems of record, identity, workflow |
| Evaluation & observability | $120k - $280k | 10 - 14% | Eval harness, tracing, on-call |
| Governance & compliance | $100k - $260k | 10 - 12% | Review cycles, audit evidence |
| Change management & training | $80k - $200k | 7 - 9% | Adoption is where ROI dies |
| Maintenance & migrations | $60k - $160k | 5 - 7% | Drift, version upgrades, deprecation |
Two patterns are worth pulling out of this table. First, the direct inference cost, the line that dominates vendor conversations, is between one fifth and one quarter of the total. If you build your business case around that number, you are wrong by a factor of four. Second, the data and integration lines together exceed 40% of TCO in every well-measured deployment we have seen. The 30% of GenAI projects Gartner predicted would be abandoned in 2025 were predominantly killed by these two lines, not the model.
The hidden costs deserve their own warning. Storage sprawl, cross-region data transfers, idle compute, and continuous retraining can together make up 60 to 80% of production spend in environments that did not budget for them. Compliance and governance infrastructure adds 10 to 20% to total project costs in regulated industries. If your business case does not have a line for each of these, it is not yet ready for finance review.
The AI ROI Calculator: A Defensible Worksheet
The calculator below is the working template Sthambh uses with clients building a business case for a single GenAI use case. It is built to survive both audit and post-launch reality. We have seen this exact structure clear procurement review at MAS-regulated Singapore banks, HKMA-supervised Hong Kong insurers, and FCA-regulated UK fintechs.
Section 1: Hypothesis and success criteria
State the hypothesis in one sentence. “Deploying a RAG-based compliance research assistant will reduce average research time per regulation lookup from 45 minutes to 15 minutes for a population of 80 compliance analysts.” State the quantified success criterion. “We define success as a 60% reduction in average research time and a 90% accuracy rate against a golden set of 200 historical lookups, measured 90 days post-rollout.”
Without these two sentences your business case has nothing to be re-measured against. The 73% of failed AI projects MIT Sloan identified all skipped this section.
Section 2: Baseline measurement
Measure the current state of the workflow. Time per unit. Cost per unit. Error rate. Volume per period. Capture this before the AI ships, with a methodology your operations team will accept. Pre-launch baselines that are estimated rather than measured are the most common reason ROI numbers fall apart in audit.
Section 3: Value translation
Convert the workflow improvement into one of the five value categories using the table below. The conversion is the load-bearing step. Without it, you have a saved-time number, not an ROI claim.
| Workflow Improvement | Value Category | P&L Translation |
|---|---|---|
| Time per unit falls 40-60% | Cost reduction or capacity | Reduced contractor spend, capped hiring against volume growth, or redirected to revenue-generating work |
| Conversion rate rises 5-15% | Revenue lift | Incremental revenue, measured against holdout cohort |
| Error rate falls 30-70% | Risk avoidance or cost reduction | Reduced rework cost, reduced loss expectancy, reduced regulatory exposure |
| Cycle time falls 30-50% | Capacity or revenue lift | Higher throughput against fixed cost, faster revenue recognition, improved customer satisfaction tied to renewals |
| Customer satisfaction rises 10-20 points | Revenue lift or risk avoidance | Reduced churn against control cohort, increased lifetime value, lower acquisition cost |
Section 4: Cost roll-up
Roll up all seven TCO categories for years 1, 2, and 3. Use vendor quotes for direct cost, internal labour rates for engineering and ops, and historical compliance spend per major change for governance. Add a 15-20% contingency line. Year 1 costs are usually 60% above year 3 steady state because of build effort. Year 3 is when you find out whether the operating cost is sustainable.
Section 5: Net present value and payback
Calculate net cash flow per year for three years. Discount at your enterprise hurdle rate, typically 10-15%. Compute net present value and identify the month payback is achieved. For most enterprise AI initiatives, payback occurs between month 14 and month 28. Anything inside 12 months is either a very narrow use case or an overly optimistic model. Anything beyond 36 months belongs on the platform investment, not the use case.
Section 6: Sensitivity and kill criteria
Model the business case at three adoption levels, 60%, 80%, and 95% of target users actively using the system 90 days post-rollout. Adoption is where most ROI dies. Define kill criteria explicitly. If adoption is below 50% at day 60, what action triggers? If accuracy on the golden set drops below 85% at day 30, what action triggers? Pre-committed kill criteria are how disciplined enterprises stop spending against the 95% of pilots that go nowhere.
How Top Performers Capture Value: What the Data Says
Two patterns separate the 5% of enterprises capturing measurable EBIT impact from the rest. The first is concentration. McKinsey’s State of AI work shows that top-quartile organisations invest several times more in data foundations, governance, and integration than they do in models. They are not buying the most expensive AI. They are buying the cheapest possible model on top of the most expensive possible data layer. This inverts the budget conversation most enterprises have.
The second is governance discipline. Only 14% of organisations have leaders consistently championing AI with a clear strategy. The ones that do show measurably better ROI capture because the strategy translates into the same boring artefacts every quarter. Prioritised use case backlog. Pre-registered hypotheses. Measured baselines. Post-launch re-measurement. Kill criteria. None of this is novel. All of it is hard. The 86% who skip it produce the 95% failure-to-impact rate.
Deloitte’s 2026 State of AI in the Enterprise survey of 1,854 executives reinforces this. Almost three-quarters of organisations report that their most advanced AI initiatives met or exceeded ROI targets. Around 20% see returns above 30%. The “most advanced” qualifier is doing the heavy lifting. Among the 5% of GenAI pilots that did show P&L impact in the MIT Sloan data, the median ROI was 188%. The dispersion between disciplined deployments and the average is enormous.
The lesson for any CTO building a 2026 business case: do not benchmark against the average. The average is a 0% return on a billion dollars of enterprise AI spend. Benchmark against the disciplined cohort and accept that joining it requires the measurement discipline most enterprises skip.
Regulatory and Audit Pressure on the Business Case
The measurement framework is not just for the CFO. As of mid-2026, three regulatory bodies are explicitly examining the ROI claims of enterprise GenAI deployments.
The Monetary Authority of Singapore’s AI Risk Management Guidelines, with enforcement on the Q1 2027 track, require institutions to demonstrate that AI deployments deliver intended outcomes and that those outcomes are continuously monitored. A business case that cannot be re-measured fails the SR/CIO sign-off requirement. Singapore banks running GenAI customer-advisor copilots have rebuilt their measurement frameworks specifically to clear this bar.
The Hong Kong Monetary Authority’s SA-2 module on AI risk requires authorised institutions to maintain documented business cases with measurable outcomes, reviewable on examination. The HKMA’s GenAI Sandbox programme, now six months into its second phase, has produced a clear signal that examiners want to see post-launch ROI measurement, not just pre-launch projections.
The UK ICO’s AI guidance and the FCA’s SS1/23 operational resilience requirements together create a similar expectation. Boards of FCA-supervised firms cannot delegate AI investment oversight without documented evidence of ongoing value and risk measurement. A business case approved in 2024 that has not been re-measured against actual outcomes is increasingly a finding waiting to happen.
For US firms, the SEC’s cyber and emerging-technology examination focus, plus sectoral pressure from HIPAA, OCC, and FINRA, push in the same direction. A defensible AI ROI measurement framework is now part of the control environment, not just the budgeting cycle. The good news is that the same artefacts that satisfy your CFO also satisfy your regulator. The framework is one investment, not two.
Real-World Examples of Defensible AI ROI Measurement
Three anonymised examples from Sthambh client work and public case studies illustrate what disciplined measurement looks like.
A Singapore retail bank deploying a customer-advisor copilot for its mass-affluent segment built a holdout framework at the relationship-manager level. Half the RMs received the copilot, half continued with standard tooling. After six months, the AI-enabled RMs showed a 14% lift in cross-sell conversion and a 9-minute reduction in average client-prep time. The bank’s CFO accepted the revenue lift number because the holdout was defensible to internal audit. The productivity number was translated into a 30% reduction in overtime claims among the AI-enabled cohort, with the cost saving journal-entered against the contractor budget line. Total annualised value capture was SGD 6.2 million against a build cost of SGD 1.8 million and an annual run cost of SGD 1.1 million. Payback at month 11.
A UK insurer running a claims-triage AI took the opposite approach to a similar problem. Rather than a holdout, the insurer measured a 6-month pre-launch baseline of claims-handling cost, fraud detection rate, and customer satisfaction. Post-launch, the same metrics were measured monthly against the baseline with macroeconomic adjustments for claims volume. The framework cleared PRA examination because it included documented uncertainty bounds. The insurer captured a 22% reduction in handling cost per claim and an 18% improvement in fraud detection accuracy, translated into GBP 4.4 million annualised value against GBP 1.5 million in TCO. The CFO sign-off depended on the uncertainty bounds being explicit, not the headline number being large.
DBS Bank’s publicly disclosed AI economic value, SGD 1 billion in 2025 against an aspiration set in 2022 of SGD 1 billion within five years, is the cleanest large-scale example of disciplined measurement in the public domain. The Harvard Business School case study authored by Professor Feng Zhu documents that DBS uses control-group benchmarking across 350 use cases supported by 800-plus models. The value tracks to three lines, increased revenue, cost savings, and risk avoidance, and is measured against matched cohorts. The structure DBS uses at scale is the same structure a 200-employee fintech should use for its first use case. The mathematics does not change with scale, only the budget does.
Common Mistakes That Kill Enterprise AI ROI
Six recurring mistakes explain most of the value gap between disciplined deployments and the average.
The first is starting with the model. The vendor evaluation conversation pulls budget toward the most visible cost line, which is the model. The 80% of TCO that lives in data, integration, and operations gets undercosted. The result is a deployment that hits production with no eval harness, no observability, and no governance, and that drifts within six months.
The second is measuring at the activity level instead of the outcome level. “Number of queries handled” is an activity. “Cost per resolved query” is an outcome. Boards do not pay for activity. They pay for outcomes.
The third is skipping the baseline. The number of enterprise AI business cases approved against an estimated baseline rather than a measured one remains the single largest source of false ROI claims. If you cannot measure the workflow today, you cannot defend the improvement tomorrow.
The fourth is ignoring adoption. The 95% of pilots that produce no P&L impact include a large cohort where the system works and nobody uses it. Adoption belongs in the business case as a sensitivity, with kill criteria attached. Productivity multiplied by 20% adoption is a rounding error.
The fifth is the single-year ROI calculation. Most enterprise AI does not deliver satisfactory ROI inside 12 months. The single-year view either rejects projects that would be transformational on a three-year horizon, or accepts projects that look strong in year one because year-two operating costs are not yet on the books.
The sixth is not re-measuring. The MIT Sloan finding that 61% of approved projects are never re-measured after launch is the single most important data point in this category. Measurement is the only forcing function that closes the gap between projected and realised value. If your enterprise does not have a standing process to re-measure every approved AI investment against its own business case at month 6, month 12, and month 24, the framework is theatre.
A Practical Three-Phase Implementation of the Framework
Most enterprises do not need to wait for a full programme to start measuring AI ROI properly. A staged implementation gets defensible numbers into the operating cycle inside one quarter.
Phase 1: Audit and Baseline (Weeks 1 – 6)
Inventory every GenAI initiative currently in pilot or production. For each, document the original hypothesis, the projected ROI, the projected cost, and the actual measured value against each. Most enterprises find that 60-80% of in-flight initiatives have no documented baseline and no defensible re-measurement plan. That is the audit finding that funds the framework.
Pick three priority workflows where AI is most likely to land at the EBIT line. Build a measured baseline for each. Define quantified success criteria. This work is unglamorous and entirely necessary.
Phase 2: Framework Rollout (Weeks 7 – 14)
Adopt the seven-category TCO template across all new investment proposals. Require every new business case to include hypothesis, baseline, value translation, sensitivity, and kill criteria. Adopt the five-value-category model so finance can see where the claimed value will land on the P&L. Train finance and AI teams together. The CFO’s office is your strongest ally if you give them a framework, and your most expensive opponent if you do not.
Phase 3: Operationalisation (Weeks 15 – 24)
Build a quarterly re-measurement cadence into the AI governance process. Every approved investment is re-scored at month 6, 12, and 24 against its own business case. Kill criteria are enforced. Investments that overperform get reinvestment priority. Investments that underperform get a remediation window and a stop-loss. This is the discipline that produces the 188% median ROI in the successful cohort. It is not glamorous and it is not optional.
How Sthambh Helps Enterprises Build a Defensible AI ROI Framework
Sthambh works with CTOs, CFOs, and Heads of AI in Singapore, Hong Kong, the UK, and the US to build the AI ROI measurement framework their enterprise can actually defend. The engagement typically starts with a portfolio audit of in-flight GenAI initiatives, surfaces the value capture gap between projected and realised, and produces a CFO-ready business case template that the next wave of investment proposals will use.
We bring the seven-category TCO model, the five-value-category benefits translation, and the pre-launch and post-launch measurement playbooks proven across regulated deployments. We also bring practitioner experience in building AI agents for enterprise, LLM integration, and the data foundation work that determines whether the business case is defensible in the first place. Our work is technical, financial, and regulatory in equal measure, because the modern AI business case is all three at once.
If you are about to submit your 2027 AI budget, are facing board questions about value capture from your 2025 and 2026 spend, or are preparing for a MAS, HKMA, FCA, or PRA examination that will ask about AI ROI evidence, we can help you build the framework before the spend, not after. Book an enterprise AI strategy call with Nikhil to walk through your portfolio.
FAQs
Q. What is the realistic payback period for enterprise GenAI projects?
A. Most disciplined enterprise GenAI deployments achieve payback between month 14 and month 28. Single-year ROI calculations either reject investments that would be transformational on a three-year horizon, or accept ones that look strong in year one because year-two operating costs are not yet on the books. A 7-to-12-month payback expectation, common in traditional software procurement, is unrealistic for transformational AI and is one of the leading causes of premature project termination.
Q. How do I separate AI lift from background change when measuring revenue impact?
A. Use a holdout cohort wherever the workflow allows it. Singapore retail banks running personalisation engines randomise at the relationship-manager level so they can compare matched cohorts of customers. UK insurers using claims-triage AI use pre-and-post baselines with macroeconomic adjustments where holdouts are not operationally feasible. The principle is the same: you need a counterfactual to defend the lift number. Without it, the CFO is right to discount the claim.
Q. What does an honest AI total cost of ownership include beyond model cost?
A. Seven categories. Direct inference and model cost (the only line vendors quote). Data foundation cost (ingestion, vectorisation, lineage). Integration and engineering cost (wiring to systems of record). Evaluation and observability cost (eval harness, tracing). Governance and compliance cost (review cycles, audit evidence). Change management and training cost. Maintenance and migration cost (drift management, version upgrades). The direct model line is usually 18 to 22% of TCO. If your business case only models that line, it understates total cost by a factor of four.
Q. How do I defend an AI ROI claim to internal audit?
A. Four artefacts. A pre-registered hypothesis with quantified success criteria. A measured pre-launch baseline of the workflow. A control mechanism (holdout cohort, matched cohort, or pre-and-post with adjustments) that lets you isolate AI lift from background change. Documented uncertainty bounds on the headline number. Internal audit teams are not hostile to AI ROI claims. They are hostile to claims that cannot be re-measured. Build for re-measurability and the audit conversation simplifies.
Q. How are MAS, HKMA, and FCA examiners looking at AI ROI evidence in 2026?
A. All three regulators now expect documented business cases with measurable outcomes that are reviewable on examination. MAS AIRG, on the Q1 2027 enforcement track, requires institutions to demonstrate that AI deployments deliver intended outcomes and that outcomes are continuously monitored. HKMA’s SA-2 module on AI risk requires the same documentation discipline. The UK FCA’s SS1/23 operational resilience requirements push boards to evidence ongoing value and risk measurement. The framework that satisfies your CFO and the framework that satisfies your regulator are the same framework.
Q. Why do productivity gains no longer carry the CFO conversation alone?
A. Boards have learned that saved hours do not appear at the P&L line unless they convert into one of four downstream effects: capped hiring against rising volume, service-level improvements that protect renewal revenue, redirected capacity that generates measurable revenue, or reduced contractor and overtime spend. Productivity research from Futurum shows productivity falling from 23.8% to 18.0% as the primary ROI metric in 2026, while direct financial impact nearly doubled to 21.7%. The shift reflects board fatigue with productivity claims that never showed up in the operating result.
Q. What kill criteria should I include in an enterprise AI business case?
A. At minimum, three. An adoption floor (for example, below 50% of target users actively using the system at day 60 triggers a review). An accuracy floor (for example, accuracy on the golden set below 85% at day 30 triggers a model or retrieval change). A unit economic ceiling (for example, cost per resolved transaction exceeding the pre-launch baseline triggers an investigation). Pre-committed kill criteria are how disciplined enterprises stop spending against pilots that go nowhere. They are also how a CTO maintains credibility with the CFO across multiple investment cycles.
Q. How does the framework change for agentic AI versus traditional GenAI use cases?
A. The framework structure does not change. The cost weights do. Agentic AI deployments shift more cost into the evaluation, observability, and governance lines because non-deterministic behaviour requires more rigorous testing and tracing. Adoption sensitivity also rises because agents can fail more visibly than chat assistants. We typically model an agentic business case with evaluation and observability at 14 to 18% of TCO rather than 10 to 14%, and we tighten the kill criteria around accuracy and unsafe-action thresholds. The five value categories remain the same. The discipline gets stricter.
Related reading: Building an AI Centre of Excellence: The Enterprise Playbook
Nikhil Khandelwal
Co-founder & CTO, Sthambh
