Corporate AI investment reached $252.3 billion in 2024, according to the 2025 AI Index Report (Stanford, 2025). Against that, a 2024 survey of C-suite executives on AI value (BCG, 2024) — 1,000 executives across 59 countries — found 74% of companies have yet to show tangible value from their use of AI. The obvious reading is that AI does not work, and it is the wrong one. That figure is not a model-quality statistic: it measures the distance between a system that runs and a business that has changed, and that distance is crossed, or not crossed, entirely on the far side of launch.
A business case has to name a date for the return. That date is the wrong variable to manage. Every dollar of an AI investment is committed before launch and every dollar of return is earned after it — in the stretch when the build is finished, the project budget has closed, and the system belongs to nobody in particular. Payoff is not a function of elapsed time; it is a function of whether anyone is still holding the system when the return is supposed to arrive.
That claim has an address. ML LABS engineered the backend of a medical-device cloud platform that reached clinical production in two countries — the cloud backend behind HeartSciences' MyoVista Insights, with AIM Consulting building the frontend — and it runs today as HeartSciences' core product across multiple healthcare organizations in US and UK production. Their Director of Software Engineering put it on the record: "Omar did outstanding work designing and building the backend for our cloud-native AI-ECG platform." What we will not do is publish their numbers — study volumes, latencies, and availability are the client's operational data, and an article about return on investment that leaks them is proving the wrong thing.
The Unowned Payoff Window
What generalizes from that build is the shape of its return, not its metrics. The platform was designed so that onboarding a new organization is configuration and never code — HL7 field mappings, enabled AI models, invoice pricing, storage provisioning — so hospitals, device vendors, and model providers get added without an engineering project each time. None of that paid anything on launch day; it pays as the platform grows, for as long as someone keeps the property true. An AI return is not an event but a position that has to be held.
Holding it is work, and the research has said so for a decade. The observation that ML systems accrue hidden debt without monitoring and feedback (NeurIPS, 2015) established that the model is a small fraction of a production ML system; the rest is the machinery that keeps it attached to reality. The model does not hold still either: research on temporal degradation in deployed models (Scientific Reports, 2022) tested 128 model-dataset pairs across healthcare, weather, finance, and transportation and found 91% of them degrading over time, many with no concept drift in the data at all. Those systems were not attacked from outside — they aged.
So the payoff window and the ownership vacuum are the same window. Costs are legible from day one: invoices, cloud bills, a line in the plan. The return is illegible unless someone builds the instrument that makes it visible and keeps that instrument honest while the system and the business move underneath it — the discipline the measurement article is about. An unmeasured return is indistinguishable, to whoever signed the check, from no return at all, and spend drifts upward in the same blind spot, which is its own failure mode with its own economics.
graph TD
A["Budget approved<br/>cost is legible"] --> B["Build<br/>staffed, owned"]
B --> C["Launch"]
C --> D["Return window<br/>no owner"]
D --> E["Model ages,<br/>spend drifts"]
D --> F["Return measured,<br/>defended, extended"]
E --> G["Investment unproven"]
F --> H["Investment pays off"]
style A fill:#1a1a2e,stroke:#0f3460,color:#fff
style B fill:#1a1a2e,stroke:#0f3460,color:#fff
style C fill:#1a1a2e,stroke:#ffd700,color:#fff
style D fill:#1a1a2e,stroke:#ffd700,color:#fff
style E fill:#1a1a2e,stroke:#e94560,color:#fff
style F fill:#1a1a2e,stroke:#16c79a,color:#fff
style G fill:#1a1a2e,stroke:#e94560,color:#fff
style H fill:#1a1a2e,stroke:#16c79a,color:#fffReturns We Can Actually Show
Fifteen-plus years and more than ten heavy-workload systems in production sit behind that argument, across healthcare, telecom, PropTech, and finance. Three of those engagements produced returns ML LABS is cleared to state, and no two of them paid off for the same reason.
The roaming optimization platform built for a top 10 global telecom company, engaged through Gigster, delivered over 12x return against engagement cost in the first year. It paid fast because its output was the decision rather than a report about the decision — routing moved from a manual review cycle the analysts could not finish before the traffic pattern shifted, to decisions the platform made on its own. It also surfaced optimization corridors across 128% more of the network than the original scope targeted, corridors the manual team had never reached because roughly 1TB of new data a day made reaching them physically impossible. Nobody had to change their behavior for that return to land; the system executed.
The property valuation engine built for a PropTech platform hit its accuracy target — within 10% of closing price in dense metro areas for 90% of cases, in seconds instead of days — and the accuracy is not what made it pay. The confidence band did: the calibration deciding which properties the system was allowed to price automatically and which ones were routed to a human appraiser. Calibrating that band correctly was harder than producing the point estimate, and it is what made the 90% commercially meaningful. Their Founder/CEO's published account is that the system "became the reason investors took us seriously" — the return was a fundraise, not a line item.
A model that knows where it is wrong can be automated around. A model that does not know is a demo with a good headline number.
The third return was never designed at all. A hedge fund was storing large volumes of unnecessary and polluted data without noticing: the aggregation step discarded information the models needed, and the storage structure duplicated what it kept. Correcting it cut storage costs by more than 60%, and their models performed 2% better afterward. Their Head of Data's published account ends "it held up in production — reliable in a way this field rarely is", and the full telling belongs with the ownership argument, because that is what it is evidence of.
The return was not created by that engagement. It had been sitting in the data platform the whole time, worth precisely nothing until somebody went looking for it.
Three systems, three returns, and in none of them did the payoff arrive because the model was good. It arrived because the output was wired straight to a decision, because the system knew when to disqualify itself, or because someone was still examining a platform everyone else had stopped examining. Research on AI and decision-making (HBR, 2021) gives the general form: value realization turns on whether the AI changes decision-making behavior, not on whether it is accurate. That is also why returns lag — research on the AI productivity paradox (NBER, 2018) traces the J-curve to intangible capital that accumulates after the technology is installed, and research on generative AI and workplace productivity (NBER, 2023) found its 14% average productivity gain emerging through use rather than at deployment.
When The Return Cannot Exist
All of this assumes there is a return there to collect. Sometimes there is not, and the most valuable thing an engineering partner can do is say so before the money moves. A major US TV network arrived at a scoping call holding a quote to build a full software system for a workflow that did not require one, and the deliverable of that $750 call was "don't build this" — their AI Program Manager's published account is that it "saved us from a $200K mistake." The best return described in this article belongs to an investment that was never made.
The gate is not whether AI can perform the task. It is whether the cost of the problem is quantifiable today. When nobody can state what the current process costs in labor, errors, delay, or lost revenue, there is no denominator and no way to prove a return later even if one shows up — which makes the quantification itself the honest first investment.
First Steps
- Name the return and its owner. Write down the business number the system is supposed to move, and the person accountable for that number once the build team is gone. A blank second half makes the first half decoration.
- Build the instrument before the handover. The measurement has to exist while the people who understand the system are still attached to it — the query, the reconciliation, the recurring number tying system output to the business outcome.
- Re-score the model against fresh outcomes. Take a recent window of live predictions, score them against what actually happened, and compare with launch-day performance. Aging is the default case, not the exception.
Ownership As The Return Mechanism
Match the move to where you actually stand. If nothing is live, the work is quantifying the problem, and no amount of ownership creates a return that was never there. If a system is live and the number it was meant to move is still theoretical, the missing piece is not a better model — it is a standing owner: someone accountable who watches drift, spend, and failure classes, keeps the measurement honest, and ships what the measurement asks for. And if several systems are live but nobody can say which of them is still earning, the gap is structural, which is where research on enterprise AI maturity (MIT CISR, 2025) gets uncomfortable: the financial performance separating leaders from everyone else shows up in the move from piloting to scaled operations, not in the models.
Returns are held, not banked. Holding them is what managed AI operations exists to do — engineering ownership that stays attached to the system through the window where the money actually arrives, keeping the measurement honest, catching decay before it becomes a rewrite, and going looking for the returns nobody budgeted for, the way a hedge fund's storage bill turned out to be one. The investment has already happened. The systems that pay off are the ones somebody is still holding when the return shows up.
References
- Stanford HAI. AI Index Report 2025. Stanford University, 2025.
- Boston Consulting Group. Where's the Value in AI?. BCG, 2024.
- Sculley, D., et al. Hidden Technical Debt in Machine Learning Systems. NeurIPS, 2015.
- Vela, D., Sharp, A., Zhang, R., Nguyen, T., Hoang, A., and Pianykh, O. S. Temporal Quality Degradation in AI Models. Scientific Reports, 2022.
- De Cremer, David, and Garry Kasparov. AI Should Augment Human Intelligence, Not Replace It. Harvard Business Review, 2021.
- Brynjolfsson, E., Rock, D., and Syverson, C. Artificial Intelligence and the Modern Productivity Paradox. NBER Working Paper, 2018.
- Brynjolfsson, E., Li, D., and Raymond, L. Generative AI at Work. NBER Working Paper, 2023.
- MIT CISR. Enterprise AI Maturity Update. MIT Center for Information Systems Research, 2025.
