on
Council Post: You're Measuring Your AI Like Software, But It's Behaving Like Labor
Ajay Pundhir, senior AI leader & founder of AskAjay.ai. Believes AI should amplify human expertise, not replace it.

getty
KPMG's Global AI Pulse survey this spring reached more than 2,000 senior leaders across 20 countries, and it found something that should stop every board in its tracks. Only 7% could say they had established a return on their AI investment, and 42% admitted they had only partial visibility into what they were even spending. Nearly a quarter said investors are already pressing them to prove the value.
Most people read those numbers as a discipline problem. Companies bought too fast and tracked too little, and now the bill has arrived. That reading is comfortable, and it is wrong. The problem isn't that finance is being sloppy. The problem is that finance is applying the right tool to the wrong category.
We are measuring AI like software. It is starting to behave like labor.
The Category Error
Think about how a CFO has evaluated software for the past 30 years. You buy a license or a subscription. You amortize the cost across its useful life. You count the seats, watch adoption and compare the per-seat price against the productivity it unlocks. The software sits on a desk. A person uses it to do work faster. The tool is a multiplier of human effort, and the entire ROI model is built on that assumption.
An agent breaks the assumption. It doesn't sit on the desk waiting to be used. It does the work. It reads the contract, drafts the reply, runs the reconciliation and closes the ticket. When the thing you bought produces output instead of enabling it, the seat is no longer the unit that matters. The unit of work is.
This is why payback-period math misleads. Traditional software costs are front-loaded and then flat. You pay, you deploy and marginal use is close to free, so the model rewards you for putting more people on the same license. Agentic systems invert that.
The cost is consumption-based and scales with every task the agent performs, and much of the value shows up as work that no longer needs redoing, which never appears as a line item anywhere. Gartner's 2026 market guide for enterprise AI coding agents describes a market entering a new phase defined by more complex pricing and ROI dynamics, with vendors shifting from seat-based licensing to consumption-based billing. That shift isn't a pricing footnote. It's a signal that the underlying economics changed and the measurement hasn't caught up.
Three Questions To Ask
If you run finance, you don't need a new framework to start fixing this. You need three questions, and you can ask them about any agent already running in your business:
1. What is the work unit this agent actually produces? Not "supports the sales team." Tickets closed. Contracts reviewed. Reconciliations completed. Invoices matched. If you can't name the countable thing the agent finishes, you cannot price it, and you are flying blind no matter how good the dashboard looks.
2. What is the fully loaded cost per work unit? Tokens are the number everyone quotes because it's the number the vendor shows you. It is also the smallest part of the bill. Load in the orchestration layer, the tooling, the integrations and, above all, the human supervision, because someone is reviewing, correcting and unblocking that agent, and their time is real money. The honest denominator is total cost divided by units of finished work, and it is usually several times what the token line suggests.
3. What is the rework rate? This is the one almost nobody tracks, and it decides everything. When the agent's output is wrong, incomplete or escalated to a human, that's a quality tax on every unit it produces. A support agent that closes tickets that reopen two days later hasn't closed them. It has deferred them and added a handoff.
The Trap This Catches
Here is the failure that software math is structurally unable to see. Picture a support agent with genuinely cheap tokens. Per interaction, it costs a fraction of what a human costs, and the per-seat comparison looks like an easy win. Leadership celebrates. Then you measure the work unit and find a 40% rework rate. Four in 10 of its "resolved" tickets come back, get re-escalated or need a human to redo them properly.
That agent is not saving money. Counting the rework, the cleanup and the customer trust it burns on the way, it is net-negative. But nothing in the traditional model surfaces that, because the traditional model was watching cost-per-seat, and cost-per-seat looked great. Measure the same agent on throughput and quality, and the picture inverts. The cheap agent is the expensive one.
This is how firms misprice agents in both directions at once. They overpay for agents that look productive and are quietly destroying value, and they kill agents that look expensive on tokens but produce clean, low-rework output that would have paid for itself many times over. Same broken denominator, opposite mistakes.
Why The 7% Won't Move
The 7% figure is not going to climb because the models get better. The models are already completing tasks that would take a human hours of focused work. The bottleneck was never the intelligence. It is the ledger.
Until finance measures agents the way it has always measured labor, by throughput, by quality and by what gets reinvested when the work comes back clean, the proof of ROI will stay stuck in single digits. The good news is that this is a solvable problem, and it lives on your side of the table, not the vendor's. Start with the work unit. Everything else follows from naming the thing the agent finishes and pricing it honestly.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?