Council Post: Scaling AI, Scaling Waste: Are You Paying Machines To Repeat Themselves?

Erum Manzoor is CEO of BOX Motorsports & Founder of VectorE Ventures, focused on motorsport, AI strategy, & high-pressure execution systems.

getty

​AI is getting expensive in a surprisingly ordinary way: We keep paying it to solve the same problem over and over again.

Inside a large company, familiar questions return with slightly different wording. Travel policy becomes reimbursement, then an exception request, then a manager asking what can be approved. The system may still load the same rules, pull the same documents and work its way back to roughly the same answer.

One request won't worry anyone. When repeated across teams, customers and agent workflows, however, it becomes part of the operating bill.

That bill is already moving quickly. A January 2026 BCG report noted that companies expected AI spending to rise from 0.8% of revenue to about 1.7% in 2026. However, a November 2025 McKinsey global survey found that only 39% of respondents had seen AI affect company-wide operating profit, and most of that group said the contribution remained below 5%. A May 2025 IBM survey of CEOs also uncovered that only 25% of AI initiatives had delivered the expected return and that 16% had scaled across the enterprise.

To me, those numbers show how easily activity gets mistaken for value.

Klarna is a useful example because the story didn't stop at the first headline. The company said in May 2024 that GenAI had helped cut marketing costs by about $10 million a year. By September 2025, its leadership acknowledged that it had leaned too heavily on AI-led cost-cutting and began shifting more attention toward products, services and growth. Klarna cut costs successfully, but the trade-off became visible once the wider business felt it.

The P&L is less impressed by the excitement around AI. A model can speed up work, reduce a unit cost or remove a recurring expense, but the "value" has to survive the wider business. Any AI saving has to be weighed against what happens to customer experience, revenue and the consequences of a poor decision.

Even when the use case works, how efficiently is the system producing that value? The simplest AI answer uses computing power to read the question and produce a response. In larger business systems, the AI may also search documents, use other tools or check its own work before answering. The waste begins when it repeats the same steps for information the company already has instead of reusing approved answers or saved context.

The scale around AI makes that difficult to dismiss. The International Energy Agency has projected global data center electricity use to roughly double to about 945 terawatt-hours by 2030, with AI being the biggest driver of the increase. McKinsey estimated in April 2025 that data centers could require about $6.7 trillion in capital spending over the same period.

Those figures cover a much larger infrastructure shift than one enterprise assistant, but they show the direction clearly. Compute, memory and energy all sit inside the AI business case whether leaders put them on the slide or not.

Much of the avoidable waste comes from architecture choices. Prompt caching can reuse repeated instructions or other identical content sent to the model. When people ask the same thing in different words, semantic caching or another application layer can recognize the shared intent and reuse an approved answer where it's safe to do so. Architecture also decides which model receives a routine task and how many steps an agent is allowed to take.

A policy answer doesn't need to be invented every time. The approved source can stay stable while the response adjusts for role, location or circumstance. A familiar customer service issue may need the right process and a little context rather than the most expensive model available. Similarly, a repeated internal request may be handled through retrieval or a smaller model before deep reasoning ever enters the picture.

Standardization sounds dull until the bill arrives. Good standardization reuses what stays the same and adjusts the answer only where the context changes. It gives companies a way to decide where fresh reasoning earns its cost and where the system can build on work already completed.

The technology already supports some of this. According to OpenAI, prompt caching can reduce input token costs by up to 90% and time to first token by up to 80% when requests reuse content the system has recently processed. A January 2026 PwC study of long-running agent tasks across OpenAI, Anthropic and Google found cost reductions between 45% and 80% when caching was used effectively.

Agents make the decision harder because they can quietly turn one request into a chain of planning, searching, retrieving, tool calls, checking and rewriting. That may make sense for difficult work. For routine questions, it can become an expensive way to reach information the company already has.

The leadership question is becoming quite practical. Does this task need reasoning, retrieval or reuse? Is the largest model actually improving the outcome? Are 10 agent steps producing value or simply producing more activity?

Companies spent the last few years asking where AI could be added. The next phase should focus on how much unnecessary compute can be removed without lowering the quality. The companies that handle this well will still use powerful models. They'll also know when the system needs to think, when it should build on work already done and when it's done enough.


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?