Is AI Going the Cloud Way? What CIOs Must Do to Control Soaring Costs

Is AI Going the Cloud Way? What CIOs Must Do to Control Soaring Costs

AI is following a familiar path. First comes the excitement, then rapid adoption, and soon after, a hard look at economics. That is exactly what happened with cloud, and many CIOs now believe AI may be heading in the same direction.

In the early days of cloud, enterprises moved aggressively to modernize infrastructure and gain flexibility. But over time, many discovered that consumption costs were harder to control than expected. The result has been a growing wave of cloud repatriation and renewed scrutiny on cost architecture. AI, especially generative AI, appears to be entering a similar phase.

Puneesh Lamba, CIO, CMR Green Technologies, believes the comparison is valid. “Yes, in some ways it will. Cloud also went through hype, over-adoption, cost issues, and then some repatriation,” he says. “However, even now, cloud adoption is still far higher than it was 20 years ago. So it has clearly stayed relevant.”

That, he argues, is exactly how AI may evolve. “AI will probably go through a similar cycle. Some companies will rush in, some will pull back, and some will rehire engineers after overestimating what AI can do,” Lamba says. The lesson for CIOs is not to avoid AI, but to adopt it with a sharper operating model.

The biggest misconception today, according to Lamba, is that AI can replace the full engineering stack end to end. “The biggest challenge with AI today is that people believe it can do complete coding end to end. We are not there yet,” he says. “There were still plenty of bugs, a number of quality concerns, a lack of documentation, concerns with scaling, and issues with integration.” 

This disconnect between promises and actual value is just one element that many enterprises have been already evaluating. Engineers aren’t merely focused on the quality, but on tangible value as well. As AI workloads expand, so do token bills, inference costs and infrastructure requirements.

Strategies to Curb Costs

Lamba points out that the market is beginning to respond. “On token costs, competition will eventually help,” he says. “We started with OpenAI, then Claude, then DeepSeek, and now newer models like Kimi K3 and other open-source options are bringing costs down.” 

But Lamba says CIOs should not think only in terms of model pricing. They must also redesign the way AI is consumed inside the enterprise. “The answer is not just about negotiating lower token prices,” he says. “It is about smart architecture — designing systems so that AI APIs are called only when absolutely needed.”

In practice, that means using the right model for the right task. “Enterprises should use cheaper models for low-end or routine work, and reserve costlier models only for complex logic or high-value use cases,” Lamba explains. He also points to the growing viability of self-hosted models such as Llama, Qwen and Gemma, which can help organizations reduce or even eliminate recurring token costs. “If the use case allows it, self-hosting is a strong option because it gives you more control over cost, deployment and governance,” he adds.

His recommendation is clear: build for flexibility. “The trick here is building a model-agnostic architecture,” he said. “You should be able to flip from LLM1 to LLM2 without changing your entire application. That interaction around prompts and context and instruction has to sit outside the model itself.” 

Tejasvi Addagada, Head of AI Governance, HDFC Bank, takes a similar view, though from a different angle. “The only thing for token costs to be controlled, in my view, is better prompting,” he says. “Prompt engineering as a discipline and as a science can get us to a better way of prompting an LLM.”

Addagada also points to the importance of scale and infrastructure efficiency. “We need to look at two aspects: infrastructure and compute,” he says. “Infrastructure is governed by certain contingencies — it’s a commodity, and contingencies define commodity pricing.” On the compute side, he believes organizations can still create meaningful efficiencies. “There are certain efficiencies that we can bring into our inferencing systems, which can ideally cut short on our inference workloads, token costs, by at least 50%,” he says.

Which leads to the conclusion that the next stage of AI will be about operational execution, not about hype. CIOs will need to be system architects, not just system shoppers. This will involve making tradeoffs among model selection, prompt quality, inference speed and efficiency, infrastructure choices and governance simultaneously.

Those enterprises that manage to be disciplined early and continue to engineer carefully for ongoing flexibility will probably derive the greatest value from generative AI. This technology is not something you can afford to try and then unplug. Like cloud, it is here to stay. But the enterprises that win will be the ones that control the cost curve before the cost curve controls them.

Share:

Author

Yashvendra Singh

Yashvendra is Editor at CIONow.in, with over two decades of experience covering enterprise technology, business, and the CIO community.

Chat with CIONow.in

Powering The Intelligent Energy Enterprise...