TL;DR — Key Takeaways

  • Cloud spending keeps rising, and AI workloads are accelerating the problem, but greater visibility alone does not remove waste or change engineering behavior.
  • The bigger opportunity is to reduce the cost of growth at the application layer, where architecture, runtime efficiency and workload design determine how much infrastructure is actually required.
  • FinOps works best when cost becomes an engineering concern embedded throughout the SDLC, allowing efficiency gains to create budget headroom for future AI investment.

Every enterprise CFO I talk to can tell me, almost to the dollar, how much cloud spend their organization is wasting. They look on as every idle workload, overprovisioned environment, and underused service pushes the number higher.

Azul’s recent CFO Cloud Cost Optimization Report found that nearly nine in 10 CFOs say cloud spending is growing, and two-thirds call it a board-level issue. The default response is to demand more visibility. Over the past few years, organizations have invested heavily in tools designed to surface waste and optimize spend. Plenty of tools now provide accurate, highly granular visibility. But visibility without incentive doesn’t change anything, because knowing where waste lives is not the same as eliminating it.

Growth Always Wins

CFOs say they care deeply about controlling cloud costs. In practice, growth wins every time.

Whether a business is actively expanding or under pressure to expand, the instinct is to provision first and optimize later. In competitive markets, the fear of losing ground outweighs any efficiency argument. Cloud cost discipline becomes a last resort. It’s something organizations get serious about only when they have no other choice.

A CTO presents a compelling case for infrastructure optimization. The board nods approvingly, then approves a new AI initiative that doubles the compute footprint. The board drives the AI investment, while finance owns cost discipline. The money can’t flow in two directions at once.

This pattern is accelerating. Gartner forecasts worldwide AI spending will hit $2.59 trillion in 2026, a 47% year-over-year increase. BCG’s 2026 AI Radar found that 94% of companies plan to keep investing in AI even without immediate returns, and most are doubling their budgets. Meanwhile, Flexera’s 2026 State of the Cloud Report shows that wasted cloud spend ticked up to 29% for the first time in five years, reversing a half-decade of improvement, as AI workloads are surging faster than organizations can manage.

Work at a Different Layer

If growth will always take priority, the only realistic path is to reduce the cost of growth itself. Not to slow it down, but to make it structurally cheaper to deliver.

Most organizations are looking in the wrong place. The visibility layer tells you what was spent. It doesn’t change what gets spent. A significant leverage point is at the application layer, where compute consumption is created and where the number of servers required, how much memory each one needs, and how hard each one works are determined.

A good example comes from the early days of online travel booking. When consumers were the ones searching for flights, the ratio of searches to actual bookings was manageable, maybe three or four to one. Then automated search engines arrived, and that ratio exploded to 1,000-to-1. It completely transformed the economics of running a travel platform. The same pressure is building now as AI agents and automated workflows multiply the demands on backend systems. The infrastructure needs to be handled more efficiently, or the economics break.

For companies running Java-based workloads—which still represent the majority of production applications in large enterprises—techniques like runtime optimization and JVM tuning can directly reduce compute usage per transaction, not just report on what’s been consumed. That translates to fewer servers, smaller footprints, and lower bills in a way that’s structural rather than temporary.

Making Efficiency Fund What Comes Next

For CFOs navigating the tension between cost control and AI ambition, the most important shift is to stop treating them as competing priorities. In the right sequence, efficiency funds innovation.

Treat efficiency as an innovation enabler, not a constraint: The organizations getting this right aren’t choosing between cost discipline and AI ambition. They leverage cloud efficiency to create budget headroom, enabling sustained AI investment.

Move FinOps from reactive to integrated: For many organizations, FinOps analysis begins after the bill arrives. Instead, cost awareness needs to be embedded in architecture decisions, development workflows, and infrastructure choices from the start.

Make cost an engineering KPI: Cost can’t remain a finance-only metric reviewed at the end of the month or quarter. It needs to become an informed, shared KPI that engineering and finance teams optimize together as part of everyday decision-making.

Cloud economics erode workload by workload, agent by agent, until the bill is everyone’s problem and no one’s responsibility. The reactive dashboard was never the answer. It’s all about integrating cost visibility and discipline into every phase of the SDLC.

Frequently Asked Questions

Why isn’t cloud cost visibility enough to reduce spending?
Visibility shows organizations where money is being spent, but it does not automatically create the incentives or engineering changes needed to eliminate waste. Cost reduction requires teams to act on that information through architecture and workload optimization.
How can application-level optimization reduce cloud costs?
Improving runtime performance, memory usage and compute efficiency can reduce the number and size of servers needed to handle the same workload. For Java environments, techniques such as JVM tuning can directly lower infrastructure consumption per transaction.
How should FinOps change as AI workloads grow?
FinOps needs to move upstream from after-the-fact bill analysis into architecture, development and infrastructure decisions. Making cost a shared KPI across engineering and finance can help organizations scale AI without allowing cloud waste to grow unchecked.