TL;DR — Key Takeaways
- Cloud spending keeps rising, and AI workloads are accelerating the problem, but greater visibility alone does not remove waste or change engineering behavior.
- The bigger opportunity is to reduce the cost of growth at the application layer, where architecture, runtime efficiency and workload design determine how much infrastructure is actually required.
- FinOps works best when cost becomes an engineering concern embedded throughout the SDLC, allowing efficiency gains to create budget headroom for future AI investment.
Every enterprise CFO I talk to can tell me, almost to the dollar, how much cloud spend their organization is wasting. They look on as every idle workload, overprovisioned environment, and underused service pushes the number higher.
Azul’s recent CFO Cloud Cost Optimization Report found that nearly nine in 10 CFOs say cloud spending is growing, and two-thirds call it a board-level issue. The default response is to demand more visibility. Over the past few years, organizations have invested heavily in tools designed to surface waste and optimize spend. Plenty of tools now provide accurate, highly granular visibility. But visibility without incentive doesn’t change anything, because knowing where waste lives is not the same as eliminating it.
Growth Always Wins
CFOs say they care deeply about controlling cloud costs. In practice, growth wins every time.
Whether a business is actively expanding or under pressure to expand, the instinct is to provision first and optimize later. In competitive markets, the fear of losing ground outweighs any efficiency argument. Cloud cost discipline becomes a last resort. It’s something organizations get serious about only when they have no other choice.
A CTO presents a compelling case for infrastructure optimization. The board nods approvingly, then approves a new AI initiative that doubles the compute footprint. The board drives the AI investment, while finance owns cost discipline. The money can’t flow in two directions at once.
This pattern is accelerating. Gartner forecasts worldwide AI spending will hit $2.59 trillion in 2026, a 47% year-over-year increase. BCG’s 2026 AI Radar found that 94% of companies plan to keep investing in AI even without immediate returns, and most are doubling their budgets. Meanwhile, Flexera’s 2026 State of the Cloud Report shows that wasted cloud spend ticked up to 29% for the first time in five years, reversing a half-decade of improvement, as AI workloads are surging faster than organizations can manage.
Work at a Different Layer
If growth will always take priority, the only realistic path is to reduce the cost of growth itself. Not to slow it down, but to make it structurally cheaper to deliver.
Most organizations are looking in the wrong place. The visibility layer tells you what was spent. It doesn’t change what gets spent. A significant leverage point is at the application layer, where compute consumption is created and where the number of servers required, how much memory each one needs, and how hard each one works are determined.
A good example comes from the early days of online travel booking. When consumers were the ones searching for flights, the ratio of searches to actual bookings was manageable, maybe three or four to one. Then automated search engines arrived, and that ratio exploded to 1,000-to-1. It completely transformed the economics of running a travel platform. The same pressure is building now as AI agents and automated workflows multiply the demands on backend systems. The infrastructure needs to be handled more efficiently, or the economics break.
For companies running Java-based workloads—which still represent the majority of production applications in large enterprises—techniques like runtime optimization and JVM tuning can directly reduce compute usage per transaction, not just report on what’s been consumed. That translates to fewer servers, smaller footprints, and lower bills in a way that’s structural rather than temporary.
Making Efficiency Fund What Comes Next
For CFOs navigating the tension between cost control and AI ambition, the most important shift is to stop treating them as competing priorities. In the right sequence, efficiency funds innovation.
Treat efficiency as an innovation enabler, not a constraint: The organizations getting this right aren’t choosing between cost discipline and AI ambition. They leverage cloud efficiency to create budget headroom, enabling sustained AI investment.
Move FinOps from reactive to integrated: For many organizations, FinOps analysis begins after the bill arrives. Instead, cost awareness needs to be embedded in architecture decisions, development workflows, and infrastructure choices from the start.
Make cost an engineering KPI: Cost can’t remain a finance-only metric reviewed at the end of the month or quarter. It needs to become an informed, shared KPI that engineering and finance teams optimize together as part of everyday decision-making.
Cloud economics erode workload by workload, agent by agent, until the bill is everyone’s problem and no one’s responsibility. The reactive dashboard was never the answer. It’s all about integrating cost visibility and discipline into every phase of the SDLC.

