Yasmin Rajabi is the Chief Operating Officer at CloudBolt Software. She is a recognized leader in the FinOps and Kubernetes communities.

​The story basically writes itself. AI demand has bent the hardware market into shapes we haven’t seen before. Big Tech is on track to spend roughly $700 billion on AI infrastructure in 2026, with some now funding that buildout through debt. Gartner analysts project a “130% surge in combined DRAM and solid-state drive (SSD) prices“ by the end of the year. And this is the cost structure enterprises will live with for the next two to three years.​

You can see it in real budgets and order queues. Lead times on big memory orders have stretched past 40 weeks. One team I talked to had to go back to the finance department and ask for six times what they’d originally scoped, not because they needed more capacity but because the same hardware now costs that much. Another had an order cancelled outright because the memory shortage has exploded and the big contracts and hyperscalers get served first.

It’s a pattern I run into constantly. A platform team walks in convinced the cluster is maxed out, and at first glance, the dashboard backs them up. Look closer, though, and the same three numbers refuse to agree. CPU utilization is in the teens, memory is below 40% and cluster allocation is above 90%. The scheduler has nowhere to place new work, yet the hardware is nowhere near exhausted. In the environments I see, the teams sitting on 70% of a cluster’s requests are often using barely a quarter of it.

That gap looks like waste, but it’s really every team doing the rational thing given how they’re measured.

The Overprovisioning Is Rational

People generally aren’t overrequesting resources because they’re careless. They do it because they remember exactly what it cost them last time. Picture an app team two years after a 2 a.m. outage. They raised the memory request, added a buffer and finally slept through the night. One OOMKill here, one peak event there and one temporary bump that never left. Over time, the cluster turns into a junk drawer of old incidents.

The incentive structure guarantees it. Being underprovisioned costs you right away and in person. You get paged, you explain the outage and your name is on the postmortem. Being overprovisioned costs nothing you can see. No invoice arrives, and no budget owner sees the line item.

What’s changed is that those buffers are usually memory, and memory is the exact commodity now repricing. Padding a CPU request was always cheap insurance. Padding a memory request at 2026 DRAM prices is a real bet. The junk drawer is now denominated in the most expensive thing in the data center.

To the team that owns the SLO, having that capacity taken back feels like ripping out the airbags because the car hasn’t crashed in a while.

Why Forced Automation Backfires

The tempting fix is to skip the negotiation. A vendor looks at the same dashboard, sees usage low and requests high, and says to turn on fully autonomous rightsizing and stop arguing with every team. But when automation gets it wrong, the people who answer for it are the engineers, not the vendor’s algorithm.

When an engineer asks what happens when this breaks (and it eventually will), that’s not being difficult. It’s a fair question—usually the right one. Unused capacity isn’t automatically reclaimable capacity; reclaiming it takes trust—trust that the recommendations are conservative, that rollback works and that no one gets blamed for going along with it.

Earn The Automation, In Order

Trust doesn’t come from a vendor’s product claim. It builds over time as the system proves itself. That means rolling automation out in stages instead of flipping it on all at once:

• Make the cost of buffers visible to the team that owns the workload.

• Surface conservative recommendations with the usage history behind them.

• Ensure intelligent, automated rollback paths exist for every change.

• Take controlled action on low-risk workloads first and allow for humans in the loop.

• Expand autonomously as it’s earned.

Every overrequested workload is a slightly different trust problem. One-size-fits-all automation treats them all the same and stalls. Automation that adapts to each team’s readiness is what truly frees up the capacity.

The Two Doors

There are really two doors. One leads to the finance department with a hardware quote, turning an operating-model problem into a capital request. The capacity you order today lands 40-plus weeks out, and the cluster absorbs it the way it absorbed the last expansion, so you’re back in the same room next year. The other leads back to the app teams. It’s slower and more political, and no purchase order fixes it, but it addresses the root cause, and it’s the only door you have to walk through once.

Finance isn’t the obstacle here. What wears them down is being asked to fund the same hole getting deeper every year. The enterprises that end the overprovisioning subsidy do three things: They make buffers visible, ask teams to own the cost and let automation earn its way in. When that happens, clusters that looked full turn out to have quarters of growth already sitting in them. That changes the conversation. The money starts going toward innovation instead of just keeping the lights on.

The crunch is real, but it didn’t create this problem. It just took away the place enterprises used to hide it.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Share.
Leave A Reply

Exit mobile version