Everyone’s Watching GPU Prices. The Memory Shortage Is Coming for Your Regular Cloud Bill.

SODIMM RAM stick

Octave Klaba, the founder and chairman of OVHcloud, told the market to expect cloud prices to rise 5 to 10 percent between April and September of this year. Not GPU prices. Not AI-cluster prices. The price of ordinary servers, the compute-optimized and general-purpose instances that run your web tier, your queues, your internal apps, and your databases. His reasoning was blunt: server hardware costs are climbing 15 to 25 percent because memory has become the single most expensive component in the box, and a hosting provider cannot absorb that forever.

Most FinOps teams spent the first half of 2026 fixated on the GPU line. That is where the drama has been, and I understand the pull. But the memory shortage that is driving GPU capacity prices up is the same shortage that is about to land on the boring 80 percent of your bill, and almost nobody has modeled it into a forecast.

What “RAMageddon” actually did to component prices

The numbers are not subtle. AI server DRAM roughly doubled in the first quarter of 2026 and is on track to quadruple across the full year, according to Deloitte’s analysis of the memory crunch. DDR4, the workhorse memory behind a large share of the instances cloud providers still run, is up more than 1,300 percent since April 2025. Samsung and SK Hynix raised server DRAM contract prices 60 to 70 percent versus the fourth quarter of 2025, per TrendForce data compiled by Hostkey.

The cause is straightforward and it is not going away on its own. Samsung, SK Hynix, and Micron have redirected fabrication capacity toward high-bandwidth memory for AI accelerators, because HBM earns three to five times the revenue per wafer that conventional DDR5 does. Deloitte estimates memory now consumes roughly 30 percent of hyperscaler data center capex in 2026, rising to 36 percent in 2027, and new fabrication plants take three to five years to build and scale. Deloitte’s own word for the situation is “RAMageddon,” and its forecast for meaningful relief is 2029 to 2030.

This is the part that matters for anyone building a cloud budget: it is a rate problem, not a usage problem, and it has a multi-year floor under it.

Why the non-GPU bill is the exposure nobody priced

Back in April I wrote a FinOps playbook for the hardware cost passthrough while the increases were still mostly a forecast. Four months later the shape is clearer, and the interesting detail is where the pain concentrates.

Break the passthrough down by service class and the segments look like this, per the TrendForce and Hostkey breakdown:

  • GPU-based servers: up 30 to 50 percent
  • CPU servers: up 10 to 15 percent
  • Dedicated servers: up 10 to 20 percent
  • Cloud compute-optimized instances: up 3 to 7 percent
  • High-memory managed services (Redis, ElastiCache, large-RDIMM database VMs): the most exposed of all

GPUs get the headline percentage. But GPU instances are a minority of most enterprise bills. The 3 to 7 percent creeping onto compute-optimized instances, and the steeper climb waiting for memory-heavy managed services, applies to the line items that make up the majority of what a typical company actually spends. A 5 percent rate increase on 80 percent of your bill moves more real dollars than a 40 percent increase on the 10 percent of your bill that is GPU.

The hyperscalers have been quiet about non-GPU pricing so far. In April, Gartner’s Tony Harvey noted to The Register that the major cloud vendors had not yet raised prices on standard servers, which is exactly why the migration-to-cloud pitch still worked. AWS has been happy to lean on that. CEO Andy Jassy has framed the memory shortage as a reason to move into the cloud, arguing AWS is “not capacity constrained” because it planned supply ahead. That framing is doing a lot of work. Where AWS has been visibly repricing is the AI corner: it raised EC2 Capacity Block rates for ML roughly 15 percent in January and roughly 20 percent again in July of 2026. Providers reprice where demand is most inelastic first. Commodity compute is simply next in line, on a delay, and Klaba said the quiet part out loud.

This is not a problem you optimize your way out of

Here is the uncomfortable framing for a FinOps team. Most of the muscle we have built over the past few years operates at the usage layer. Right-size the instance, kill the idle resource, cache the response, route the cheap model, delete the orphaned volume. All of that reduces how much you consume.

A memory-driven rate increase does not care how efficient your consumption is. If the per-hour price of a compute-optimized instance rises 5 percent, your beautifully right-sized fleet costs 5 percent more. You cannot prompt-cache your way out of a DDR5 shortage. This lands squarely in the rate optimization half of FinOps, which is the half most teams under-invest in because it is less satisfying than hunting waste.

In 20-plus years running IT operations, the pattern that held every time a component market tightened was the same: the organizations that got hurt were the ones carrying everything on on-demand rates, betting that prices only ever fall. Rate discipline is boring insurance right up until the moment the market moves, and then it is the only thing that matters.

Four moves for the second half of 2026

Reprice the forecast, separating rate from usage. Most cloud forecasts implicitly assume flat unit rates and model only usage growth. Break those apart. Run a scenario where non-GPU compute rates rise 3 to 7 percent in the back half of the year and memory-heavy services rise more. If your forecast has no rate variable in it at all, that is the first thing to fix, and it is a one-afternoon change to the model.

Treat commitment coverage as a rate hedge, not just a discount. A Savings Plan, Reserved Instance, or committed use discount does one thing people forget in calm markets: it freezes today’s rate for one to three years. In a rising-rate environment, locking current pricing on your stable, predictable baseline is worth more than the headline discount percentage. This is the moment the commitment-versus-on-demand math tilts hardest toward committing. The discipline is to commit on the baseline you are confident about and leave headroom on the volatile top, not to over-commit to instance families you may still want to right-size.

Audit memory you are paying for and not using. Memory-optimized services are the sharpest exposure, so rightsizing memory is now higher-leverage than rightsizing CPU. Find the ElastiCache and Redis nodes provisioned for a peak that never comes, the database VMs sitting on oversized RDIMM footprints, the caching layers nobody has revisited since launch. Every gigabyte of provisioned memory you shed is a gigabyte insulated from the steepest part of the curve.

Re-run the repatriation math, but read it honestly. The instinct when cloud gets pricier is to look at buying your own hardware. That door is narrower than it looks right now. Servers cost roughly four times what they did a year ago, hard drive makers have reportedly sold out their entire 2026 output, and Meta has extended its own server lifecycle from six years to seven because supply is that tight. Jassy is not wrong that the cloud looks attractive against that backdrop. The catch is that you are renting into a market where rent is rising. Neither escape is clean, and a repatriation decision made in a panic about this quarter’s prices will not survive the next two years.

The line item to watch

The State of FinOps 2026 survey found 84 percent of organizations naming cloud cost management as their top challenge, and it found teams pouring attention into AI spend, which 98 percent now manage in some form. Both are real. But the memory shortage is a reminder that the discipline’s oldest job has not gone away. The commodity compute line, the one that is not exciting and not AI and not on any conference slide, is where a multi-year supply shock is about to show up on the majority of bills.

Watch the GPU number if you want. Just do not let it be the only number you are watching when the invoice for your regular fleet comes in a few points higher and your forecast said flat.

ty247

Ty Sutherland is the Chief Editor at Kost Kompass. With 25 years of experience in enterprise strategy and financial management, Ty Sutherland is the driving force behind kostkompass.com. Specializing in helping Finance and Technology Managers optimize costs in servers, cloud, and SaaS, Ty combines technical acumen with financial discipline to deliver actionable insights for cost-effective solutions.

Recent Posts