Flexera’s 2026 State of the Cloud report put wasted cloud spend at 29%, up from 27% the year before. That is the first increase in five years, snapping a steady decline. For an enterprise spending $10 million a year on cloud infrastructure, $2.9 million goes to resources nobody is using, instances nobody right-sized, and commitment discounts nobody purchased.
The reversal has a cause, and the cause is AI. The FinOps Foundation’s State of FinOps 2026 survey, built on 1,192 respondents representing more than $83 billion in annual cloud spend, found that 98% of teams now manage AI spend. That number was 63% a year earlier and just 31% two years ago. No category the Foundation tracks has moved that fast. GPU instances that cost $30 to $50 an hour sit idle at single-digit utilization while teams hold capacity they grabbed during the scarcity of 2023 and 2024. The traditional waste categories did not improve; they were simply overshadowed by a new one that nobody had tooling for.
This guide breaks cloud waste into six auditable categories, gives you a measurement framework with specific thresholds, and maps the native and third-party tools worth using in late 2026. In twenty years running IT operations budgets across organizations spending anywhere from $500K to $50M a year, the pattern has held: most waste is findable within 30 days and eliminable within 90. What changed this year is that one of the six categories now moves faster than you can tag it.
Six Categories of Cloud Waste
1. Idle Resources (8 to 15% of Spend)
Provisioned resources running with minimal utilization. EC2 instances at 2% CPU for months. Unattached EBS volumes accumulating storage charges. Load balancers forwarding zero traffic. AWS Cost Optimization Hub, a free feature, consolidates idle resource detection across Compute Optimizer, Trusted Advisor, and Cost Explorer into a single deduplicated dashboard. AWS has since layered a Cost Efficiency score on top: a 0-to-100 metric now surfaced in the Billing and Cost Management console. AWS’s State of Cost Efficiency Report, drawn from more than 71,000 anonymized opted-in customers, pegged the median score at 83 and the mean at 79, which gives you an external benchmark for whether your cleanup is holding or quietly regressing.
The threshold: flag anything with average utilization below 20% and P95 below 40% over a 30-day window. Well-optimized environments maintain 40 to 60% average CPU across their compute fleet. Treat a rising Cost Efficiency score as confirmation, not a finish line; the metric only grades the waste AWS already knows how to detect.
2. Oversized Resources (10 to 20% of Spend)
Over-provisioning driven by the “just in case” habit engineering teams carry over from on-premises infrastructure. The gap between provisioned capacity and actual utilization commonly exceeds 60%. An m5.2xlarge running a workload that peaks at 15% CPU and 30% memory should be an m5.large, cutting the cost by 75%.
Right-sizing is the single highest-impact optimization most organizations can make. AWS Compute Optimizer, Azure Advisor, and GCP Recommender all provide instance-level recommendations with projected savings. The challenge is governance, not detection: every cloud provider can tell you which instances are oversized. The open question is who acts on the recommendation and how quickly. This is exactly the gap the new native agents are trying to close, more on that below.
3. Pricing Model Misalignment (15 to 25% of Spend)
Running steady-state workloads on on-demand pricing when Reserved Instances or Savings Plans would cost 40 to 72% less. This is the largest single category of waste in most organizations.
The benchmark: mature FinOps programs reach 60 to 70% commitment coverage across stable workloads (FinOps Foundation, 2026). Organizations using less than 60% of their reserved capacity are effectively paying more than on-demand on the committed portion, which means the cure became the disease.
There is a 2026 wrinkle worth naming. With capacity genuinely scarce (see category six), commitments have taken on a second job. A Zonal Reserved Instance or an EC2 Capacity Block no longer buys only a discount; it reserves physical hardware you might otherwise queue for. The commitment decision has shifted from a pure discount calculation to a capacity hedge. Misalignment still wastes money, but under-committing on scarce GPU families now risks something worse than overspend: not getting the hardware at all.
4. Architectural Inefficiency (5 to 15% of Spend)
Data transfer costs that compound silently. Cross-AZ traffic at $0.01/GB on AWS adds up when services talk to each other millions of times a day. Synchronous processing patterns that keep expensive compute running during I/O waits. Batch jobs scheduled during peak-rate hours when off-peak pricing exists.
These require deeper analysis than the earlier categories but produce sustainable, compounding savings. A single data pipeline rearchitected to avoid cross-region transfers can save more annually than months of instance right-sizing. AWS PrivateLink and GCP Private Service Connect can eliminate egress charges for internal API traffic. Understanding egress cost dynamics matters even more now that the EU Data Act is reshaping which transfer fees vendors can charge and which they cannot.
5. Zombie Resources (3 to 8% of Spend)
Resources from decommissioned projects, departed employees, or failed experiments. Enterprises commonly carry 10 to 15% of their cloud footprint as zombies. These are the easiest to find (query for resources with zero connections, zero requests, or no active users over 90 days) and the hardest to delete, because nobody wants to own the decision.
Tag enforcement through AWS Service Control Policies or Azure Policy prevents the next generation of zombies. The cleanup is a one-time project; the governance is permanent. Zombie AI resources are a growing subset: orphaned vector databases, abandoned fine-tuning jobs, and inference endpoints left running after a proof of concept died. Roughly a third of AI spend sits outside normal IT visibility, which makes AI zombies harder to see than the EC2 kind.
6. GPU and AI Workload Waste (10 to 30% of AI Spend)
This category barely existed before 2024. It is now the fastest-growing source of cloud waste and the single reason the industry waste rate ticked upward in 2026.
Cast AI’s 2026 State of Kubernetes Optimization Report, drawn from tens of thousands of real workloads, found average GPU utilization of just 5%. The breakdown by provider is bleak and consistent: AKS clusters average 2%, EKS 5%, GKE 6%. At 5% utilization, the effective cost of a GPU runs roughly 20 times its nominal hourly rate. A cluster of 20 H100s at 20% utilization still burns around $200,000 a year in idle compute alone, and most fleets do far worse than 20%.
Here is where the common story is now wrong. Through 2024 the problem was defensive over-provisioning during a shortage; teams reserved capacity before workloads existed and the industry assumed the scarcity would ease. It did not ease. It intensified. AWS reported a backlog above $240 billion, up roughly 40% year over year; Microsoft disclosed tens of billions in unfilled Azure orders, with GPUs sitting in warehouses it lacks the power to install; Blackwell lead times run 36 to 52 weeks, and hyperscaler capex is tracking past $700 billion for the year. Supply constraints are projected to run into 2027 and 2028.
That changes the stakes of an idle GPU. In 2024 a hoarded GPU wasted money. In 2026 it also strands scarce capacity that another team inside your own company is queuing for. GPU cost optimization now depends on workload-aware scheduling (fractional sharing via NVIDIA MPS or MIG), idle detection with automatic shutdown, and Kubernetes-native cost controls for containerized AI. Organizations that implement idle GPU detection typically recover 20 to 35% of total GPU spend, and in a scarcity market they recover capacity, not just dollars.
Measuring Waste: The Five-Week Audit
A structured audit takes five weeks and produces a baseline waste rate you can track over time.
Week 1: Establish utilization baselines. Pull 30-day averages and P95 metrics for CPU, memory, network, and storage IOPS across every account and region. AWS Cost Optimization Hub centralizes this across services in one console. The Azure FinOps Toolkit, an open-source collection of workbooks, optimization engines, and PowerShell modules, does the same for Microsoft estates. GCP’s FinOps Hub aggregates recommender data with billing exports in BigQuery.
Week 2: Map commitment coverage. Calculate your effective savings rate: (on-demand equivalent cost minus actual cost) divided by on-demand equivalent cost. Healthy coverage sits between 60 and 80% for stable workloads. Below 50% means you are overpaying on steady-state infrastructure. Above 85% usually means you have over-committed and are paying for reserved capacity you are not fully using. For GPU families, weigh the capacity-hedge value separately; the discount math alone will understate what the commitment is buying.
Week 3: Hunt orphaned resources. Query for unattached volumes, unused Elastic IPs (AWS charges $0.005/hour per unattached public IPv4 address), stale snapshots older than 90 days, and load balancers with zero healthy targets. This is where the zombie category lives.
Week 4: Analyze data transfer patterns. Map cross-AZ, cross-region, and internet egress flows. AWS Cost and Usage Reports, aligned with the FOCUS specification for cross-cloud normalization, provide the line-item detail to trace transfer costs to specific services and pipelines. FOCUS 1.4, ratified in June 2026, added invoice reconciliation and richer commitment detail; the 1.5 release ratifying on December 3, 2026 adds native token tracking (model ID and family, plus input, output, and cached token fields), which finally brings AI billing into the same dataset as everything else.
Week 5: Calculate your waste rate. Waste Rate = (Identified Waste / Total Cloud Spend) x 100. The Flexera 2026 benchmark is 29% across the industry. Organizations without FinOps programs typically run 32 to 40%. Mature programs operate at 15 to 20%. Set a target 5 to 10 points below your current rate and work toward it over 90 days.
The Tooling Shift: Native Agents Arrive
The biggest change since this guide last covered tooling is not a new dashboard. It is that the detect-to-action gap, the one that made category two mostly a governance problem, is now being attacked by native AI agents.
Native detection matured, then went agentic. AWS Cost Optimization Hub still aggregates and deduplicates recommendations from Compute Optimizer, Trusted Advisor, Cost Explorer, and Budgets into one free console. On top of it, the AWS FinOps Agent entered public preview on June 9, 2026. It investigates cost anomalies down to root cause, answers plain-language cost questions, and posts into Slack and Jira where engineers already work. Azure and Google shipped their own cost agents in the same window. The pitch is that the recommendation no longer waits for someone to open a dashboard; it arrives in the ticket queue. Useful, but treat agent output as a prioritized worklist, not an autopilot. The agents surface the waste they recognize and leave the judgment calls to you.
Where third-party tools still win. Multi-cloud visibility remains the primary differentiator. If you run across two or more providers, no native tool gives you a unified view. FinOps platforms like CloudZero, Finout, and IBM Cloudability provide cross-cloud normalization, usually aligned with FOCUS, plus commitment portfolio optimization and business-unit allocation that native showback cannot match. Several now ingest token-level AI spend directly, which native cost consoles still handle unevenly.
What changed on the vendor map. Apptio Cloudability is now IBM Cloudability, folded into an IBM FinOps Suite that converges Cloudability, Turbonomic, and Kubecost (cross-module integration was still in progress as of mid-2026). CloudHealth left direct Broadcom sales entirely: since May 2024, Arrow Electronics is the sole global distributor, selling it through the ArrowSphere marketplace, and Broadcom shipped a refreshed Tanzu CloudHealth with Intelligent Assist and Smart Summary in June 2025. Spot by NetApp still leads on spot automation but lags on commitment management.
The decision rule: use native tools until your monthly cloud spend passes roughly $100,000 or you operate in more than two clouds. Below that line, the 2026 generation of native tooling (now with agents) drives meaningful savings without a platform subscription. Above it, or across clouds, a third-party platform earns its fee through cross-cloud visibility, commitment portfolio work, and AI cost attribution.
Cutting Waste: Priority Order
Immediate (first two weeks):
- Delete unattached storage volumes: 2 to 4% of storage spend recovered.
- Schedule non-production shutdowns outside business hours with AWS Instance Scheduler or Azure Automation: typically 8 to 15% of total spend within 30 days.
- Release unused Elastic IPs and delete snapshots older than 90 days.
- Implement idle GPU shutdown policies for AI and ML workloads: 20 to 35% of GPU spend recovered, plus reclaimed capacity in a scarce market.
Medium-term (month one through three):
- Right-size the 20 most expensive instances using 14 days of monitoring data. Systematic right-sizing typically yields 20 to 30% compute spend reductions.
- Purchase or convert commitment instruments based on six or more months of historical usage, weighing capacity-hedge value for GPU families.
- Deploy tag enforcement to stop future zombie creation.
- Implement fractional GPU sharing for development and inference workloads using Kubernetes resource quotas.
Strategic (month three through six):
- Migrate fault-tolerant workloads to spot or preemptible instances: 60 to 90% savings off on-demand.
- Rearchitect data transfer patterns to cut unnecessary cross-AZ and cross-region traffic.
- Evaluate serverless for variable-traffic workloads. Lambda costs nothing at zero traffic; EC2 costs the same at 1 request or 10,000.
- Assess neocloud providers for GPU workloads where the hyperscalers charge a scarcity premium.
Making Waste Governance Stick
Detection without governance is a one-time project. The organizations that hold low waste rates across multiple years share three habits.
They assign waste targets to engineering managers, not a central FinOps team. When cost per transaction sits on the same dashboard as uptime and latency, engineers treat efficiency as a first-class engineering concern. The FinOps Foundation’s 2026 survey found practitioners with executive alignment show two to four times more influence over technology selection. That influence starts with putting cost data in front of the teams that create it.
They automate the guardrails. Maximum instance sizes gated behind approval. Auto-termination for resources whose expiration tag has passed. Alerts when utilization drops below threshold for seven straight days. Budget thresholds that trigger investigation, not just a notification. Native agents help here, but only if someone owns the tickets they open.
They track the waste-rate trend, not the absolute dollar figure. A waste rate holding steady at 18% while total spend doubles means new provisioning is repeating old mistakes. The trend line tells you whether governance is working; the absolute number only tells you how much you spend. Flexera’s data shows the industry rate climbing from 27% to 29% in 2026, driven almost entirely by AI workloads that most organizations have not yet brought under FinOps discipline. If your own rate is flat while AI spend climbs, you are quietly winning.
Frequently Asked Questions
What percentage of cloud spend is wasted on average?
The Flexera 2026 State of the Cloud report measured 29% average waste across enterprise portfolios, up from 27% in 2025. That reversal ended five straight years of decline. Organizations without formal FinOps programs waste 32 to 40%; mature practices run 15 to 20%. Zero waste is unrealistic given the need for capacity buffers, burst headroom, and development environments.
How do I calculate cloud waste in my organization?
Sum four inputs: identified idle resource costs, the delta between current and right-sized instance costs, the gap between on-demand and optimal commitment pricing, and orphaned resource costs. Divide by total cloud spend. For AI workloads, add GPU idle time (hours provisioned minus hours of active training or inference, times the hourly rate) as a fifth input. At an industry-average 5% GPU utilization, that fifth input is often the largest.
What is the fastest way to reduce cloud spend?
Scheduling non-production shutdowns outside business hours delivers the quickest return: 8 to 15% of total spend within 30 days, at minimal risk. Next, buy Compute Savings Plans for workloads running steadily for six or more months. Third, delete unattached EBS volumes, aged snapshots, and unused Elastic IPs. For AI-heavy estates, idle GPU shutdown policies usually beat all three.
Do the native cost agents replace third-party FinOps tools?
Not yet. The AWS FinOps Agent and its Azure and Google counterparts close the detect-to-action gap inside a single cloud by pushing root-cause analysis and recommendations into Slack and Jira. They do not give you a unified multi-cloud view, portfolio-level commitment optimization, or mature token-level AI attribution. If you run one cloud under roughly $100,000 a month, native agents may be enough. Across clouds or above that line, a platform still earns its keep.
Why did cloud waste increase in 2026 after years of decline?
AI workloads. Generative AI became one of the most widely used public cloud services in 2026, GPU instances cost 5 to 10 times more than equivalent CPU compute, and most organizations provision them without the utilization monitoring, right-sizing, or commitment strategy they spent a decade building for ordinary compute. Average GPU utilization of 5% (Cast AI, 2026) created a waste category large enough to overwhelm steady gains everywhere else. Scarcity made it worse: teams hoard GPUs they cannot replace quickly, so idle capacity now strands hardware as well as money.
