Amazon Web Services announced stricter enforcement of CPU usage limits as engineers face pressure to reduce EC2 waste amid agentic AI demand surge.
Amazon Web Services is cracking down on internal use of EC2 instances among its engineers. In May, the company reportedly met with engineers and told them to reduce CPU waste to ensure AWS has enough CPU capacity to meet customer demand, according to The Information. The directive reflects an unprecedented shift in data center economics: the traditional eight-to-one or four-to-one ratio of GPUs to CPUs is moving closer to parity.
What once required hours is now taking days. Engineers report waiting several days to get access to EC2 instances when they previously could provision them within hours. One engineer told The Information they had never had to wait this long for an instance, even after several years at Amazon.
EC2 instances power a large chunk of the modern internet and are used in both cloud and private deployments. Traditionally, AWS engineers leveraged the relatively low CPU utilization of web infrastructure to spin up multiple virtual machines for development. The internal squeeze signals unprecedented external demand.
Amazon deploys several CPU architectures in EC2, including AMD, Intel, and its newer Graviton5 chip—the company's most powerful CPU to date, which uses an Arm-based architecture similar to Nvidia's Vera CPU and Arm's own AGI.
The CPU shortage has been driven by the explosive growth of agentic AI workloads. Unlike traditional inference, which is GPU-accelerated, agentic AI involves complex orchestration and tool calls that run on CPUs. This has upended the previous model where CPUs served mainly to feed GPUs. The demand is so intense that Intel, AMD, and others report that customers will accept whatever capacity is available.
The scale of the problem is evident in recent incidents. A coding agent at Amazon consumed $1.8 million in token costs last month, surpassing its development budget by 860%. AMD has responded by unveiling its Zen 6 "Venice" CPUs for the data center—the first time AMD has launched a new architecture for data centers before the client market in decades. Nvidia has similarly pivoted its messaging from accelerators toward its new Vera CPU to capture the agentic AI infrastructure opportunity.
Shortages appear concentrated in spot instances. A consultant told The Information that contracted capacity has not experienced meaningful shortages, suggesting Amazon's internal squeeze stems from demand spikes rather than absolute supply constraints.
Following publication, an AWS spokesperson stated: "Demand for AWS services, including EC2, is incredibly strong and growing. Even with this heavy demand, we continue to satisfy the overwhelming majority of compute needs for both our internal and external customers. We work closely with internal teams to meet their compute needs while ensuring they use EC2 resources as efficiently as possible, such as reclaiming idle instances, right-sizing, and scaling up and down as needs change—just as we've always done. These efficiencies help manage capacity for internal and external customers alike, and any suggestions that these long-standing efforts reflect new capacity constraints is simply wrong. This premise is sensationalized. Frugality is in our DNA since Day 1. We have always encouraged our teams to operate efficiently. We also share best practices for how customers can optimize resources to external customers. This isn't a new directive, and encouraging efficient use of resources isn't unique to Amazon."