Amazon Web Services and NVIDIA have announced one of the largest AI infrastructure commitments either company has made public, agreeing to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure between 2027 and 2028. The expanded partnership goes well beyond raw chip counts, reaching into CPUs, networking hardware, robotics, and government computing in a deal that both companies say reflects demand running ahead of even their own forecasts.
The announcement builds on a relationship that stretches back sixteen years, making it one of the longest running technology partnerships in cloud computing. Just months earlier, at NVIDIA’s GTC 2026 conference, AWS had already committed to adding more than 1 million NVIDIA GPUs starting in 2026. According to both companies, that earlier estimate has already been exceeded by actual customer demand, prompting this newer and considerably larger follow-up commitment.
The GPUs themselves span three of NVIDIA’s most advanced chip families. AWS plans to deploy Blackwell Ultra, Rubin, and Rubin Ultra accelerators across its infrastructure, giving the cloud provider access to NVIDIA’s current flagship silicon alongside chips still ramping toward full production. That mix matters because it signals AWS isn’t simply stockpiling one generation of hardware. Instead, the company is positioning itself to absorb whichever NVIDIA architecture proves most efficient for a given workload as the AI compute market keeps shifting year over year.
NVIDIA founder and CEO Jensen Huang framed the expansion as a full stack partnership rather than a hardware order. He described the collaboration as scaling across GPUs, CPUs, networking, open models, and software simultaneously, aiming to make agentic and physical AI achievable at a scale he said only AWS and NVIDIA together could deliver. That language points to where much of this deal’s real substance lives, since the GPU number, however large, is only one piece of what the two companies are building together.
On the CPU side, NVIDIA’s Vera processors are coming to AWS, giving customers a new compute option specifically aimed at agentic AI workloads, the kind of systems that plan, reason, and take multi step actions rather than simply respond to a single prompt. AWS and NVIDIA are also extending NVIDIA NVLink Fusion technology with custom high bandwidth memory, a move designed to reduce bottlenecks between processors and accelerators as workloads scale across increasingly large clusters. Networking gets similar attention, with NVIDIA Spectrum technology being integrated more deeply into AWS’s infrastructure to keep massive GPU clusters communicating efficiently, a challenge that becomes harder rather than easier as deployment size grows into the millions of chips.
Government computing represents one of the more notable additions to this partnership. AWS and NVIDIA plan to build AI factories specifically for the U.S. government, including 100,000 GPUs running on secure AWS infrastructure dedicated to federal and national security workloads. That figure alone would have counted as a major standalone deployment just a couple of years ago, and its inclusion here shows how quickly government demand for dedicated, secure AI compute has scaled alongside commercial demand.
Robotics is another area where the partnership is deepening rather than simply expanding in size. Amazon Robotics is working directly with NVIDIA on next generation robots built around NVIDIA’s physical AI platform, which includes the Jetson hardware platform, the Omniverse simulation libraries, and the Isaac robotics development platform. This is where the concept of physical AI enters the conversation, referring to AI systems that operate in the real world through robots and automated hardware rather than purely digital environments. For Amazon specifically, improvements here could eventually filter into its own warehouse and logistics operations, which already rely heavily on robotics at massive scale.
The companies also detailed new work on data processing and open model access. NVIDIA’s Nemotron family of open models is now available through Amazon Bedrock and Amazon SageMaker, giving AWS customers another option beyond proprietary models when building AI applications. Separately, AWS and NVIDIA are collaborating to bring GPU accelerated data processing to Amazon EMR using new Amazon EC2 G7 instances, while also speeding up vector indexing on Amazon OpenSearch using NVIDIA’s cuDF and cuVS libraries. AWS says its new G7 instances deliver 4.6 times the AI inference performance and roughly double the graphics performance compared to the previous generation G6 instances, and the company noted it is the first major cloud provider offering compute instances built around NVIDIA’s RTX PRO 4500 Blackwell Server Edition GPUs.
Matt Garman, CEO of AWS, tied the announcement back to a simpler idea that has increasingly shaped how enterprises buy AI infrastructure. He said customers want the freedom to choose the best tools for their AI workloads while having confidence that everything works together seamlessly, a statement that gets at why this deal covers so much more than chip volume. Enterprises adopting AI at scale are no longer satisfied with fast GPUs alone. They need CPUs, networking, storage, and software that all work in concert without forcing internal teams to stitch together mismatched components themselves.
The timing of this expansion lines up with a broader industry pattern playing out across major cloud providers this year. Companies including Microsoft, Google, and Meta have all disclosed sharply rising capital expenditure tied to AI infrastructure buildouts, even as some investors have started questioning how quickly that spending will translate into returns. AWS and NVIDIA’s announcement suggests that, at least from the supply side, demand for AI compute capacity shows no sign of leveling off through the next two years, with customers moving from early pilot projects into full production deployments across agentic AI, scientific research, enterprise automation, and robotics.
Whether this level of infrastructure investment proves sustainable will likely depend on how quickly enterprises can turn these massive GPU deployments into workloads that generate real revenue, rather than sitting as underutilized capacity waiting for demand to catch up. For now, both AWS and NVIDIA are betting heavily that the answer is yes, committing to a scale of deployment that would have seemed almost implausible in the cloud computing industry just a few years ago.