AWS and NVIDIA have announced plans to deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure in 2027 and 2028. The August 26 announcement spans Blackwell Ultra, Rubin and Rubin Ultra GPU families, alongside work on Vera CPUs, networking, open models, data processing and robotics. The scale is striking, but the important analytical distinction is timing: this is a future deployment plan, not a count of GPUs available for customers to use today.
The plan reflects an infrastructure market that is becoming more integrated. Training a large model requires more than accelerators. It also requires high-bandwidth memory, interconnects, networking, storage, cluster scheduling, security, power, cooling and software that keeps the machines busy. AWS and NVIDIA are trying to coordinate more of that stack. The stated work includes NVIDIA NVLink Fusion integration with AWS custom silicon, use of NVIDIA’s custom high-bandwidth memory technology, and deeper integration of GPU infrastructure with AWS Nitro and Elastic Fabric Adapter.
For customers, the most tangible implication is not only the eventual number of chips, but the breadth of deployment choices. AWS says it will offer NVIDIA Vera CPU-based infrastructure for agentic AI workloads and continue to support NVIDIA Nemotron open models in Amazon Bedrock and Amazon SageMaker. This positions the cloud as a venue where organizations can choose between managed models, custom training and a mixture of GPU and custom-silicon infrastructure. That optionality can reduce architectural lock-in, but it also makes cost, portability and governance decisions more complex.
The announcement also extends AI infrastructure into data engineering. AWS says its GPU-accelerated data processing on Amazon EMR can deliver up to 3.7 times faster processing and 30% better price-performance than CPU-based configurations. It says GPU-accelerated vector indexing on Amazon OpenSearch can be up to nine times faster at one-quarter the cost. Those are company-reported benchmarks, not universal outcomes. Actual results depend on dataset shape, software configuration, query patterns, instance availability and the engineering skill required to redesign a workload around accelerators.
The inclusion of 100,000 planned GPUs on secure AWS infrastructure for U.S. federal and national-security workloads broadens the significance beyond commercial cloud demand. AI infrastructure is increasingly a procurement and sovereignty issue: public-sector users need controlled environments, specialized compliance postures and dependable long-term capacity. Yet a plan to build secure capacity is not the same as a completed authorization, installation or workload deployment. Readers should separate planned infrastructure, contracted capacity and active production use when measuring the real pace of adoption.
There is also a physical-AI dimension. Amazon Robotics is working with NVIDIA on simulation, synthetic data, training, routing, safety and real-to-simulation validation. That is a reminder that the GPU race is not limited to chat interfaces. Warehouses, industrial processes and scientific workflows can consume substantial compute when they rely on simulations and iterative model training. The commercial hurdle is proving that these systems generate operating value that exceeds the ongoing cost of data, compute, integration and maintenance.
The practical conclusion is that this partnership is an infrastructure roadmap with unusually broad technical scope. Its success should be judged by deployment cadence, availability, utilization and customer economics—not headline GPU totals alone. The next signal to watch is whether customers can access this capacity predictably and whether the promised full-stack integration lowers the cost and complexity of putting agentic and physical AI into dependable production.