AWS and NVIDIA have announced plans to deploy two million additional GPUs across AWS infrastructure during 2027 and 2028. The planned rollout focuses on NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra silicon across global cloud infrastructure and specialized AI factories. The commitment represents scheduled future procurement rather than active data-center capacity, expanding on an earlier target announced at NVIDIA GTC 2026 to introduce more than one million GPUs starting in 2026.

Compute, Memory, and Rack Interconnects

Beyond future GPU deliveries, the agreement introduces several hardware and processor integrations across AWS:

  • Host processors: The companies are working to introduce NVIDIA Vera CPU-based infrastructure to AWS to support agentic AI workloads requiring high-performance host compute paired with accelerators.
  • Custom interconnects and memory: Amazon’s Annapurna Labs is collaborating with NVIDIA and memory suppliers to pair NVIDIA NVLink Fusion with custom high-bandwidth memory (NVHBM) on next-generation Trainium chips, for future Trainium-based infrastructure.
  • Networking and security: NVIDIA Spectrum networking will be used across large GPU clusters, continuing integrations with the AWS Nitro System and Elastic Fabric Adapter (EFA).
  • Near-term graphics and inference: AWS will expand NVIDIA Blackwell capacity, including Amazon EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. AWS claims these instances deliver 4.6 times the inference performance and 2.1 times the graphics throughput of previous-generation G6 instances.

Dedicated Federal Clusters and Robotics

AWS and NVIDIA plan to build AI factories for the U.S. government, including 100,000 GPUs on AWS infrastructure designed for federal and national-security workloads classified at Impact Level 6 (IL6) and above.

For physical AI, Amazon Robotics is incorporating NVIDIA’s robotics stack—including the Jetson platform, Omniverse simulation libraries, and the Isaac robotics development platform—into its warehouse automation workflow. The workflow runs across GPU-accelerated EC2 instances to handle physics simulation, synthetic training data generation, route optimization, functional safety, and validation.

Managed Software and Data Processing

The collaboration also details data-processing and model-serving integrations across AWS services:

  • Managed models: NVIDIA’s Nemotron family of open models remains supported on Amazon Bedrock as serverless managed models and on Amazon SageMaker for self-hosted deployment and fine-tuning.
  • Vector indexing: The companies described GPU-accelerated vector indexing using NVIDIA cuVS for Amazon OpenSearch Service and OpenSearch Serverless. AWS claims up to 9 times faster indexing at a quarter of the cost.
  • Data processing: AWS and NVIDIA are collaborating on GPU-accelerated Amazon EMR processing using EC2 G7 instances and NVIDIA cuDF. AWS claims up to 3.7 times faster processing and 30% better price performance than CPU-based configurations.

The announcement combines managed model access with further integration work and hardware expansion. The two-million GPU rollout is scheduled for 2027 and 2028.