→ Back to Home
AI Hardware

AWS and NVIDIA Expand Alliance to Deploy 2 Million Rubin and Blackwell Ultra GPUs

Amazon Web Services and NVIDIA have announced a major multi-year expansion of their infrastructure collaboration, committing to deploy an additional 2 million NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs across AWS Global Infrastructure across 2027 and 2028. The agreement introduces NVIDIA Vera CPU-based instances to the cloud provider, deepens deployment of Amazon EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, and extends NVIDIA NVLink Fusion interconnect and custom NVIDIA high-bandwidth memory (NVHBM) technology into Amazon's Annapurna Labs Trainium roadmap. For DevOps leaders, infrastructure architects, and AI practitioners, this scale-out roadmap addresses persistent compute allocation constraints while fundamentally altering cluster engineering. By bringing NVLink Fusion and NVHBM into AWS's custom Trainium ecosystem, the collaboration reduces the architectural friction historically encountered when mixing merchant GPU nodes with hyperscaler-specific ASICs. Workload orchestrators will gain unified memory semantics and higher cross-chip bandwidth, which are vital for mitigating latency during complex agentic workflows, large-scale model fine-tuning, and multi-agent reasoning tasks. This development fits into a broader macro trend across cloud computing where memory wall limits and power constraints dictate data center evolution. As models evolve from single-prompt generation to multi-step agentic coding and complex tool execution, input context windows expand rapidly, shifting hardware pressure toward memory bandwidth and intra-node communication speeds. Rather than treating proprietary silicon and commercial accelerators as mutually exclusive silos, cloud hyperscalers are increasingly adopting hybrid hardware topologies that share advanced packaging, specialized memory, and high-speed fabrics. In practice, engineering teams should begin planning workload scheduling strategies that capitalize on this hybrid accelerator model. Platform teams managing Kubernetes clusters with custom device plugins must prepare for topologies where frontier reasoning runs on Rubin architectures while high-volume batch processing routes to NVLink-connected Trainium nodes. Additionally, organizations migrating inference workloads to Amazon EC2 G7 instances can leverage a reported 4.6x improvement in AI inference efficiency over prior G6 baselines, lowering operational token costs and allowing more complex agent loops within production budgets.
#nvidia#aws#gpus#ai accelerators#trainium
Read original source