AWS and NVIDIA Expand AI Infrastructure with 2M Next-Gen GPUs and Vera CPUs
Amazon Web Services and NVIDIA have announced a major expansion of their strategic infrastructure partnership, committing to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure throughout 2027 and 2028. The planned rollout focuses on next-generation Blackwell Ultra, Rubin, and Rubin Ultra architectures. Crucially, the agreement extends beyond raw accelerator counts: AWS will integrate NVIDIA Vera CPUs designed specifically for AI agent workloads, implement NVLink Fusion with custom high-bandwidth memory (NVHBM), construct dedicated AI factories including a 100,000-GPU environment for secure government workloads, and deepen data processing acceleration across Amazon EMR and OpenSearch with CUDA-X libraries.
For ML platform leads and infrastructure architects, this commitment addresses long-term GPU availability while redefining the architectural baseline of enterprise cloud computing. Agentic workflows require substantial CPU orchestration and memory bandwidth alongside raw matrix acceleration, making the introduction of Vera CPUs and NVHBM directly relevant to real-time agent latency and throughput. Organizations building production AI pipelines are no longer just leasing isolated accelerator instances; they are orchestrating co-engineered systems where compute, networking fabrics like AWS Elastic Fabric Adapter (EFA), and security enclaves like the AWS Nitro System operate as unified supercomputing units.
This development reflects a decisive trend across hyperscalers: the coexistence of proprietary custom silicon and dominant commercial accelerators. While AWS continues heavy investment in its internal Trainium line, the surging demand for enterprise agentic AI and physical robotics mandates sustained, massive access to NVIDIA's software-hardware ecosystem. Infrastructure operators have learned that software compatibility and CUDA-based tooling remain decisive factors in enterprise adoption. Furthermore, the explicit inclusion of federal AI factories underscores how strict regulatory and data sovereignty compliance frameworks are reshaping cloud data center topologies.
In practical terms, DevOps and platform teams must prepare their orchestration layers for deeply heterogeneous clusters. Integrating Vera CPUs alongside Rubin-generation GPUs will require updating scheduler configurations, leveraging dynamic resource allocation (DRA) in container runtimes, and benchmarking workload profiles to determine whether memory-bound agentic tools belong on standard instances, Vera-equipped nodes, or custom Trainium clusters. Teams should also audit data preprocessing stages, as taking advantage of GPU-accelerated cuDF and cuVS in pipelines like Amazon EMR and OpenSearch can yield immediate cost and runtime efficiencies before scaling downstream model compute.
Read original source