→ Back to Home
Oracle Cloud

OCI Scales Superclusters to 131K Blackwell GPUs for Zettascale AI Workloads

Oracle Cloud Infrastructure (OCI) announced orders for massive AI supercomputing clusters powered by NVIDIA's Blackwell architecture, scaling up to 131,072 GPUs. The new OCI Supercluster instances deliver up to 2.4 zettaFLOPS of peak computing capacity, configured with OCI Compute Bare Metal, high-throughput HPC storage, and ultra-low latency networking via RoCEv2 (using ConnectX-7 and ConnectX-8 SuperNICs) or NVIDIA Quantum-2 InfiniBand fabrics. For enterprise infrastructure architects and AI engineering leads, this announcement directly tackles the inter-node communication wall that threatens large-scale distributed training runs. As parameter counts for frontier foundation models and multi-modal architectures expand, standard ethernet-based cloud topologies often introduce unacceptable synchronization latency during all-reduce operations. By engineering a high-density, bare-metal fabric capable of scaling to six figures of interconnected GPUs, OCI addresses critical cluster-level efficiency (MFU—Model Flops Utilization) for massive parallel training jobs. This development highlights a broader industry shift: cloud providers are rearchitecting data centers from general-purpose virtualized compute platforms into specialized, tightly coupled distributed supercomputers. Hyperscalers are competing aggressively to host sovereign AI projects and commercial foundation model builders. Rather than focusing solely on managed high-level APIs, Oracle has differentiated its IaaS layer by offering raw bare-metal hardware and flexible networking choices, allowing teams running custom distributed frameworks (such as Megatron-LM or DeepSpeed) to optimize execution close to the silicon. In practice, infrastructure teams evaluating OCI Superclusters must balance the raw throughput advantages against operational realities. Managing clusters of this magnitude requires advanced checkpointing strategies, automated node health remediation, and rigorous network topology monitoring to mitigate the cost of hardware failures at scale. Organizations should assess whether their training workloads actually saturate smaller multi-node setups before committing to zettascale footprints, while reviewing power density, colocation constraints, and multi-cloud interconnect architectures to prevent data gravity lock-in.
#oracle cloud#oci#supercomputing#nvidia blackwell#ai infrastructure#gpus
Read original source