Cisco and AMD Unveil High-Performance AI Infrastructure Solution
As AI and Machine Learning workloads continue their rapid expansion, organizations face increasing pressure to equip their data centers with robust infrastructure capable of delivering consistent performance and efficiency. Scaling AI/ML workloads across distributed GPU clusters often introduces complexities in coordination and data transfer, leading to underutilized GPUs and extended job completion times. To tackle these critical infrastructure challenges, Cisco and AMD have partnered to deliver a unified, high-performance AI infrastructure solution.
The joint solution is built upon the Cisco Unified Computing System (Cisco UCS) and features a core architecture centered on AMD EPYC CPUs, AMD Instinct GPUs, and AMD Pensando Pollara 400 AI NICs. Complementing these components is Cisco Nexus One for AI Networking, which includes Cisco N9000 Series Switches powered by Cisco Silicon One G200 switching processors. This integrated approach ensures balanced compute, high-memory capacity, and an open, Ethernet-based AI networking layer, enabling customers to scale AI workloads predictably and maximize GPU productivity.
A key element of this offering is the Cisco UCS C885A M8 Rack Server, a dense GPU rack server specifically engineered for large-scale AI training, fine-tuning, and inference. It combines 5th Gen AMD EPYC CPUs, AMD Instinct MI350X GPUs, and AMD Pensando Pollara 400 AI NICs for efficient east/west GPU interconnected fabric. The solution also incorporates UEC-ready RDMA, bringing intelligent load balancing and path-aware congestion control to dynamically manage traffic and maintain consistent performance as clusters grow. Cisco Nexus One for AI networking provides a comprehensive, integrated stack of silicon, switches, optics, and software, all managed through a unified operating model, offering advanced congestion-control mechanisms and robust security.
Read original source