→ Back to Home
Observability

Cisco Expands AI Infrastructure Stack with Hardware-to-Job Observability

Cisco expanded its Secure AI Factory with NVIDIA through a partnership with Supermicro, introducing high-density liquid- and air-cooled GPU rack architectures scheduled for availability in October 2026. Beyond expanding physical compute capabilities to support trillion-parameter model training and high-throughput inference, the release embeds integrated day-two telemetry and management via AgenticOps within Cisco Cloud Control and NVIDIA AI Enterprise. This operational framework correlates AI training and inference job health directly with granular hardware telemetry across compute nodes, network interface cards (NICs), optical interconnects, and switch fabrics. For DevOps, platform engineering, and observability teams, managing distributed AI workloads introduces failure modes fundamentally distinct from traditional microservices. Distributed GPU clusters are exceptionally vulnerable to gray failures—such as intermittent packet drops, degraded optical transceivers, or thermal throttling—where a single impaired link can stall an entire collective communication phase and halt distributed training runs. By coupling network fabric telemetry across Cisco Silicon One and NVIDIA Spectrum-X switches directly with host-level GPU performance and job health metrics, engineers can swiftly determine whether a performance slowdown stems from model logic, framework overhead, or underlying physical network degradation. This development reflects the broader industry convergence of physical infrastructure monitoring, network telemetry, and AI application observability. Historically, data center operators monitored switches and servers in isolation, while machine learning practitioners relied on separate profiling tools to track job throughput and loss curves. However, as modern AI architectures push rack densities beyond 200 kW and mandate rack-to-fabric liquid cooling, operational health can no longer be abstracted away from physical and environmental variables. Unifying telemetry across silicon, cooling systems, and high-bandwidth fabrics represents a critical maturation in observability, shifting the discipline from discrete component monitoring toward full-stack physical-to-job correlation. In practice, engineering teams running large-scale AI infrastructure should audit their observability pipelines to eliminate blind spots between hardware layers and job schedulers. Practitioners must ensure that high-frequency telemetry from NICs, optics, and GPU engines can be ingested, normalized, and correlated in real time without creating telemetry bottlenecks. Teams should focus on configuring automated alarms that link distributed job slowdowns to specific fabric anomalies, allowing site reliability engineers to proactively cordon degrading nodes or optical paths before unrecoverable training checkpoint failures occur.
#observability#ai infrastructure#network monitoring#telemetry#agenticops
Read original source