→ Back to Home
Oracle Cloud

Oracle Adds Multi-Node H100 Serving and Day-0 Nemotron Support to OCI Enterprise AI

Oracle announced key upgrades to its OCI Enterprise AI platform in its August 2026 update. Headlining the release is Day-0 availability of NVIDIA's Nemotron 3.5 Lightning, an open model engineered specifically for low-latency, always-on AI agent workflows. In parallel, OCI rolled out multi-node serving support across NVIDIA H100 GPU clusters for imported models, native OCI Identity and Access Management (IAM) authentication on hosted application endpoints, and a new asynchronous Background Mode in the OCI Responses API. The platform also expanded on-demand model availability and added flexible hardware unit shapes for Meta and Cohere foundation models. Deploying cutting-edge open-weight models at enterprise scale typically requires engineering teams to manage complex distributed inference runtimes, handle cross-node interconnects, and stitch together custom identity layers. With native H100 multi-node serving, platform engineers can now host high-parameter models that exceed the memory capacity of a single GPU node directly within OCI's managed perimeter. Furthermore, integrating IAM authentication directly into AI application endpoints eliminates rogue API tokens and aligns model consumption with enterprise role-based access policies, significantly lowering the barrier to deploying enterprise-grade agentic architectures. This release reinforces Oracle's ongoing strategy to differentiate OCI as an open, high-performance host for sovereign and enterprise AI workloads. Rather than restricting users to walled-garden proprietary models, major cloud providers are competing to provide the fastest execution environments for open-source foundation models. Following earlier additions of model import capabilities for families like Cohere and Meta Llama, OCI's Day-0 support for specialized models such as Nemotron 3.5 Lightning and its investment in asynchronous inference APIs mirror the broader industry shift toward persistent, autonomous AI agents running on optimized cluster networking. Practitioners building agentic systems should evaluate whether Nemotron 3.5 Lightning can replace larger general-purpose models for high-frequency subtasks to optimize inference budgets. DevOps and MLOps teams hosting custom large models can now migrate off self-managed Kubernetes clusters on compute instances to native OCI Enterprise AI multi-node endpoints, streamlining maintenance and node health management. Teams should also update their client integrations to leverage the new Responses API Background Mode for long-running workflows, and audit existing endpoint authentication to transition from static keys to native OCI IAM policies.
#oracle cloud#oci#artificial intelligence#nvidia#cloud infrastructure
Read original source