OpenAI Details Jalapeño ASIC at Hot Chips 2026 to Optimize Inference Latency and Energy per Token
OpenAI disclosed the microarchitecture and cluster topology of its first custom-designed inference ASIC, codenamed Jalapeño, during the Hot Chips 2026 symposium. Built in collaboration with Broadcom, the 700-watt processor integrates 216 GB of HBM4 memory delivering up to 15.4 TB/s of bandwidth, alongside compute engines rated for up to 13.4 MXFP4 PFLOPS and 3.4 MXFP8 PFLOPS. Developed in a nine-month RTL-to-tapeout window assisted heavily by AI-driven hardware synthesis tools and the XLS framework, the chip scales from a 128-accelerator single-rack domain to a 2,048-chip pod fabric powered by Broadcom Tomahawk 6 switches in a half-flattened two-level Clos topology, delivering 27 aggregate EFLOPS and 32 PB/s of memory bandwidth.
This architecture is significant because inference and multi-turn agentic workloads—not pre-training—now account for the dominant share of AI infrastructure expenditure. Traditional merchant GPUs are often overprovisioned for raw matrix training throughput, incurring steep thermal, electrical, and licensing overheads when serving long-context token generation. By orienting Jalapeño strictly around low-latency decode performance and tokens delivered per joule, OpenAI aims to compress its largest recurring cost center: API and agent inference serving. For enterprise consumers and cloud operators, this proves that purpose-built inference ASICs can structurally alter the cost floor of production LLM serving.
Jalapeño accelerates a broader hyperscaler pivot toward custom captive silicon, mirroring Google's TPU series, AWS's Trainium and Inferentia lines, and Meta's MTIA deployments. As context windows expand beyond hundreds of thousands of tokens and reasoning loops demand iterative tool execution, the conventional approach of stacking massive general-purpose GPUs hits severe memory-bandwidth walls and facility power limits. OpenAI's benchmark comparisons against Nvidia's GB200 on the InferenceX framework highlight that hardware efficiency is no longer governed merely by peak FLOPs, but by NUMA-style spatial memory hierarchies and rack-level interconnect design.
For infrastructure leads and platform architects, Jalapeño demonstrates that inference infrastructure must be optimized around end-to-end request latency (time-to-last-token) and memory bandwidth rather than raw compute density. In practice, organizations should expect custom ASICs to rapidly diverge from merchant GPUs in price-to-performance for specialized, high-volume workloads. While proprietary ASICs remain captive behind closed APIs, their architectural patterns—specifically MXFP4 precision formats, direct Ethernet scale-up fabrics, and HBM4 spatial locality—will define the baseline standards for next-generation inference runtimes and accelerator procurement.
Read original source