→ Back to Home
AI Funding

Gimlet Labs Lands $300M Series B at $3B Valuation to Disaggregate AI Inference

Gimlet Labs announced a $300 million Series B funding round at a $3 billion valuation, led by Andreessen Horowitz with participation from Sapphire Ventures, Arm, and Microsoft's venture fund M12, alongside existing investors Menlo Ventures and Factory. The round brings the San Francisco-based startup's total capital raised to $392 million, following commercial adoption across hyperscalers and frontier AI labs. The company builds a multi-silicon inference cloud engineered to disaggregate and execute AI model workloads across diverse hardware platforms—including silicon from Nvidia, AMD, Intel, Arm, Cerebras, and d-Matrix. This raise underscores an inflection point where inference efficiency, rather than raw model training capacity, dictates enterprise AI unit economics. Interactive agents and real-time generation pipelines require consistent throughput and low latency, yet scaling dedicated GPU clusters for around-the-clock inference strains capital and data center power budgets. Gimlet Labs addresses this friction by dynamically breaking down model execution stages across heterogeneous chips best suited for specific compute or memory-bandwidth requirements. For enterprise platform engineering teams, this breaks monolithic hardware dependence, enabling organizations to optimize compute utilization across multiple silicon vendors without re-architecting their underlying inference runtimes. The funding lands amid a broader re-architecting of the AI infrastructure stack. Over the past several quarters, the rapid growth of agentic workflows has created exponential token volumes, shifting compute consumption from batched training jobs to persistent online inference. Simultaneously, major cloud providers and chipmakers have introduced alternative accelerators, yet adoption has historically stalled due to software compilation hurdles and proprietary toolchain locks. Gimlet's multi-silicon orchestration software acts as an abstraction tier analogous to what Kubernetes did for distributed container management, insulating operational workloads from underlying physical hardware heterogeneity. Practitioners managing MLOps and cloud infrastructure must evaluate whether their serving architectures are locked into single-silicon constraints. While multi-silicon orchestration offers attractive utilization gains and lower cost per token, it introduces operational trade-offs around cross-hardware kernel latency, pipeline orchestration, and unified observability. DevOps teams should begin auditing inference pipelines to separate latency-critical token generation phases from throughput-oriented processing steps. Standardizing model packaging and exploring hardware-agnostic runtimes will make it easier to capitalize on multi-accelerator routing fabrics as heterogeneous clouds mature.
#ai infrastructure#inference#multi-silicon#venture capital#mlops
Read original source