Rising Token Costs Drive Enterprise AI Workloads to Open-Weight Models, Shifting MLOps Focus to FinOps
A recent report by McKinsey, presented at Tech Week Singapore, indicates a significant trend in enterprise AI: most workloads are migrating towards open-weight models rather than relying on proprietary frontier models from providers like OpenAI and Anthropic. This shift is primarily driven by the rapidly increasing token consumption of agentic AI systems, which is pushing AI spending past allocated budgets for a staggering 93% of enterprises in the past six months.
This development is critical for MLOps practitioners as it fundamentally alters the calculus for model selection and deployment. The focus is no longer solely on achieving the highest possible model performance, but also on the economic viability and sustainability of AI solutions in production. While the cost per token has seen a substantial decrease—around 90% since 2023—the exponential increase in token usage by agentic models, estimated by Gartner to be five to 30 times more per task, negates these per-token savings. This means that even with cheaper individual tokens, the overall AI bill is skyrocketing, forcing organizations to seek more cost-effective alternatives.
This trend fits into the broader, well-established movement within cloud and DevOps towards FinOps, or Cloud Financial Management. Just as cloud infrastructure costs required careful monitoring and optimization, the operational costs of AI are now demanding similar scrutiny. The article suggests that core, routine traffic where unit cost is paramount, and domain-specific tasks requiring fine-tuning on proprietary data, are ideal candidates for self-hosted open-weight models. Conversely, spiky, low-sensitivity workloads without strict data residency requirements could leverage open-weight models via hosted APIs. Only the most complex reasoning tasks, where the cost of failure is exceptionally high, would remain with frontier APIs.
In practice, this means MLOps teams must develop stronger competencies in cost analysis and optimization. They need to implement robust monitoring for token usage, evaluate the trade-offs between proprietary and open-weight models based on specific workload characteristics and budget constraints, and potentially invest in the engineering skills required to build and manage sophisticated "harnesses" around LLMs. A harness is the software layer that controls the context, token expenditure, and task execution of an LLM or agent. This shift implies a greater emphasis on architectural design for cost-efficiency, potentially leading to hybrid deployment strategies that intelligently route different types of AI tasks to the most economically appropriate model and infrastructure. Practitioners should actively explore open-weight model ecosystems, understand their self-hosting implications, and integrate FinOps principles deeply into their MLOps workflows to ensure the long-term financial sustainability of their AI initiatives.
Read original source