The management issue is less about model fascination and more about when AI demand becomes infrastructure planning. Once AI use shifts from pilots to recurring operational workloads, CIOs need a funding lens that distinguishes exploratory consumption from predictable, business-critical demand. The key decision is not cloud versus owned capacity in principle, but which workload patterns justify moving from variable unit pricing to a capacity-based investment model.
That makes portfolio governance central. Leaders should require each major AI workload to show expected demand shape, utilization assumptions, service-level needs, and business outcome measures before approving a different sourcing model. Without that discipline, organizations risk locking in capital for workloads that never scale, or staying on pay-per-use pricing long after costs become difficult to forecast and control.
The article also points to an operating-model challenge that many AI programs underestimate: owned or reserved capacity only creates value if the enterprise can continuously fill it with productive use. That requires clear decision rights across platform teams, business owners, finance, and architecture leaders; a governed intake process for new use cases; and active benefits tracking tied to adoption, throughput, or customer-service outcomes rather than technical utilization alone.
Practical next questions for IT leaders include:
- Which AI workloads are persistent enough to model as strategic capacity rather than discretionary consumption?
- What utilization threshold would trigger a sourcing review?
- Who is accountable for matching capacity decisions to realized business value six and twelve months later?
The broader implication is that AI cost management is becoming a portfolio and operating-discipline problem, not just a procurement or engineering optimization exercise.
When customers talk about AI costs, the conversation usually starts with token prices and ends with access to the latest, most capable model in the cloud. Do they always need that level of capability? Not necessarily. But that is often where the conversation goes.
As AI moves from experimentation to production, model choice is only part of the equation. When demand becomes steady and business-critical, a consumption-only approach can turn AI spending into a variable monthly line item that is difficult to forecast as usage, workloads, and model requirements change.
At that point, the question is no longer simply which model to consume, or which provider offers the lowest token price: It is how to run AI economically, predictably, and at sustained scale.
AI is moving from isolated pilots into production portfolios: assistants, retrieval-and-knowledge systems, and agentic applications. Customer-service, IT, research, and business-process agents can execute multi-step workflows across enterprise systems, creating recurring demand across models, data, and tools.
This is already starting to happen. Deloitte’s 2026 State of AI in the Enterprise reflects what many leaders are seeing: worker access to AI rose 5% in 2025, and the share of companies with at least 40% of their AI projects in production is expected to double within six months.
When AI becomes a portfolio of always-on workloads, not a collection of experiments, the economics change. Consumption pricing gives teams flexibility and limits commitment. But when usage becomes steady, predictable, and large enough to keep capacity productive, leaders need to ask a different question: Does it still make economic sense to buy AI one request at a time, or is it time to invest in capacity they can optimize and control?
This is not an abstract cloud-versus-on-premises debate. It is a workload-by-workload business decision. Over the next 12 to 18 months, how much AI demand can the company reasonably expect? How consistently will that capacity be used? When multiple workloads share infrastructure, the enterprise can spread fixed costs across more productive use—improving the economics of ownership.
The question is how much you run
Ownership is not automatically the lower-cost answer. It only makes sense when an enterprise can keep capacity productive.
Every organization has a crossover point, the level of sustained use at which owning capacity can become more economical than buying it one request at a time. There is no universal number. It depends on the models being used, the balance of input and output tokens, performance requirements, system design, energy costs, and the operating model required to support it.
A retrieval-heavy knowledge system can have a very different cost profile from a simple assistant because it may process far more context for every interaction. Agentic workflows can be different again: a single business task may involve repeated reasoning, retrieval, model calls, and tool use. That is why generic cost benchmarks are not enough. Enterprises need to model their actual workloads, understand expected demand, and size capacity accordingly.
At the right utilization level, the benefit is not only lower effective cost, but also greater predictability: the ability to manage AI capacity as a strategic infrastructure investment rather than watch a monthly spend line fluctuate with model use and workload demand.
Ownership only works when it is put to work
The capital decision is only half the equation. Even when the economics support ownership, capacity creates value only when the business gets workloads into production quickly and keeps them running.
That takes more than installing infrastructure. It takes an operating model that connects the technology to adoption and business outcomes: bringing users and workloads on board, governing how AI is used, reviewing utilization, and continually identifying the next high-value use case.
The goal is to create value early, then build on it. That means measuring use, identifying underutilized capacity, and bringing additional high-value workloads onto the platform over time. Without that discipline, the business may never realize the economic value that justified the investment. With it, AI capacity becomes a productive asset the business can optimize, expand, and use to create measurable value.
Three questions to ask
Before committing capital, leaders should ask three questions:
- Is demand becoming steady, predictable, and large enough to justify dedicated capacity?
- At what level of usage does ownership make economic sense?
- Can we keep that capacity productive through adoption, governance, and continued use-case expansion?
Make the shift deliberately
As AI moves into production, the organizations that create the most value will look beyond token prices and the latest model. They will know when recurring demand calls for a different economic model—and they will have the operating discipline to make that capacity productive.
That is when AI stops being an expense and becomes an asset.
This content was produced by HPE. It was not written by MIT Technology Review’s editorial staff.
Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.

