Gartner’s $23.3B Inference Forecast Shows AI Cloud Demand Turning Always-On

Rows of GPU servers in a cloud data center illustrating Gartner's AI-optimized IaaS spending forecast and the shift to inference

TL;DR · 30-second read

The Short Version

Businesses are on track to nearly double what they spend renting powerful computers for artificial intelligence this year, to about $42 billion, according to research firm Gartner.

The bigger change is what that money pays for. For the first time, more will go to running artificial intelligence tools every day — answering customers, handling tasks at work — than to building new ones.

Building happens in bursts. Running never stops. So the giant computer warehouses behind these tools increasingly have to stay busy, and powered, around the clock.

Gartner said in a forecast published August 10, 2026 that worldwide spending on AI-optimized infrastructure as a service (IaaS) — rented cloud computing built for artificial intelligence workloads — will grow 96.4% in 2026 to $42.3 billion, up from $21.5 billion in 2025, and reach $66.1 billion in 2027, a further 56.5% increase.

The research firm also forecast that in 2026 spending on inference (running trained AI models in production) will reach $23.3 billion and surpass spending on training (building those models) at $19 billion. Inference is projected to account for 55% of AI-optimized IaaS spending in 2026 and 59% in 2027.

Executive Summary

The headline number is the near-doubling of rented AI cloud spending, but the more consequential figure is the crossover: in Gartner’s forecast, 2026 is the year inference overtakes training as the larger use of AI-optimized cloud capacity. Hardeep Singh, Sr Principal Research Analyst at Gartner, attributes the shift to fine-tuned and domain-specific models being built into customer-facing and operational systems that require “continuous, real-time execution rather than periodic training.”

That distinction matters for anyone who builds, powers or connects AI infrastructure. Training demand arrives as large, time-bound projects; inference demand scales with everyday usage and persists. A market increasingly weighted toward inference behaves less like a series of construction jobs and more like a utility load — steady, growing with adoption, and tied to where users and applications are.

Gartner’s numbers also put the AI boom in proportion: AI-optimized services are forecast at 14.7% of total IaaS spending in 2026, and conventional cloud infrastructure keeps growing at roughly 20% a year alongside it. Data center operators face rising demand on both fronts, not a wholesale swap of one for the other.

Inference Overtakes Training, and Demand Stops Arriving in Bursts

Training is the process of building an AI model by feeding it vast amounts of data; inference is the work of using that finished model to answer a question, draft a document or complete a task. Gartner forecasts inference spending of $23.3 billion in 2026 against $19 billion for training, with inference’s share of AI-optimized IaaS rising from 55% in 2026 to 59% in 2027. Singh ties the shift to models moving into production systems that require “continuous, real-time execution rather than periodic training.”

The mechanism is the shape of the demand. A training run is a project: a large block of accelerators — the specialized chips, usually GPUs, that do AI math — is reserved, runs for weeks, and is then released. Inference consumption instead scales with use: every customer query and every step an AI agent takes draws compute. Gartner specifically flags agentic AI, software that carries out multistep tasks on its own, as amplifying compute intensity. Demand that tracks usage behaves like a utility load — it persists, rises with adoption and follows the rhythm of the businesses it serves.

That changes the planning problem for several groups. Cloud providers get fleets that are more consistently busy, which improves the economics of expensive accelerators. Data center operators and utilities see more load that draws close to its contracted power much of the time, rather than in project-driven peaks. And because inference feeding customer-facing systems has to reach those customers quickly, well-connected sites near users and networks tend to gain value relative to remote campuses chosen purely for cheap power.

The Percentages Are Slowing. The Dollars Are Not.

On growth rates alone, AI-optimized IaaS looks like it is cooling: 180% in 2025, 96.4% in 2026, 56.5% in 2027. But the absolute increments rise each year — about $20.7 billion of new annual spending in 2026 and roughly $23.9 billion in 2027. For people who build capacity, the increment is what counts: under Gartner’s forecast, each year requires more new AI-ready capacity than the year before, not less.

The forecast also keeps AI in proportion. AI-optimized services are about 9.7% of total IaaS spending in 2025, 14.7% in 2026 and 18.4% in 2027. They account for roughly a third of the market’s growth in each forecast year. The remainder — general-purpose rented infrastructure — still grows by about 22% in 2026 and about 20% in 2027, on Gartner’s figures. AI is not crowding conventional cloud out of these numbers; builders face demand for both high-density AI halls and standard cloud capacity, which have very different power and cooling requirements.

Rented AI Capacity Is Not the Same as Built Capacity

Gartner’s figures measure what customers pay to rent AI-optimized infrastructure, not what cloud providers spend building it or what enterprises spend on hardware for their own facilities. The forecast therefore shows strong and rising demand for rented AI capacity, but on its own it cannot show AI investment shifting away from ownership toward the cloud, because it does not measure the ownership side.

What it does indicate is where enterprise AI is being consumed as it moves into production. Singh says the shift toward deployment is “accelerating the cloud consumption patterns.” Rental spending is the revenue that providers weigh their construction against, and a forecast in which that revenue nearly doubles and tilts toward always-on inference strengthens the case for facilities designed for sustained, high-density load — heavier power delivery and the advanced cooling dense AI racks often require. The caveat is that this is a forecast: whether chips, power and finished data center space arrive fast enough to meet it is a separate question.

Background

Gartner, Inc. (NYSE: IT) is a Stamford, Connecticut-based research and advisory firm whose market spending forecasts are widely used by technology buyers, vendors and investors. Infrastructure as a service is the part of cloud computing in which customers rent raw computing, storage and networking from providers such as Amazon Web Services, Microsoft Azure and Google Cloud, paying for what they use instead of running their own hardware.

The generative AI boom created a distinct segment within that market: cloud capacity built around accelerator chips and the high-speed networking needed to train and run large language models. Gartner’s figures show that segment grew 180% in 2025 to about $21.5 billion, and the firm now expects spending to shift from building models toward running them in everyday business applications.

Sources

Source: Gartner Forecasts Worldwide AI-Optimized IaaS Spending to Grow 96% in 2026 – Gartner — Gartner’s August 10, 2026 forecast of AI-optimized cloud infrastructure spending, including its inference-versus-training split.