Data Center Knowledge reported on May 23, 2026, that AI inference — the day-to-day serving of trained AI models to end users — is pulling infrastructure investment back toward metro data centers, reversing years of momentum toward remote hyperscale campuses. The driver, per the report’s framing, is latency: inference workloads live and die by response time, and response time is a function of physical distance to users.
Executive Summary
The trade publication’s thesis is straightforward: the AI buildout’s first act was dominated by training — the compute-intensive process of creating models — which rewarded remote sites with cheap land and abundant power, because training does not care where it runs. The second act is inference, the phase where those models actually answer queries for businesses and consumers, and inference is latency-sensitive in a way training never was.
If the thesis holds, it matters for nearly everyone in the infrastructure value chain. Metro colocation operators, carrier hotels, and interconnection-rich urban facilities — assets many analysts treated as yesterday’s story during the gigawatt-campus land rush — would regain strategic relevance. Site-selection criteria, capital allocation, and power procurement strategies would all tilt back toward proximity to population centers, precisely where power and real estate are scarcest.
Training Built the Campuses; Inference Pays the Bills
Training and inference are economically different animals. Training is a batch job: it runs for weeks or months, consumes enormous power, and produces a model. Because no end user is waiting on it in real time, operators could chase the cheapest available megawatt — which pushed campuses into rural and exurban regions with land, transmission access, and accommodating utilities. Inference is the opposite: it is the recurring, revenue-generating workload, triggered every time a user prompts a chatbot, a copilot drafts an email, or an application calls a model behind the scenes.
As AI products mature from demos into production services, the share of total AI compute devoted to inference grows structurally. That shifts the industry’s center of gravity from “where is power cheapest?” to “where are the users?” — a question metro data centers were built to answer. The report’s framing suggests the market is beginning to price this in.
Why Latency Is Redrawing the Map
Latency — the delay between a request and its response — is bounded by physics. Data cannot travel faster than light through fiber, and every additional kilometer between user and server adds round-trip time. For a monthly batch job, that is irrelevant. For an interactive AI assistant, a fraud-check API, or a voice agent, tens of milliseconds are perceptible and, at scale, commercially meaningful.
Newer AI application patterns compound the effect. Agentic and multi-step systems chain many model calls together to complete a single task, so per-call latency multiplies. Retrieval-augmented applications shuttle data between models and enterprise systems that already live in metro colocation facilities. Placing inference capacity near users and near enterprise data reduces both delay and data-transit cost — a pull toward the very urban markets the hyperscale era had de-emphasized.
Winners, Losers, and the Assets in Between
The clearest beneficiaries of a metro revival would be operators holding interconnection-dense urban facilities: carrier hotels, established colocation campuses in major metros, and providers with existing utility relationships in constrained markets. Those assets are hard to replicate — urban land, fiber density, and grid connections accumulate over decades. Enterprises also stand to gain optionality, since inference capacity near their existing colocation footprints simplifies hybrid architectures.
This is not, however, a zero-sum reversal. Remote hyperscale campuses remain essential for training and for latency-tolerant inference, and the report’s headline says infrastructure is being pulled “back into” metros, not out of the hinterlands. The more defensible reading is bifurcation: a two-tier geography where massive remote campuses handle training and batch work while a distributed metro layer serves real-time inference. The open question is how capital gets split between the tiers — and whether metro grids can absorb their share.
The Constraint That Follows the Workload: Power
The uncomfortable irony is that inference demand is heading toward the places least prepared to power it. Major metros already contend with constrained grids, long interconnection queues, and community resistance to new data center construction. AI inference hardware, while less power-dense per site than a training cluster, still pushes rack densities well beyond what many legacy urban facilities were engineered for, often requiring liquid cooling retrofits and electrical upgrades.
That constraint cuts both ways. It limits how fast the metro shift can happen, but it also makes existing permitted, powered metro capacity more valuable — scarcity is a landlord’s friend. Expect the competition for metro megawatts, substation capacity, and retrofittable urban shells to intensify if the trend the report describes continues.
Background
Data center geography has swung on a pendulum for two decades. The early internet clustered compute in urban carrier hotels where networks met; the cloud era then pushed capacity outward to remote regions where land and power were cheap, and the AI training boom of the mid-2020s accelerated that outward push into multi-hundred-megawatt and gigawatt-scale campuses.
Data Center Knowledge, the source of this report, is a long-running trade publication covering the data center industry. Its May 2026 piece captures a question the industry has been circling as AI products move from development into production: once models are built, the economics of serving them — inference — may favor a very different map than the one training drew.
Source: AI Inference Pulls Infrastructure Back Into Metro Data Centers — Data Center Knowledge, May 23, 2026, on how latency-sensitive AI inference workloads are shifting data center demand back toward metropolitan markets.

