Vista’s 50-Metro Inference Cloud Bets Enterprise AI Won’t Run on GPUs Alone

Vista Equity and Cambian Capital VC2 inference cloud server racks mixing Intel, Nvidia and SambaNova chips in Los Angeles

TL;DR · 30-second read

The Short Version

Two investment firms, Vista Equity Partners and Cambian Capital, have opened a computing center in Los Angeles built for running AI tools day to day. That is the moment a chatbot answers or software sorts through data, as opposed to building new AI.

They plan similar sites in more than 50 American cities. The twist: instead of relying only on Nvidia’s well-known graphics chips, they mix in chips from Intel and a smaller company, SambaNova. Their bet is that matching each job to the right chip makes AI cheaper to run.

Vista Equity Partners and Cambian Capital have launched an enterprise inference cloud built for what they call disaggregated inference, with its first facility in Los Angeles, Fierce Network reported on June 3, 2026. The platform, called Vector Core Compute (VC2), combines three kinds of processors: Intel Xeon CPUs, Nvidia Blackwell GPUs and SambaNova RDUs.

Additional sites are in development in Chicago, Seattle and Phoenix, and the companies say sites are planned across more than 50 U.S. metros. They describe the distributed model as designed to scale into strategic markets globally, with future regions shaped by strategic partners and customer demand.

Executive Summary

VC2 is a cloud for inference, the work of running already-trained AI models to answer queries, process data and carry out tasks. It is not a cloud for training new models. Its defining choice is architectural. Rather than standardizing on a single accelerator, it spreads work across general-purpose CPUs, Nvidia GPUs and SambaNova’s dataflow chips. Vista CEO Robert F. Smith framed the logic as cost: distributing workloads like “always-on monitoring, high-volume data processing, and complex multi-step orchestration across specialized hardware to reduce cost.”

Two further features make this more than another GPU cloud launch. The first is the footprint: many smaller sites in metro areas rather than one giant campus, with a stated plan of more than 50 U.S. metros. The second is the demand source. Vista, one of the largest investors in enterprise software, is positioning the platform as early-access infrastructure for its own portfolio companies. If the cost argument holds, it strengthens the case that enterprise inference will be served by mixed fleets placed close to users, not only by GPU clusters in remote mega-campuses.

Three Chip Families, One Cost Argument

The headline claim of VC2 is that enterprise AI workloads do not all belong on a GPU. Disaggregated inference means breaking an AI job into its component steps and sending each step to the hardware best suited to it. The workloads Smith named show why that could matter. Always-on monitoring, high-volume data processing and multi-step orchestration are not single large model calls. They are long chains of control logic, data movement and retrieval, punctuated by bursts of heavy model math. An AI agent that checks a system, pulls records, decides what to do and then drafts a response spends much of its time on steps that do not need a top-tier accelerator.

The hardware mix maps onto that pattern. Intel Xeon CPUs are general-purpose server processors, well suited to branching logic and data handling. Nvidia Blackwell GPUs, Nvidia’s current data center generation, excel at the massively parallel arithmetic at the heart of large models. SambaNova’s Reconfigurable Dataflow Units (RDUs) are a different accelerator design, which SambaNova positions as efficient at serving models. The implied economics are simple: stop paying GPU rates for work a cheaper processor can do, and reserve the most expensive silicon for the steps that need it.

What the launch does not yet show is proof. No pricing, benchmarks or utilization figures accompanied it, so the claim that a mixed fleet undercuts a GPU-only deployment remains an assertion. Heterogeneity also carries its own costs: three vendor software stacks, three supply relationships, and a scheduling layer that must decide in real time where each step runs. In practice, VC2’s product is that orchestration layer. The chips are commodities that others can buy. Whether the savings survive the complexity is the thing to watch.

Mixed Racks, Many Cities: The Facility Brief Changes

For the people who build and operate data centers, a three-architecture inference cloud is a different tenant from a GPU cluster. CPU servers typically run in conventional air-cooled racks. Nvidia’s densest Blackwell rack systems are designed for liquid cooling and draw far more power per rack. A site hosting both, plus a third accelerator type, needs rooms that can handle uneven power density and mixed cooling side by side. The companies have not said which Blackwell configurations they deploy, but the design choice pushes toward facilities with flexible power distribution rather than uniform halls.

The metro strategy multiplies that brief. One site is live, in Los Angeles. Three more are in development, in Chicago, Seattle and Phoenix. The 50-plus figure is a plan, not a footprint. Executing it would mean dozens of separate agreements for space, power and network connectivity, many of them in cities where grid capacity for new data center load is already contested. Colocation providers (companies that rent out space, power and cooling inside shared data centers) with available power in major and second-tier metros are the natural beneficiaries if the rollout proceeds. Each additional city is also an additional execution risk.

A Portfolio as the First Customer

Filling capacity is among the hardest problems for any new AI cloud. Hardware depreciates whether or not anyone rents it. Vista’s framing addresses that directly: early access for portfolio companies that, in Smith’s words, are “doing this work today.” A private equity owner with a large stable of enterprise software businesses can, in effect, supply a starting pool of demand. It also has a direct interest in lowering the AI running costs that now sit inside those companies’ cost of goods.

That structure is a real advantage, but it also raises fair questions. Captive demand from sister companies does not by itself demonstrate a market among outside enterprises. How pricing and commitments work between Vista-owned software companies and a Vista-backed cloud has not been described. The model is best read as a test of whether an owner of software businesses can bring down those businesses’ AI costs by building the infrastructure itself rather than renting it from hyperscalers, the largest public cloud providers.

Background

Vista Equity Partners is a private equity firm that specializes in enterprise software, with a portfolio of software companies that increasingly embed AI features into their products. Those features must run somewhere, and the cost of running them, known as inference, has become a significant line item for software businesses. Cambian Capital is Vista’s partner in launching the VC2 platform.

The AI infrastructure market has so far been dominated by large GPU clusters, many built in remote campuses sized for training ever-larger models. As AI moves into everyday business software, attention has shifted toward inference. Inference runs continuously, serves live users, and rewards low cost per task and proximity to customers. That shift has opened space for specialist clouds and for alternative chip designs, such as SambaNova’s, that compete with GPUs on serving efficiency.

Sources

Source: Vista Equity Partners, Cambian Capital light up new inference cloud in Los Angeles, Fierce Network’s report on the launch of the VC2 inference cloud and its first Los Angeles facility.