Everyone’s Watching New AI Silicon. CoreWeave Found 19.8% in Racks It Already Had.

NVIDIA GB300 NVL72 rack in an AI data center, illustrating CoreWeave MLPerf Inference v6.1 per-GPU throughput gains

TL;DR · 30-second read

The Short Version

CoreWeave rents out computers that run artificial intelligence. On September 16 it published results from an independent industry test and said its machines handled AI requests faster than those of rival rental companies.

The more interesting result is quieter: on machines it installed months ago, it got about 20 percent more work out of each chip purely by improving its own software. No new hardware.

That matters because those chips cost a fortune. Squeezing more out of the ones you own, instead of buying more, is where the profit in this business actually sits.

CoreWeave on September 16, 2026 published its submissions to MLPerf Inference v6.1, the industry benchmark round released the same day by MLCommons. Writing on the company blog, Shadi Saba said CoreWeave entered the Datacenter Closed division’s Available category across four NVIDIA platforms — HGX B200, HGX B300, GB200 NVL72 and GB300 NVL72 — and four model families: Qwen3-VL-235B-A22B, DeepSeek-R1-671B, GPT-OSS-120B and Llama 2 70B. Headline figures include 16,635 tokens per second per GPU on GPT-OSS-120B in the offline scenario, more than 1.19 million tokens per second from a single GB300 NVL72 rack, and 944,902 tokens per second on Llama 2 70B in the server scenario. NVIDIA, in a blog post of its own, announced that its next-generation Vera Rubin NVL72 made its MLPerf debut in the same round.

One day later, on September 17, CoreWeave filed an 8-K disclosing a proposed $3.0 billion offering of convertible senior notes due 2033, with an option for a further $500 million, alongside an Equity Distribution Agreement covering at-the-market share sales and collared forward sales through eleven banks. The two disclosures, a day apart, describe the same business from opposite ends: what the fleet produces, and what it costs to keep building.

Executive Summary

CoreWeave’s MLPerf claims break into two categories, and they are not equally interesting. The absolute throughput records — 1,196 queries per second on the multimodal Qwen3-VL model, over 1.16 million tokens per second on GPT-OSS-120B from one rack — are a function of running NVIDIA’s newest shipping silicon, GB300 NVL72, at scale. Any operator with the same racks and the same power envelope is in that fight.

The second category is harder to replicate. CoreWeave says that comparing its 72-GPU v6.1 submission against its 64-GPU v6.0 submission on GB200 NVL72, derived per-GPU server throughput on DeepSeek-R1-671B rose 19.8% in the five months between rounds. That gain came from serving-stack tuning, scheduling and operations, not from a new chip. If it holds in production, it changes the revenue a depreciating asset can generate across its life — which is the number that underwrites everything else.

That includes the financing. The 8-K filed September 17 shows CoreWeave raising $3.0 billion in converts on top of an existing stack of senior notes carrying coupons from 8.500% to 9.750%. Capital that expensive is serviced by output per installed accelerator, and software gains on hardware already paid for are the cheapest output there is.

The 19.8% That Didn’t Require a New Chip

Strip the benchmark round down and one number does real work. CoreWeave reports that its DeepSeek-R1-671B per-GPU server throughput on NVIDIA GB200 NVL72 rose 19.8% between MLPerf Inference v6.0 and v6.1 — five months apart — through changes to its serving stack, scheduling and operations rather than new hardware. The mechanisms it names are specific: topology-aware scheduling through SUNK, which pins an inference workload inside a single high-bandwidth NVLink domain so that tokens moving between accelerators never leave the fastest interconnect in the rack, and fleet-level health monitoring through Mission Control. For a mixture-of-experts model like DeepSeek-R1 — an architecture that routes each token through a subset of specialised sub-networks, so traffic between chips is constant — placement is not a detail. It is the workload.

Who this affects is everyone financing GPUs on a multi-year schedule. An accelerator bought in early 2026 has a fixed capital cost and a depreciation clock that does not care how efficiently it runs. If per-GPU output rises roughly a fifth without capex, the cost per token served falls by a comparable amount across the remaining life of that asset, and the residual economics of the installed base improve rather than decay. That is a materially different asset story from the one where value arrives only with the next generation of silicon.

The claim deserves its caveats stated plainly, and CoreWeave states some of them itself. Per-GPU throughput is a derived metric — total system throughput divided by accelerator count — and the company notes it is not verified by MLCommons. The comparison also spans two differently sized submissions, 64 accelerators in v6.0 against 72 in v6.1, so scaling behaviour and configuration changes are folded into the same percentage as the software work. The direction is credible and the mechanism is named; the precision of the figure is CoreWeave’s own arithmetic, not the benchmark’s finding.

The Unit of Account Moved From the Chip to the Rack

The GB300 NVL72 is not a server. It is a rack in which 72 accelerators share one NVLink domain, meaning they address each other at memory-like speeds rather than over conventional networking. CoreWeave’s results read accordingly: over 1.16 million tokens per second in the server scenario and over 1.19 million offline on GPT-OSS-120B from a single rack, 944,902 and 1,136,100 tokens per second on Llama 2 70B, and 1,196 queries per second on Qwen3-VL-235B-A22B. The company’s GB200 NVL72 submission reached 901,058 tokens per second on GPT-OSS-120B in the server scenario — the same architecture, one generation back, and a visible step down.

The consequence lands on the people who build and power facilities rather than the people who buy models. When the smallest sensible deployment unit is a rack drawing far more than a legacy hall was designed for, throughput per rack becomes a proxy for how much revenue a given square metre and a given megawatt can carry. Operators comparing a rack-scale NVL72 deployment against clusters of HGX B200 or B300 servers — which CoreWeave also submitted, and which it says led all HGX entries from cloud providers — are really comparing two different power-density and liquid-cooling commitments, not two spec sheets.

Available Is the Word Doing the Most Work

CoreWeave’s submissions sit in the Datacenter Closed division’s Available category. Closed means every entrant runs the same reference implementation, so results are comparable rather than a contest in custom optimisation; Available means the system can actually be obtained. That second qualifier matters this round, because NVIDIA used the same release to announce that its next-generation Vera Rubin NVL72 made its MLPerf debut. A benchmark headline generated by a platform that is debuting is not a price a buyer can procure against today; CoreWeave’s four platforms are all Blackwell and Blackwell Ultra, the generations in production now.

For anyone running a procurement process, the practical read is to check three fields before treating any v6.1 result as a quote: which silicon generation produced it, whether the entry sits in the Available category, and whether the metric quoted is a submitted number or a derived per-GPU figure. CoreWeave also stresses that it ran on the same production clusters and images customers get rather than a benchmark-tuned rig — a genuinely meaningful distinction, though one the benchmark itself does not audit.

Throughput Is the Collateral Story

The day after the benchmark post, CoreWeave’s 8-K disclosed a proposed $3.0 billion offering of convertible senior notes due April 1, 2033, sold under Rule 144A to qualified institutional buyers, with an option for up to $500 million more and capped call transactions to limit dilution on conversion. A concurrent Equity Distribution Agreement adds at-the-market share sales and collared forward sales through a syndicate that includes Deutsche Bank, Goldman Sachs, J.P. Morgan, Morgan Stanley and Citigroup. Interest and conversion terms, the release states, are set at pricing.

The comparison worth making is internal. CoreWeave’s own press release lists the existing obligations the new notes will rank alongside: senior notes at 9.250%, 9.000%, 9.750%, 9.625% and 8.500%, and earlier convertibles at 1.75%. Converts price far below straight debt because the buyer is paid partly in equity optionality rather than cash coupon — cheaper carry now, potential dilution later, which the capped calls are bought to blunt. For a business whose costs are dominated by accelerators, the servicing capacity for any of it comes back to tokens per installed GPU.

That is the honest link between two documents filed a day apart, and it is a structural one rather than a causal one: a benchmark result is a performance claim, not a revenue disclosure, and nothing in the filings ties the offering to the MLPerf round. But the neocloud model runs on both halves at once — demonstrable output per accelerator to justify the capital, and access to capital to buy more accelerators. Software gains of the kind CoreWeave reported on already-deployed GB200 racks are the half that does not require either.

Background

CoreWeave, based in Livingston, New Jersey and listed on Nasdaq as CRWV, is a specialist cloud provider built around NVIDIA accelerators for AI training and inference — a category often called a “neocloud,” distinguished from general-purpose hyperscalers by its narrow focus and its heavy reliance on debt and equity markets to fund hardware. The company was recently named a Visionary in the Gartner Magic Quadrant for Cloud AI Infrastructure.

MLPerf, administered by the MLCommons consortium, has become the closest thing the industry has to an audited comparison of AI serving performance, which is why vendors submit and why the categories matter. Inference — running a trained model to answer requests — is now the larger and more cost-sensitive half of the AI workload, because a model is trained once but served continuously. That shift is what makes tokens per accelerator, rather than raw chip specifications, the metric operators and their lenders increasingly watch.

Sources

Source: MLPerf® Inference v6.1 Results: CoreWeave Leads Providers — CoreWeave’s September 16, 2026 post detailing its Datacenter Closed, Available submissions across four NVIDIA Blackwell platforms. See also NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut, NVIDIA’s account of the same benchmark round.

Primary sources: CoreWeave, Inc. Form 8-K filed September 17, 2026 (Items 7.01 and 8.01, disclosing the proposed convertible notes offering and the Equity Distribution Agreement); Exhibit 99.1 — CoreWeave Announces Proposed $3.0 Billion Convertible Senior Notes Offering; and Exhibit 99.2 — CoreWeave Investor Presentation, September 2026. Full benchmark tables are published by MLCommons at mlcommons.org.