brain/
sourcestock-marketartificial-intelligence

2026 09 15 Vera Rubin Nvl72 Agentic Inference 67X Better

Early AgentX results on Vera Rubin NVL72: ~67x tokens per TCO vs GB300 at 170 TPS on TRTLLM NVFP4; ~2x modeled profit per GW vs Blackwell. Jensen GTC 2026 3x/MW vs measured up to 7x.

view source ↗
Source

Summary

SemiAnalysis (September 14, paid; free portion filed) publishes the first verified AgentX agentic-inference results for Vera Rubin NVL72. Even on early pre-release software they report ~67x total throughput per TCO vs GB300 Dynamo TRTLLM at 170 TPS, and ~2x modeled annual profit per gigawatt vs the strongest Blackwell configuration. Jensen's GTC 2026 "3x performance per MW" is called a sandbag against measured up to 7x. This is hardware-mix and profit-per-GW color for existing AI-infra spines — not a tenth AI-infrastructure chain.

Article

Rubin is the first platform co-designed across six products for the agentic era: Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and Spectrum-6. Today we are publishing the first verified agentic inference results for Rubin, measured on our agentic inference benchmark, AgentX. Even on early pre-release software, the results already show why extreme co-design was necessary.

At GTC 2026, Jensen presented this graph that VR NVL72 achieved 3x performance per MW compared to Blackwell on O(1-3 Trillion) parameter model around 200 TPS. But when compared to the real world performance of Rubin already on prelease software, we are already seeing up to 7x better token throughput per megawatt. Jensen needs to stop sandbagging his performance claims at GTC. The last time he did this at GTC 2024, when he claimed GB200 NVL72 would deliver 30x Hoppers performance, but when we tested, it was 98x better performance than Hopper.

Our estimates on the performance results show that even on early software builds, Rubin can earn over 2x more profit per gigawatt than the Blackwell platform. As the Rubin software stack and kernel libraries mature, and as the developer community builds experience optimizing for Rubin, we expect that gap to widen further.

We evaluate the performance using the industry standard agentic inference benchmark scenario called AgentX. This replays real world agentic traffic across our fleet of thousands of chips. The results can thus be holistically referenced by actual inference providers and hyperscale AI labs to decide which chips are most efficient in what scenarios.

Vera Rubin is a substantial improvement over GB300 in terms of performance per dollar. At 170 TPS, On Apples to Apples TRTLLM NVFP4 Dense, Vera Rubin NVL72 delivers ~67x the total throughput per TCO of GB300 Dynamo TRTLLM under the owning cost assumptions. Furthermore, at the part of the frontier where most providers would actually serve this model (60-100 TPS), Vera Rubin achieves between 1.4x and 3x the throughput per TCO compared to the latest and greatest GB300 TRTLLM configuration.

Our rack uses the production SKU of 2300W TDP & 1.5TB of CPU LPDDR5X per compute tray.

Vera Rubin achieves approximately 61% higher maximum P90 interactivity than GB300 Dynamo TRTLLM, reaching 276.24 versus 171.53 P90 TPS. However, when using the open-source SGLang stack, GB300 can achieve similar interactivity as Vera Rubin.

When we consider the 3 year rental cost which in July 2026 was over $8.5/hr/chip for Rubin and $5/hr/chip for Blackwell Ultra NVL72, the upgrade is still more than justified. At 80 TPS P90 interactivity, Vera Rubin can produce 62% more total tokens for the same rental TCO. Across the higher interactivity portions of the frontier, Vera Rubin realizes up to 16x more tokens per 3 year rental TCO.

On agentic workloads, Vera Rubin makes H200 look about as competitive as a TI-84 calculator. At a P90 interactivity target of 80 TPS, Rubin delivers 18x as many tokens per dollar. At 120 P90 TPS, that advantage widens to 39x.

At 100 TPS, Rubin delivers approximately 59.4 million total tok/s/MW, compared with 28.5 million for GB300 Dynamo SGLang and 21.1 million for GB300 Dynamo TRTLLM. That is a 2.09x advantage over the stronger GB300 engine at this target.

At a fixed power budget, Rubin’s advantage is not simply that it can process more tokens. It gives an operator more capacity to monetize demand without securing additional utility power. At 75 TPS interactivity, 60% utilization and no model-license fee, Vera Rubin generates $159.5 billion in annual revenue and $149.9 billion in modeled profit per all-in utility GW. The strongest GB300 configuration in this annual revenue comparison, Dynamo SGLang, generates $114.9 billion and $105.3 billion respectively. Rubin therefore delivers approximately 39% more revenue and 42% more modeled profit from the same power allocation.

The same advantage also creates pricing headroom. Holding workload mix, throughput and billable utilization constant, Rubin could charge approximately 28% less across cached-input, uncached-input and output tokens while matching the revenue per GW of GB300 Dynamo SGLang at the displayed prices.

(Paywalled after this point. partial: true.)

Referenced by