brain/
Semis · Software moat

Nvidia’s inference moat

CUDA built the training landscape. Feldman says it has no role in inference. A late-August agentic-coding bench still has Nvidia winning. The same shop’s token-per-watt bench, on a different test, does not.

Covers stock-market wiki · pages updated through September 2026

Andrew Feldman, who sells a chip that does not run CUDA, put the claim in a sentence. CUDA was decisive in creating the AI landscape. It is not important now, he said, and it has no role whatsoever in inference. Moving a model off Nvidia’s GPUs onto Cerebras is “10 keystrokes. Just move point to our API.”

CUDA was really important in the creating of the AI landscape, but it's not important now and it has no role whatsoever in inference… we can move it in 10 keystrokes.

Andrew Feldman, Cerebras

A year ago, in his telling, every major frontier model was built on a CUDA foundation. Today two of three are not. Gemini trains on Google’s TPUs. Anthropic trains on Amazon’s Trainium, “no CUDA.” Only GPT remains CUDA-trained. Feldman calls that a hemorrhaging of share — “they lost 70% market share.” Count the labs, not the wafers. The number is his.

A mid-September check left the Gemini half standing and weakened the Anthropic exclusive. Google’s own Gemini 3.1 Flash-Lite model card says the model “was trained using Google’s Tensor Processing Units (TPUs).” SemiAnalysis, in November 2025: “Gemini 3 is one of the best models in the world and was trained entirely on TPUs.” JAX and Pathways, not CUDA. Anthropic, on 20 April 2026, said it “currently use[s] over one million Trainium2 chips to train and serve Claude,” with AWS still “primary training and cloud provider.” The same firm’s own posts also name a three-platform mix: “Google’s TPUs, Amazon’s Trainium, and NVIDIA’s GPUs.” Feldman’s “no CUDA” sentence does not survive that primary. Record it as a contradiction. Do not silently merge it. OpenAI’s latest named pretrain is still Nvidia: Aidan Clark on GPT-6 Astra said it was “the first time we’ve pretrained on more than 100,000 GPUs at our Stargate site in Texas.” Jalapeño remains an inference accelerator, not a training replacement. “Seventy percent / two-of-three” is still competitor framing, not an audited FLOP-hour share.

J.P. Morgan Equity Research, via Michael Cembalest, gives the economics a non-competitor source: hyperscalers using in-house chips report total-cost-of-ownership reductions of 30 to 40 percent. Anjney Midha, from the buyer’s side of the table, puts about 80 cents of every research dollar a lab spends in the pocket of a chip provider like Nvidia. That is why Microsoft builds Maia, and why every lab wants a second stack.

None of this touches the other moat. Nvidia still takes more than half of TSMC’s CoWoS — on the order of 800,000 to 850,000 wafers a year. Stacks and packaging are a physical queue. CUDA is a software argument. They can both be true at once.

The multiple already moved

Cembalest has Nvidia’s forward price-to-earnings around 18 to 20, “lowest it’s been.” On a July Compound conversation, Jonathan Thomas put it at 18 — below the S&P. Josh Brown’s line: the stock no longer reacts to earnings upside, higher guidance, or upgrades. “People are acting as though it’s been disrupted already.” Brown does not call it a bubble. A bubble, he said, would have been the multiple going from 30 to 90. This is a discount.

Thomas also offers a second explanation for the same print: a worry that the spending ends sooner than people thought. Competition and capex duration both fit an 18-times multiple. Neither explanation has been retired.

Jensen said niche

Jensen Huang, on Nvidia’s May earnings call, called LPX and other SRAM-based decode accelerators a niche product for some time. That is the company’s own answer to Feldman, and it is why the erosion chain sits at low-to-medium conviction. GPT is still trained in CUDA. SemiAnalysis still lists CUDA as one of only two stacks with Day 0 support for DeepSeek V4 — a friction that sits against “10 keystrokes.”

In December 2025 Nvidia licensed Groq’s language-processing unit, the inference architecture that senators Warren and Blumenthal later described as genuinely faster and more efficient for certain workloads, in a $20 billion acqui-hire. That is a defense of the inference franchise, not proof it was never threatened. SemiAnalysis has since upgraded AMD to “a great chance of success” at the software layer — conditional on a Helios rack ramp and stable internal clusters — and then added the sentence that keeps the pie intact: AMD share need not mean Nvidia does poorly if the market is growing.

Then the bench talked back

On 24 August SemiAnalysis published AgentX, a measured counter to Feldman’s “no role whatsoever.” On realistic multi-turn agentic coding, after 21 August vLLM optimizations, Nvidia’s B200 surpassed AMD’s MI355X on DeepSeek V4 Pro performance per dollar. Qwen3.5 on SGLang is the hold SemiAnalysis calls strong: more than 20 times at 90 tokens per second per user, “zero competition from AMD.” GLM 5.3 at 150 tokens per second per user: Nvidia up to five times as cost-efficient — “even if the competitor chip hardware was sold for free,” because the datacenter still pays for power.

AMD’s ATOM wins some total-cost slices — part of a 40-to-60-second end-to-end band against GB300 — but labs will not run it in production, SemiAnalysis says, besides one small Alibaba advertising unit. The main Qwen org does not use it. The gap, in that read, is software rather than silicon: agentic inference is a systems and key-value-cache problem, not a kernel problem. The same frictional-moat sentence the shop has carried since DeepSeek V4’s Day-0 CUDA advantage.

The next morning Neil Movva, hiring for batch agents that crawl at one to ten tokens a second, talked from the other side of the table. “I don’t look for CUDA experience at all. That’s actually a huge red herring.” NVLink, he said, is mandatory for low-latency inference, and Nvidia is excellent at that — “and I’m telling you that we don’t really care that much about low latency inference.” Other vendors’ chips that cannot talk to their peers quickly can still be “a really, really good compute per dollar option.” He is optimizing throughput cost for long-running agents, not the interactive AgentX bench. The wiki records both. It does not merge them. It does not flip the re-rate.

A day later the same shop published a different bench. OpenAI’s Jalapeño A0 — Broadcom silicon, TSMC N3P, HBM4 likely Samsung — beats Blackwell on InferenceX tokens per megawatt without speculative decode. SemiAnalysis verified the runs in the lab. They did not run AgentX. The numbers are OpenAI-provided. Eight-thousand-in, thousand-out is an easier prompt than multi-turn agentic coding. Rubin is shipping; Jalapeño is engineering samples, production in 2027. The author’s own inference: the CUDA moat is potentially dead given how fast OpenAI can bring up models on its silicon. That is a sentence, not an AgentX result. The wiki leaves the pair unmerged.

The August 26 print does not close it. Revenue $96.2 billion, data center $89.0 billion, a third-quarter guide of $108.0 billion plus or minus 2 percent with no China data-center compute in the guide. Vera Rubin, Huang said, is in full production. That is hardware demand. His newsroom line — “compute is revenue,” CUDA-X inside an Agent Toolkit — is not a bench against Feldman or against Jalapeño. AgentX and InferenceX stay unmerged. Conviction on the erosion chain stays low-to-medium.

The All-In hosts, first on an incomplete clip and then on the speaker-labeled full episode, treated the quarter as a blowout that shreds the AI-capex-bubble story. Jason Calacanis, now attributed: $96.2 billion against a $92 billion Street number, a 70 percent growth guide. David Sacks: the capex-bubble narrative is “getting shredded”; Nvidia is “only trading at 12 times earnings.” Those sit next to the newsroom print — $96.221 billion of revenue, $89.0 billion of data center, 75.0 percent gross margin, a $108 billion plus-or-minus-2-percent guide, $279 billion of commitments — not on top of it. Calacanis also recalled Sam Altman buying $100 billion of Nvidia in 2025, then AMD, then “jalapeno” inference chips. That is recollection, not a bench. It still does not merge AgentX with Jalapeño. Conviction on the erosion chain is held.

September 8 added a third inference print from the same shop, and still did not merge the benches. SemiAnalysis’s first third-party InferenceX Official Preview for Google’s TPUv7 Ironwood — a partial note; the total-cost model sits behind a paywall — put aggregated FP8 serving at “up to 50% better performance per dollar” versus B200 and B300. At 100 tokens per second per user: Ironwood $0.181 per million total tokens against B200 $0.222 and B300 $0.276. TorchTPU is expected to leave private beta around October; the shop thinks vLLM and SGLang may put TPU on the day-0 list. Anthropic, on the same grain, is “over one million” TPUs — about 400,000 purchased, about 600,000 rented. Nvidia still leads native FP4 until TPUv8i Boardfly. GB300 NVL72 still wins a disaggregated-versus-aggregated comparison by about 30 percent in the mid-latency band until TPU disaggregation is optimized. A third alternative. Not an adjudication of AgentX versus Jalapeño. The wiki does not re-date Nvidia.

September 15 was the same shop again, on AgentX, and still not a merger. SemiAnalysis’s Vera Rubin NVL72 note — partial, early software — models about 67 times the total throughput per TCO versus GB300 running Dynamo TRT-LLM at 170 tokens per second. Modeled annual profit per gigawatt is about twice the strongest Blackwell. Same vendor. A generational print on the bench that still had Nvidia winning in August is not a sentence about Jalapeño, Ironwood, or Feldman. Conviction on the erosion chain stays low-to-medium, below the 0.4 floor. The wiki does not re-date Nvidia.

Three days later the same shop, on Engrams, put CUDA back on the Day-0 side of a different model. InferenceX on DeepSeek-V4.1 Flash had Nvidia’s vLLM image on H100 through GB300 from the first day. AMD’s vLLM image was missing for 23 hours, then two to four times worse performance per dollar versus B200 — up to 14.8 times versus H200 and 42 times versus B200 and B300 on the late image. SemiAnalysis: “the CUDA Moat is still mogging MI355X.” Stack readiness, not a merge of AgentX and Jalapeño. Conviction held. Do not re-date Nvidia.

Wiki this weaves