brain/
sourcestock-market

2026 07 25 Feed Semianalysis CAN AMD Break THE Cuda Moat

SemiAnalysis upgrades AMD from '0% chance' (early 2025) to 'a great chance of success' at eroding the CUDA moat — conditional on fixing internal-cluster instability and Helios rack production. The frontier has moved from single-node to disaggregated (multi-node) inference; AMD's software-composability problem is the gate. Anthropic (2GW), Microsoft/Azure (OpenAI end-customer) committed.

view source ↗
Source

Summary

The load-bearing claim: SemiAnalysis upgrades AMD's odds of eroding Nvidia's software moat from "0% chance" (early 2025) to "a great chance of success"conditional on solving two risks (unstable internal GPU clusters for ROCm dev/CI, and a slow Helios rack production ramp). The competitive frontier has shifted from single-node aggregated inference to multi-node disaggregated inference, where AMD's "software composability problem" (individual optimizations work; combining them breaks) is the binding gate. Hyperscaler commitments (Anthropic 2GW; Microsoft/Azure Helios with OpenAI as end-customer; Meta's "about-face" ordering a cut-down MI455X) give early revenue. Tradeable: independent written corroboration for cuda-moat-erosion-at-inference / cuda-moat-erosion-to-nvda-rerate and open-source-share-shift-bullish-for-compute — but the erosion is inference-specific and conditional, not a training-side or unconditional NVDA short. Part 3 (TCO / the "up to 105% equity-rebate discount" for OpenAI/Meta) is paywalled.

Article

(Written analysis — source-attributed. Verbatim quotes where captured; partial: Part 3 TCO section is behind a paid-subscriber wall. Published 2026-07-25.)

Central argument. The analysts upgraded AMD's competitive position: "Based on our experience on AMD software stack this year, we update our view again from non-zero percentage chance to now a great chance of success as long as AMD solves the two major risks we outline below."

Risk 1 — Helios manufacturing delays. "AMD's first AI rack scale system, Helios, is … currently going through a slow rack production ramp given that it is not using a cableless tray design." Requires 550+ Broadcom ethernet retimers per rack (85% of backplane links); flyover-cable design (which Nvidia has abandoned for cableless) creates assembly reliability issues; 1,728 flyover cables per rack.

Risk 2 — unstable internal GPU clusters. "The chief complaint from most AMD engineers internally is that there is a persistent lack of stable GPU clusters for internal software development teams and a lack of stable GPU clusters for automated testing CI." This bottlenecks distributed-inference optimization and AI-agent-driven dev velocity.

MI455X silicon / Helios architecture. First datacenter 2nm silicon (vs Nvidia N3); 3,470mm² logic per package; 432GB HBM4 (12 stacks) vs Nvidia 288GB; 23.3 TB/s bandwidth vs Rubin's 22. But "MI455X is lacking the 3 bit LUT tensor cores that Rubin SM107 has … AMD is forced to be aggressive on silicon to compensate for this deficiency." Helios: 72 GPUs, 12× Broadcom Tomahawk 6 switches (merchant, not custom), 1.8 TB/s scale-up per GPU via UALoE. Competitive on peak specs but "lacks Nvidia's co-design rigor."

ROCm software progress. Improvements: vLLM added AMD mirrors + gating for eight test groups (June 2026); SGLang merged disaggregated nightly tests for DeepSeek-V4; up to 18× Kimi K2.5 latency improvement via AITER; MiniMax M3 now competitive with B200; official ROCm inference docs + upstream alignment. Gaps: Kubernetes Pollara NIC CI at "0% parity with Nvidia's ConnectX"; vLLM gating "massively regressed" due to cluster instability; MI455X (gfx1250) bring-up lacks distributed testing. "Gating/blocking tests mean the tests are of the highest quality because PRs cannot merge unless the test is passing."

Disaggregated-inference composability problem. AMD can execute individual optimizations (WideEP, DP-attention, tensor-parallel) but combining them frequently breaks — e.g. DeepSeek-R1 DP-attention scored near 0% on GSM8K until patched; an accuracy cliff at concurrency 64 remains unfixed. "The competitive frontier has moved from single node aggregated to multi node disaggregated inference." AMD's MoRI (RDMA framework) relies on ~5–6 engineers and lacks production-scale CI; upstreaming to Nvidia's NIXL transport (June 2026) was necessary but late.

Hyperscaler adoption. Anthropic: 2GW of AMD chips, using Claude agents to bring up the inference stack on AMD hardware. Microsoft/Azure: MI455X Helios deployment with OpenAI as primary end-customer. Meta: "about-face" after skipping MI325X/MI355X; now orders a cut-down MI455X (4 compute dies vs 8, 6 HBM stacks vs 12) for recommendation systems — which the analysts criticize: "AMD needs to … make sure they get the normal MI455 instead of the gimped Meta custom version."

Pricing / TCO (paywalled header claim). "AMD gives Meta and OpenAI close to a 105% equity rebate discount … Helios performance per TCO is so great that … cost per million tokens is practically negative cost." Mechanism: stock options triggering at AMD $600/share plus volume purchase — "AMD is practically giving away Helios racks and [a]n 5% extra on top of that to an SF-based nonprofit called OpenAI." Full TCO math truncated: "This post is for paid subscribers."

Tradeable implications. AMD: conditional upgrade to "great chance" if the two risks resolve. NVDA: "just because AMD gains market share, that doesn't mean that Nvidia will do poorly. The pie is growing rapidly for everyone, and Nvidia will continue to massively grow revenue." But "Jensen will need to cut bureaucracy and flatten the layers … if he wants Nvidia to move faster and defend their lead." Analysts frame disaggregated inference + WideEP as Nvidia's emerging moat, with Nvidia's GTC 2026 Groq 3 LPX (SRAM decode LPU) a counter-move.

Referenced by