2026 06 09 Feed Semianalysis Deepseekv4 Inference Performance Over Time
SemiAnalysis benchmarks DeepSeek V4 1.6T inference across accelerators: GB300 NVL72 unbeatable; Huawei Ascend 950 (CANN) hits Day-0 parity — one of only two Day-0 stacks alongside CUDA; AMD ROCm collapses Day-0 then recovers 100x. CUDA's Day-0-support moat persists.
view source ↗Summary
SemiAnalysis tracks DeepSeek V4 1.6T inference performance across accelerators over 43 days, and the load-bearing causal claim is about software readiness as the moat: CUDA and Huawei's CANN are "the only two stacks with Day 0 Support for DeepSeekV4," while AMD's ROCm "delivered confusing results with unusably low interactivity (1–2 tokens/second)" on Day 0 before a dramatic recovery. The post bears directly on cuda-moat-erosion-at-inference / cuda-moat-erosion-to-nvda-rerate (the moat persists via Day-0 support and the vLLM/SGLang ecosystem) and on the China-substitution risk (Huawei Ascend 950 "David" achieving day-zero parity on a Chinese model is "unprecedented," though "whether it can fell a moving giant is yet to be seen"). inference-speed-as-a-pricing-premium gets quantified data: GB300 NVL72's single-NVLink-domain scale-up advantage gives ~$0.156 / M output tokens at 50 tok/s/user.
Provenance / paywall note. This clipping is the free (non-paywalled) portion of the SemiAnalysis post, extracted via WebFetch; the article is truncated at "This post is for paid subscribers" before the total-cost-of-ownership / cost-per-token section.
partial: true. Quoted passages are verbatim; section structure is the extractor's.
Article (free portion — extracted)
Performance hierarchy (as of June 6, 2026). GB300 NVL72 decisively leads across all tested interactivity levels. With Multi-Token Prediction (MTP) enabled, the authors state: "serving with the GB300 is unbeatable across all interactivity levels." Cost reaches approximately $0.156 per million output tokens at 50 tokens/second/user.
Secondary tier:
- B300 (single-chip): 3× throughput improvement over the first week via grouped FP4 MoE optimizations.
- B200: similar trajectory to B300; TensorRT-LLM superior at lower interactivity but requires manual fixes; native CUDA vLLM/SGLang work "out of the box."
- MI355X (AMD): "over 100x performance improvement in less than a month" from a Day-0 baseline, eventually matching or exceeding H200 at lower interactivity — but only after substantial engineering catch-up.
- H200: baseline; outperformed by optimized MI355X and all Blackwell variants at equivalent batch sizes.
CUDA / Nvidia software advantage. "With CUDA, distributed inferencing tends to be supported near Day 0 for the latest open models." Native vLLM and SGLang implementations shipped immediately; TensorRT-LLM required engineering fixes (the authors' own PR #13710 corrected a hardcoded hidden-size bug). The vLLM/SGLang ecosystem is spawning startups (Inferact, RadixArk).
Chinese chip catch-up (Huawei Ascend 950DT). Huawei achieved day-zero parity with NVIDIA — described as unprecedented: "The Huawei CANN stack is one of only two stacks with Day 0 Support for DeepSeekV4, the other being Nvidia's CUDA." The Ascend 950's internal codename "David" is characterized as aspirational: "whether it can fell a moving giant [Nvidia] is yet to be seen," given Nvidia's annual architecture cadence.
AMD software weakness. ROCm suffered an initial collapse on Day 0 (MI355X achieved only 1–2 tokens/second interactivity). While AMD's team (HaiShaw) recovered dramatically, the article criticizes strategic misallocation: "AMD is re-focusing on ATOM (an inference engine that serves 0 production tokens) instead of focusing on native vLLM."
Power efficiency. B200 vLLM improved ~1.7× from Day 0 to June 5 — from ~300k to ~500k tokens/second/MW — reflecting pure software gains despite a fixed hardware power envelope (~2.17 kW/GPU).
Scale-up vs scale-out. GB300's NVL72 puts 72 GPUs in a single NVLink domain, enabling wide expert parallelism without spilling to slower scale-out fabric — favoring single-rack GB300 NVL72 over distributed B200/B300 deployments on cost-per-token for MoE workloads.
[Paywall — total cost of ownership / cost-per-token section not available.]