brain/
Semis · Cloud economics

Inference beyond the GPU

When inference stops renting Nvidia's stack, silicon demand does not shrink — it broadens across nodes, instruction sets, and cloud business models. Independent lab numbers still do not make inference the bulk of 2026 compute spend.

Covers stock-market wiki · pages updated through September 2026

Training built the GPU empire. A wave of challengers is building inference on supply chains the GPU flagship does not consume. Wafer-scale SRAM chips skip HBM and CoWoS packaging. Custom ASICs run on TSMC 4nm while Rubin sits on 3nm. Agentic workloads are reviving the CPU socket. Each path adds silicon demand rather than substituting for it, which is why TSMC's CEO says the foundry wins regardless of which architecture prevails. That is not the same claim as “inference is already most of the compute bill.”

The ASIC wave on second-tier supply

Etched, one of a cohort of inference ASIC builders, claims order-of-magnitude concurrency advantages over GPUs at a given level of interactivity — translating directly into tokens per watt and tokens per dollar. Its first-generation product runs on TSMC 4nm and a different HBM generation than Rubin. Co-founder Rob Wachen: "It's not a decision between a gigawatt of a GPU and a gigawatt of us, it's 2 gigawatts."

OpenAI validated the workflow-specific ASIC thesis from the demand side. Sam Altman described the company's "Jalapeno" chip as "really good at a specific workflow" and said Jalapeno and its successors would be "a huge competitive advantage." Jalapeño is widely described as Broadcom-co-designed, TSMC-fabbed, leading-edge packaged, eight stacks of HBM — the same sandwich, not an escape from it. TSMC’s second-quarter print still will not size custom ASIC or XPU versus GPU. HPC was 66 percent of wafer revenue, from 61 in the first quarter; 5-nanometer was 33 percent. SemiAnalysis later scored the part the same way Cerebras scores itself: tokens per megawatt, because the halls are “currently limited by datacenter power, not by budget or floorspace.” Jalapeño is a custom ASIC, not wafer-scale. What it implies for Cerebras sits behind a paywall. Attach only. Do not re-rate the wafer-scale name on it.

Epoch’s independent lab numbers do not make inference the bulk of near-term compute spend. OpenAI in 2024: about $1.8 billion of inference against about $5 billion of research-and-development compute. OpenAI in 2025: a roughly even split. Anthropic in 2025: $2.7 billion of inference against $4.1 billion of training. The Information / DECODER put 2026 training inside a $50 billion compute bill at about $32 billion. Speed is sold as a SKU — OpenAI Fast mode, Anthropic Fast mode at two-and-a-half times output tokens per second and twice the list price — but tokens per dollar and concurrency still price the volume market. The wiki keeps that step partial.

Cerebras offers a different architectural escape. Its wafer-scale engine uses fast on-chip memory at scale rather than slow HBM — the reason, CEO Andrew Feldman says, GPU inference is slow. Cerebras claims 15 times faster inference than the fastest GPU, sidestepping all three binding AI-silicon constraints: no HBM, no TSMC CoWoS packaging, and TSMC 5nm rather than the most-constrained 3nm node.

There are three areas right now that are limiting vendors and building AI Compute. Number one is HBM... We don't use it. The second part that's limiting is a process inside of TSMC called COAS [CoWoS]... We don't use it. The third thing is that at TSMC the factory that is under most pressure is their 3 nanometer factory. We don't use it. We use 5 nanometer.

Andrew Feldman, Cerebras

On Cerebras's Q2 2026 earnings call, Feldman restated the constraint routing as a public company: core revenue reached $209.9 million, up 103% year-over-year; cloud revenue hit $127.7 million, up 287%; remaining performance obligations stood at $25.4 billion. The binding constraint, he said, is data-center space — more than 600 megawatts live or contracted through end-2027, still "not nearly enough." TSMC has given Cerebras as many wafers as needed; powered buildings are the wall.

The additivity claim holds only while challengers stay off the leading node. At gigawatt-per-month scale, they plausibly compete for the same wafer, HBM, and CoWoS pools — collapsing the "more is more" logic. SemiAnalysis also flagged that Cerebras's next-generation CS-4 pairs with HBM-based XPUs for disaggregated prefill, a scoped concession on the "we don't use HBM" purity at the system level.

The CPU leg

TSMC's Q2 2026 call opened a fourth demand leg that every other AI-silicon chain in the book had missed. C.C. Wei: "The emergence of agentic AI is leading to a resurgence in the role of CPUs in AI data centers, which drive more silicon demand in addition to AI accelerators." Because x86, Arm, and RISC-V are "almost all TSMC's customers," the foundry captures the CPU leg regardless of which instruction-set architecture wins.

The independent corroboration is a scarcity tell from the buyer side. Server-CPU shortages in Q2 2026 were severe enough to strand DRAM on US cloud-provider balance sheets — you cannot ship a rack without the CPU to pair the memory with. TSMC raised its 2026 capex budget to $60–64 billion from $52–56 billion, naming the "newly emerging agentic AI market" as a driver, and guided full-year revenue growth to slightly above 40%.

Wei declined twice to size the CPU leg — no breakout of CPU versus GPU versus XPU, no updated AI CAGR. The leg is established; its magnitude is not.

Arm royalties and merchant silicon

If the CPU resurgence skews Arm-based, it lands on Arm Holdings' royalty model. Jensen Huang declared at GTC Taipei that the standalone CPU product market, including Nvidia's Vera CPU, would grow to $200 billion. Arm's datacenter royalty revenue "continues to more than double year-over-year" through Q1 FY2027. Rene Haas revised the CPU opportunity upward to $220 billion from the $100 billion discussed in March.

Arm is no longer only a royalty-taker on Nvidia's Vera. It is shipping its own datacenter CPU — the Arm AGI CPU — with demand exceeding $2 billion and manufacturing capacity secured for a $1 billion opportunity. Customers named include Cerebras and OpenAI, Meta and Cloudflare, and Oracle Cloud Infrastructure. Selling silicon captures far more per socket than a 1–2% royalty.

Wei's architecture-agnostic framing cuts against treating the CPU resurgence as purely an Arm thesis. He named x86, Arm, and RISC-V as equal claimants. The bigger pie is confirmed; Arm's slice is a conditional bet.

Cloud rent when tokens replace compute hours

As inference demand scales, the economic model shifts from infrastructure-as-a-service toward token-as-a-service. Amazon Web Services structurally gains because Anthropic's Bedrock distribution pays AWS twice — an infrastructure fee plus a revenue share — without AWS absorbing model development cost.

AWS AI revenue mix, Q1 2026

Bedrock · 37% of AWS AI (was 9%) AWS AI · 10% of rev (was 2%)

SemiAnalysis estimate. Bedrock grew 170% quarter-over-quarter in Q1 2026. Azure and GCP remain 80%+ IaaS.

Bedrock's run-rate revenue reached roughly $5.5 billion at an estimated 55% EBIT margin. AWS EBIT margins expanded 213 basis points quarter-over-quarter in Q1 2026, primarily from customers spending on Claude through Bedrock rather than traditional IaaS. Trainium custom silicon processes more than 50% of Amazon Bedrock token usage at structurally lower cost than Nvidia H100/H200 — a vertical integration advantage Azure and Google Cloud cannot replicate on the same terms.

Anthropic reached $30 billion in annual recurring revenue, adding $21 billion in net new ARR in Q1 2026, predominantly on AWS Bedrock. The +213 basis points figure is an analyst-computed estimate; AWS does not formally break out Bedrock.

Wiki this weaves