Concurrency per megawatt
Concurrency per megawatt
One-line summary: As raw inference speed commoditizes across a new wave of ASICs, the buying criterion shifts from "how fast" to "how many users can I serve at a guaranteed speed per megawatt" — power, not silicon, is the denominator, so tokens-per-watt hardware efficiency becomes the differentiator.
The insight
The Etched founders argue the industry is exiting the phase where speed alone wins (inference-speed-as-a-pricing-premium) and entering one where speed is table stakes: customers fix the interactivity their product needs, then maximize concurrent users inside a fixed power envelope. Because power availability — not chips — is the binding constraint ("the more power you want, the more shortage of the rates"), the hardware race becomes tokens-per-megawatt. This connects the inference-silicon theses to the power-scarcity theses already tracked (operating-profit-per-gigawatt, nuclear-baseload-for-ai-data-centers).
Caveat: articulated by founders of a private inference-ASIC company (etched) whose product is designed to win exactly this metric, on a podcast hosted by their own investor. Directionally plausible, self-servingly framed.
Evidence
- rob-wachen in 2026-06-30-podcast-invest-like-the-best-etched-building-ai-hardware-to-make-inference: "Whatever the speed is, this is my speed. The question is, in a given amount of power, how many users can I serve while guaranteeing that speed?... There's an entirely new wave of AI chips, us being one of them, that are all going to be able to hit these speeds. The question then is, if you're hitting these speeds, what is the number of users you can serve at the same time?"
- rob-wachen in 2026-06-30-podcast-invest-like-the-best-etched-building-ai-hardware-to-make-inference: "Our hardware is going to generally be able to get you an order of magnitude more concurrency at a given level of interactivity. That directly translates into tokens per watt, tokens per dollar."
- rob-wachen in 2026-06-30-podcast-invest-like-the-best-etched-building-ai-hardware-to-make-inference: "One of the things we need to think about is how do we get way more juice out of each megawatt... there's a lot of people trying to scale in their data centers as much as trying to scale them out."
- rob-wachen in 2026-06-30-podcast-invest-like-the-best-etched-building-ai-hardware-to-make-inference (the futurist version): "right now we measure productivity as a society as GDP per capita, but really it's going to look much more like agents per megawatt."
- gavin-uberti in 2026-06-30-podcast-invest-like-the-best-etched-building-ai-hardware-to-make-inference (the physical headroom claim): "chip to chip latencies on an Nvidia product, you're looking at 4,000 nanoseconds... What's the mathematical limit is speed of light. You can do it in just a handful, like 2, 3 nanoseconds."
Design implications
- If speed commoditizes, the inference-speed-as-a-pricing-premium moat (Cerebras/CBRS thesis leg) has a shelf life; the durable metric becomes tokens-per-MW.
- Reinforces power-scarcity beneficiaries (generation, cooling, 800VDC distribution) as the constant winners regardless of which silicon vendor takes share.
- Tradeable angle runs through the canonical mechanism: inference-asic-wave-to-tsm-demand-broadening.
Contradictions / tensions
- Interested-party sourcing (see etched sourcing caveat); no independent benchmark of the "order of magnitude more concurrency" claim.
- jensen-huang's counter (via inference-demand-to-wafer-scale-advantage): specialized decode accelerators stay "niche for some time" — if GPUs hold the bulk of inference, per-MW concurrency differences may not re-route demand.
Open questions
- Do neoclouds/hyperscalers actually publish or procure on concurrency-at-interactivity-per-MW metrics, or does TCO-per-token remain the criterion?