brain/
conceptstock-market

Concurrency per megawatt

Notes

Concurrency per megawatt

One-line summary: As raw inference speed commoditizes across a new wave of ASICs, the buying criterion shifts from "how fast" to "how many users can I serve at a guaranteed speed per megawatt" — power, not silicon, is the denominator, so tokens-per-watt hardware efficiency becomes the differentiator.

The insight

The Etched founders argue the industry is exiting the phase where speed alone wins (inference-speed-as-a-pricing-premium) and entering one where speed is table stakes: customers fix the interactivity their product needs, then maximize concurrent users inside a fixed power envelope. Because power availability — not chips — is the binding constraint ("the more power you want, the more shortage of the rates"), the hardware race becomes tokens-per-megawatt. This connects the inference-silicon theses to the power-scarcity theses already tracked (operating-profit-per-gigawatt, nuclear-baseload-for-ai-data-centers).

Caveat: articulated by founders of a private inference-ASIC company (etched) whose product is designed to win exactly this metric, on a podcast hosted by their own investor. Directionally plausible, self-servingly framed.

Evidence

Design implications

  • If speed commoditizes, the inference-speed-as-a-pricing-premium moat (Cerebras/CBRS thesis leg) has a shelf life; the durable metric becomes tokens-per-MW.
  • Reinforces power-scarcity beneficiaries (generation, cooling, 800VDC distribution) as the constant winners regardless of which silicon vendor takes share.
  • Tradeable angle runs through the canonical mechanism: inference-asic-wave-to-tsm-demand-broadening.

Contradictions / tensions

  • Interested-party sourcing (see etched sourcing caveat); no independent benchmark of the "order of magnitude more concurrency" claim.
  • jensen-huang's counter (via inference-demand-to-wafer-scale-advantage): specialized decode accelerators stay "niche for some time" — if GPUs hold the bulk of inference, per-MW concurrency differences may not re-route demand.

Open questions

  • Do neoclouds/hyperscalers actually publish or procure on concurrency-at-interactivity-per-MW metrics, or does TCO-per-token remain the criterion?

Related

Referenced by