Compute utilization overhang as latent supply
Compute utilization overhang as latent supply
One-line summary: A large fraction of installed AI compute sits idle — independent data centers run <70% chip utilization and far lower per-chip utilization — so a software layer that makes heterogeneous GPUs fungible can unlock 30-40% of "missing" capacity without a single new fab, capping the scarcity premium that underpins the GPU/neocloud/HBM-bottleneck theses.
The insight
The wiki's AI-capex cluster (hbm-cowos-as-binding-bottleneck, ai-capex-to-power-and-materials-cascade, cowos-packaging-capacity-crunch, datacenter-construction-electrical-picks-shovels) rests on a physical-scarcity premise: demand outruns the supply of chips, packaging, power, and shells, so capex routes to whoever owns the constrained input. anjney-midha (founder of AMP PBC, first Anthropic check, ex-a16z) argues a meaningful slice of that "scarcity" is not physical at all — it is stranded utilization that software can reclaim:
- The average independent data center runs <70% chip utilization; xAI's Colossus 2 (Memphis, ~500K GB300) ran "<60% node utilization and <11% MFU [model-flop utilization]." Google, by contrast, runs ~99% (AMP's co-founder ran the Borg utilization project there, taking it from ~62% to ~99%).
- The gap is structural, not lazy: research demand is spiky and hard to forecast, so labs over-provision for peak, not base load, then sit on idle capacity between training spikes. The marketed GPU rate is ~$2.50/hr, but the effective rate after wastage is "$25-28/hr."
- AMP's claim is that a software "grid" — a Borg-for-everyone translation layer that makes any chip type (Nvidia, AMD, other) a fungible "grid credit" — lifts incubated labs from 50-60% to 95-96% utilization, and fills idle pockets with inference while reserving the rest for training.
If even partly true, 30-40% latent supply can be unlocked from the installed base by software, which is a deflationary force on compute price and a cap on the scarcity premium the capex-bottleneck theses price in. This is why it lives here as a tension, not a buy: the cleanest beneficiary (AMP) is private and not tradeable, so the concept's job is to bound conviction on the scarcity side, not to name a new long.
The chain
Installed AI compute is badly under-utilized (<70% at independent DCs, <11% MFU at the chip level) → research demand is spiky so labs over-provision for peak → a software fungibility layer (AMP's "grid") reclaims idle capacity, lifting utilization toward Google's ~99% → effective compute supply expands without new fabs/packaging/power → compute price deflates at the margin → the scarcity premium in GPU/neocloud/HBM-bottleneck valuations is capped. No canonical mechanism page (no tradeable endpoint — AMP is private); this is a bound on the scarcity cluster.
Evidence
- anjney-midha in 2026-06-13-podcast-odd-lots-anjney-midha-s-plan-to-radically-lower-the-price: "Today the average data center in the independent ecosystem is running at less than 70% utilization. The Colossus 2 in Memphis, Elon's 500,000 GB300s, was running at less than 60% node utilization and less than 11% MFU... At Google the utilization is roughly 99%."
- anjney-midha in 2026-06-13-podcast-odd-lots-anjney-midha-s-plan-to-radically-lower-the-price: "We turn all that unutilized compute, no matter what format it is, into one fungible resource... we improve utilization, sometimes from 50, 60% at labs we have incubated to close to 95, 96%."
- anjney-midha in 2026-06-13-podcast-odd-lots-anjney-midha-s-plan-to-radically-lower-the-price: "The effective price per hour you're paying is closer to 25 to $28, whereas the marketed rate you think you're paying is $2.50. That spread due to wastage is just insane."
Design implications
- Bounds conviction on the GPU/neocloud-scarcity leg. A software-reclaimable 30-40% latent supply is a structural headwind for any thesis that prices permanent compute scarcity (neoclouds on take-or-pay; the most aggressive HBM/CoWoS-bottleneck reads). The hardware-supply chokeholds (cowos-packaging-capacity-crunch, hbm-supply-bottleneck) are less exposed than the raw-GPU-count scarcity story, because packaging/HBM is the binding input regardless of how well the GPU is scheduled.
- Corroborates the custom-silicon driver. Midha's "80 cents of every R&D dollar flows to Nvidia, so labs build their own chips for unit-economic + supply-chain independence" is the same force tracked in cuda-moat-erosion-at-inference — another margin-driven erosion of the Nvidia take-rate.
- Watch item, not a position. If a credible public compute-utilization/orchestration name emerges (vs. private AMP/OpenRouter), this concept graduates toward a tradeable; today it is a conviction-bound only.
Contradictions / tensions
- Source is talking his own book — Midha sells the utilization fix, so the <70%/<11% figures are directional, not audited. xAI/Colossus utilization numbers are unverified third-party claims.
- Demand is "perpendicular." Midha himself says long-term-rental compute prices are up 2x Jan→June 2026 and demand is "perpendicular" — so even if utilization rises, demand may swamp the unlocked supply, and the scarcity premium holds. The tension is two-sided: latent supply caps price, runaway demand floors it.
- HBM/CoWoS is orthogonal. Better GPU scheduling does nothing to relax the packaging/HBM physical bottleneck (hbm-cowos-as-binding-bottleneck) — that chokepoint is on making the chip, not using it efficiently. So this tension applies to the raw-compute-scarcity narrative, not the picks-and-shovels packaging leg.
- Google already did it internally (99%), and it did not collapse Nvidia demand — internal hyperscaler utilization has been high for years without deflating the GPU market, which argues the marginal independent-DC reclaim may be smaller than framed.
- Direct contradiction on Colossus economics (surfaced 2026-07-09, unresolved). This page's low-utilization reading is corroborated by rajiv-jain, who put the same cluster at "11% capacity utilization" and concluded "the economics are really bad." Three days later, gavin-baker in 2026-06-11-podcast-bg2-pod-the-spacex-ipo-fable-5-ai-capex-update-market cited the opposite: "she calculated a 55% ARR on Colossus 1" — and put 2026 inference revenue at "well over 200 billion," against Jain's "maybe 70, 80 billion." Utilization and return are not the same metric and may both be true; but if Baker's ARR holds, the latent-supply bound this page places on the scarcity thesis weakens considerably. Neither conviction was moved. Tracked at ai-inference-revenue-run-rate-dispute. See also token-price-inflation-favors-asset-heavy-compute, which argues compute price is inflating, not deflating.
Open questions
- compute-chip-futures-market — Midha procures via call options on compute clusters but warns against financialization; how the forward/derivatives market develops bears directly on whether utilization gains get arbitraged into price.
- csp-capex-cycle-peak-or-sustained — if software unlocks latent supply, does the marginal 2027 capex dollar still get spent, or does it defer?