AI ROI reckoning: token-cost blowup vs. flat realized productivity
AI ROI reckoning: token-cost blowup vs. flat realized productivity
One-line summary: A cluster of same-week July-2026 sources argues AI token spend is compounding vertically while realized enterprise ROI is near-zero and marginal model gains have flattened — and that the market has already stopped rewarding the capex spenders, the classic late-cycle tell.
The insight
Three independent July-2026 sources converge on the same warning: the AI-capex build-out may be running ahead of its returns. Chamath cites token costs "doubling every 45 days" against realized enterprise ROI of "0 to 2%," with model improvement "asymptoted." DataTrek's Josh Brown frames "the number one question facing investors" as how long the market rewards hyperscaler capex — and observes the spenders (Meta, Oracle) already aren't being rewarded, the 1999 pattern where spender stocks roll before supplier/"darling" stocks do. Nick Colas' first-hand 1999 correction is the load-bearing nuance: spender weakness alone did not end that cycle — leadership rotated (to B2B) and the actual kill-switch was the Fed / cost of capital. So this is a demand-side skepticism overlay on the AI-capex long cluster, and a watch-item, not yet a confirmed break: the falsifier that turns rotation into a rout is a rate shock, not the ROI number itself.
This is the equity-side twin of the credit-side mega-issuance-peak-to-ai-capex-derate and the demand-side counterpart to the supply-driven compute-as-financialized-commodity.
Evidence
- chamath-palihapitiya in 2026-07-11-podcast-all-in-podcast-more-trillion-dollar-ipos-anthropic-3t-zuck-s: "Our token costs are doubling every 45 days. Honestly, what we're finding out is that you need to use a lot more tokens to get to this next iteration of improvement because we've effectively already asymptoted."
- chamath-palihapitiya in 2026-07-11-podcast-all-in-podcast-more-trillion-dollar-ipos-anthropic-3t-zuck-s: "The actual ROI was somewhere between 0 and 2%. At some point you're going to have to show an ROI that's above the risk free rate of return."
- michael-cembalest in 2026-07-10-podcast-the-compound-and-friends-the-real-ticking-time-bomb-with-michael-cembalest: (per the GS-financing mechanism) hyperscalers moved from funding capex out of internal cash flow to borrowing to do it — the financing side of the same reckoning.
- nick-colas in 2026-07-13-podcast-the-compound-and-friends-the-number-one-question-facing-investors-how-the: "The cause of that implosion was 110% the Fed. It was the realization like, oh my God, the cost of capital was going to go away very quickly." (i.e., spender roll-over alone is not the sell signal.)
- sam-altman in 2026-07-28-podcast-invest-like-the-best-sam-altman-how-to-make-an-abundant-future-invest (the demand-side principal, talking his book): names two paths to oversupply — (a) a demand ceiling ("if the bounds of our attention are such that they just cannot absorb more than… a fairly limited amount of compute can do, then we can get into oversupply"), and (b) a scaling wall ("if we don't drive the cost curve down because we hit some sort of scaling wall, we could also get into oversupply"). Framing: "the observation about uncapped demand implies a certain price" — i.e., demand is uncapped only at a sufficiently low price, so the reckoning is a cost-curve/price question, not a level question.
- From 2026-07-28-autoresearch-ai-compute-oversupply-vs-sustained-capex-2027 (external corroboration, both sides): the consensus is sustained, not peaking — ~$725B 2026 hyperscaler capex (+77% YoY), DC vacancy at a record ~1.0–1.6% with ~92% pre-committed, "short on compute" fear dominant. The genuine risk is timing, not a supply glut — capex is outrunning cloud revenue (Amazon FCF projected negative 2026). This pressures the near-term timing of this reckoning: the physical glut isn't near, but the FCF/return-timing strain is the real trigger (routes to mega-issuance-peak-to-ai-capex-derate, which fires on a financing break, not physical oversupply).
Design implications
- Treat the AI-capex long cluster (semis, power, neoclouds) as carrying a demand-ROI risk overlay, not just a supply story. The tell to watch is whether capex announcements stop being rewarded (Meta/Oracle already) and then whether that spreads to the suppliers.
- Consumer-AI revenue (many small buyers) is framed as more resilient than enterprise (fewer, more demanding buyers) — Chamath calls enterprise AI revenue "brittle."
Update (2026-07-16) — the dispute largely dissolves into a measurement error; two of the bear case's load-bearing numbers do not survive
A dedicated gap-fill (2026-07-16-autoresearch-ai-roi-dispute-seat-vs-token-divergence) tested this concept's two headline claims. Both are weaker than the 07-15 ingest implied — but the phenomenon they point at is real and better explained by a different mechanism.
1. ⚠ "Token costs doubling every ~45 days" is inverted. Blended AI cost fell 67% YoY, from $18.40 to $6.07 per million tokens (Q1'25→Q1'26), while 73% of enterprises report AI costs exceeded projections. Both are true simultaneously: unit price is deflating fast; volume is what's doubling, and it outruns the price decline. Chamath's observation (bills are rising) is correct; his stated cause (unit costs rising) is backwards. This is the classic Jevons shape — falling price per unit, rising total spend — and it is a materially more bullish mechanism for asset-heavy compute than the one the bear case asserts. Calibration-worthy (see CALIBRATION).
Corroborated independently, on the same day, by two operators who are not talking the same book:
- pat-gelsinger in 2026-07-15-podcast-all-in-podcast-former-intel-ceo-on-what-went-wrong-what-s-next: "I have to make AI 10,000x better. Right. You know, it's way too expensive today. You know, we want to drop, you know, by five orders of magnitude the cost per token… so that we really do have Jevons Law, that we just explode the access to AI" (note: 10,000x is four orders of magnitude, not five — his own numbers disagree).
- michael-batnick in 2026-07-14-podcast-the-compound-and-friends-ibm-warns-apple-sues-openai-big-bank-earnings: "with every step, function, increase in the efficiency and whatever these models are able to do, people are spending way more, not less."
2. ⚠ "Realized ROI 0–2%" traces to an ~11-month-old study with contested methodology and a disclosed conflict. The 95%-of-pilots-fail figure is MIT NANDA's "The GenAI Divide," published July–August 2025 — predating the agentic wave that produced the very consumption growth in dispute. Its success definition is narrow ("deployment beyond pilot phase with measurable KPIs," ROI measured six months post-pilot — excluding efficiency, churn, conversion, pipeline velocity), and NANDA is an MIT Media Lab project reportedly charging ~$250,000 for corporate memberships that promotes the NANDA protocol as the fix; critics call it "marketing disguised as science." Within it, 5% of integrated pilots are "extracting millions in value," and ~80% of pilot-to-production work is data engineering/governance/integration — i.e. the failures are implementation, not model capability.
Independent bearish reads that do not depend on MIT survive: S&P Global — 42% of companies abandoned most AI projects in 2025; IBM — 25% of initiatives delivering expected ROI; only 7% report an established measurable return; KPMG — spending climbing while ROI stays elusive.
3. What actually explains the evidence: the seat is being replaced by the meter. Per-seat spend is genuinely decelerating (Forrester: 25% of planned AI spend postponed to 2027; Microsoft killed internal Claude Code licenses at $500–2,000/engineer/month; Uber capped agentic tools at $1,500/employee after burning its 2026 AI budget in four months) while per-token consumption compounds (Google >3.2 quadrillion tokens/month, 7× YoY; OpenRouter 25T/week, 5× in six months; 67% of enterprises >1B tokens/month; Anthropic $30B run-rate, 30× in 15 months, ~85% enterprise/dev). The link: Anthropic killed fixed-seat bundles in favor of "a low base seat fee, plus full token consumption billed at API rates." A CFO cancelling licenses while the token bill explodes isn't losing faith in AI — they're migrating contracts. Both facts print in the same quarter. Gartner still has total AI spend at ~$2.59T in 2026 (+47% YoY); Morgan Stanley has inference at 70–80% of AI compute spend by 2027.
Note on Uber specifically: burning a full-year AI budget in four months is demand exceeding budget, not demand failing. The $1,500 cap is rationing — what a shortage looks like from the buyer's side.
Net read for the book. This concept survives as a pilot-implementation critique and a rate-shock vulnerability (per Colas, 07-15) — not as evidence of demand roll-over. It should stop being carried as a bear case on AI-capex demand. Its two strongest new implications point the other way:
- Constructive the consumption/asset-heavy leg (MU, TSM — see hbm-cowos-as-binding-bottleneck).
- The bear evidence here is simultaneously bull evidence for the seat-erosion short (agentic-ai-seat-erosion-to-saas-rerate): the same Forrester/Microsoft/Uber facts that read as "AI disappointing" read as "the seat is dying." The book already has that thesis; today it went to
medium-high. - ⚠ It also qualifies token-price-inflation-favors-asset-heavy-compute — that concept's endpoint survives, but its stated driver (token price inflation) is contradicted by the 67% price decline. It needs re-arguing on volume growth outpacing price deflation, or it rests on a premise the evidence cuts against.
The bear case's best surviving form is not Chamath's — it is Josh Brown's, which is about price, not demand: josh-brown in 2026-07-14-podcast-the-compound-and-friends-ibm-warns-apple-sues-openai-big-bank-earnings: "A seismic shift in the pricing of AI due to more efficient models that rely on less token use and less memory." … "If they decide that they're going to use these open weight models for 95% of the workflows and then only send the most critical 5% to the more expensive frontier models… a fifth of the cost… that changes all of a sudden the earnings expectations." That is a coherent, falsifiable bear mechanism that the volume data does not refute — see open-source-share-shift-bullish-for-compute.
⚠ Source-quality caveat: the load-bearing 67% price-decline figure rests on a single secondary source and should be treated partial until corroborated against a primary price series (vendor API price sheets, or an index like Epoch AI's). It is the first gap for a follow-up pass.
Update (2026-07-17) — the adopter leg gets its counter-thesis, and the two are directly opposed
The bear case here rests on realized enterprise ROI being near-zero. Two same-day sources put the opposite claim on the table, and the disagreement is now explicit enough to be worth resolving rather than just recording.
- The counter-thesis — jonathan-thomas (CEO, American Century Investments) in 2026-07-17-podcast-the-compound-and-friends-you-re-about-to-see-the-real-ai-winners-stand-up: "the creators and the enablers of AI were just soaring. And now what's starting to happen is the adopters are starting to receive the benefit from it... If the adopters realize the expected benefit, productivity — because productivity is what drives margin, profits, gdp, the economy, everything — if they realize those productivity benefits that I think are out there, this will tick back up." Note the conditional and the hedge: "that I think are out there" is an assertion, not a measurement — which is exactly this concept's complaint about the bull case. See ai-creators-to-adopters-rotation.
- The stakes — josh-brown in 2026-07-17-podcast-the-compound-and-friends-you-re-about-to-see-the-real-ai-winners-stand-up: "is there a handoff where we don't have to automatically just have a bear market?... That is the handoff where the S&P493 start to outearn." And his own honesty: "I don't know if it's gonna work."
- The earnings-bubble framing this concept implies, named — josh-brown in 2026-07-17-podcast-the-compound-and-friends-you-re-about-to-see-the-real-ai-winners-stand-up: "The new meme going around now is that it actually it's an earnings bubble. They can't say it's a stock bubble because the biggest, most visible growth stocks in the market have shrinking multiples... So what the bears have pivoted to is, oh, no, no, we're not saying the stock's in a bubble now. We're actually saying profitability is. And maybe they'll be right." And the mechanism: "The companies are over earning relative to what happens when this capex normalizes." jonathan-thomas: "That's exactly right." See steroid-era-earnings-inflation.
- The cost-collapse accelerant — jack-farley in 2026-07-17-podcast-forward-guidance-the-ai-unwind-is-forcing-a-historic-market: a new Chinese open-weight model, "its capabilities are right up there with the frontier models... I can get 80% of the capability of Fable 5, but in this open model that costs 10%, that starts to put at question the entire proposition that the NASDAQ is built on right now." If intelligence gets 10x cheaper, the ROI arithmetic this concept tracks changes on the cost side without any productivity improvement at all. See open-source-share-shift-bullish-for-compute.
- The rate/liquidity kill-switch this concept already names, restated — tyler-neville in 2026-07-17-podcast-forward-guidance-the-ai-unwind-is-forcing-a-historic-market: "Right now we're tightening financial conditions and it's going to other sectors that have better growth... when does the liquidity come back?" Consistent with Nick Colas' 1999 read already on this page — the actual kill-switch is the cost of capital, not the ROI number. But Tyler adds a political floor: "AI, like Trump tweeted the other day, it's a national security imperative. I have to imagine that trumps everything at some point — they can't let this derail. They need the debt markets to provide capital for this stuff."
Where this leaves the concept. Unresolved, and now sharply so. Thomas asserts adopter productivity is "out there"; Chamath's cited figure is realized enterprise ROI of "0 to 2%". They cannot both be right, and neither is measured. The falsifiable test is Brown's — count healthcare/financial earnings calls attributing beats to AI workflow investment. Recorded on ai-creators-to-adopters-rotation as the resolving observable.
Update (2026-07-18) — the token-spend blowup gets a hard growth number (Ramp 21×) and a named earnings-miss channel
The All-In panel supplies the cleanest datapoint yet on the volume leg — and reframes the bear case as a CFO-visibility/earnings-timing risk, consistent with the 07-16 "volume outruns price deflation" correction above.
- The 21× print (source-attributed, Ramp CEO Eric Gliman clip in 2026-07-18-podcast-all-in-podcast-can-the-ai-industry-regulate-itself-stripe-wants): "token spend among ramp customers has grown by 21 times" over the last year, launching a "token Spend Management" product because "CFOs... are often very surprised by the bill... every time they're introducing new models, the rates often go up." This is the Jevons volume shape (unit price falling, total spend 21×-ing) as a live CFO pain point.
- The named earnings-miss channel (source-attributed, Chamath, same source): "if things are 21xing every few months, somebody's going to miss a quarter... it could be as much as dollars [of EPS]... unless you get a control of this... This is a bridge to nowhere. It is a money burning furnace." The mechanism: unguided engineer token spend (they "want to use the latest greatest model," are "not tied to the money") → uncontrolled opex → a public-company CFO misses a quarter and attributes it to token spend.
- The cost-disparity that makes it acute (source-attributed, Chamath): Chinese models ~$0.50/M input tokens vs. Fable ~$56 — "you're paying 56 bucks as well per million input tokens for that risk? That is insanity." The open-weight substitution (see open-source-share-shift-bullish-for-compute) is the pressure-relief valve CFOs will reach for.
Read for the book: this strengthens the volume-compounding fact (bullish asset-heavy compute) while giving the bear case its most falsifiable near-term form — watch 2H-2026 for a public-company earnings miss explicitly attributed to token/AI opex. That is Chamath's trigger, now with a growth number attached.
Contradictions / tensions
- Directly contradicts the AI-capex bull thesis — this is the calibration-worthy flag: several credible voices now argue the spend is outrunning returns. But it is balanced by Gerstner ("intelligence is the largest TAM we've ever seen"; the experimental spend "nobody cares" about yet) and Colas' point that leadership rotates rather than dies absent a Fed shock.
- The real falsifier for the whole complex is a Warsh-driven rate shock (see us-recession-resistance-regime), not the ROI number in isolation.
Open questions
- Does an AI-capex earnings miss (the trigger Chamath names) actually materialize in 2H-2026, or does the TAM narrative absorb the ROI gap for another cycle?