brain/
sourcestock-marketartificial-intelligencetechnology-adoption-s-curves

Autoresearch: AI inference revenue run-rate vs realized enterprise AI ROI

The bear and bull camps are measuring different things: per-SEAT/license spend is genuinely decelerating (Forrester 25% postponed; Microsoft and Uber capping licenses) while per-TOKEN consumption is compounding (Google 3.2 quadrillion tokens/mo, 7x YoY; OpenRouter 25T/week, 5x in 6mo). Blended token unit price FELL 67% YoY ($18.40→$6.07/M) — which directly contradicts the 'token costs doubling every ~45 days' claim; bills rise because volume outruns a falling price. The '0-2% ROI' stat traces to an Aug-2025 MIT NANDA study with contested methodology and a disclosed conflict of interest.

Source

Autoresearch: AI inference revenue run-rate vs realized enterprise AI ROI

Generated by /autoresearch on 2026-07-16. Synthesized across 3 rounds from web sources. No Grokipedia anchor attempted (fast-moving contemporary topic — encyclopedia coverage lags by design). Treat as raw material — review before promoting. Context: vault/projects/stock-market No priors captured (headless run — no interactive user turn available).

Summary

The AI-ROI dispute flagged in the 2026-07-15 dispatch (ai-roi-reckoning) largely dissolves into a measurement error: the bear case and the bull case are counting different line items. Both sets of facts are true simultaneously.

  • Per-seat / license spend is genuinely under attack. Forrester says enterprises will postpone ~25% of planned AI spend to 2027; Microsoft killed internal Claude Code licenses at $500–$2,000/engineer/month; Uber capped agentic coding tools at $1,500/employee/month.
  • Per-token consumption is compounding violently. Google processes >3.2 quadrillion tokens/month (7× YoY); OpenRouter runs 25T tokens/week (5× in six months); 67% of enterprises consume >1B tokens/month.
  • The two are causally linked, not contradictory. Anthropic restructured enterprise contracts to kill fixed-seat bundles in favor of a low base seat fee plus token consumption at API rates. The seat is being replaced by the meter. Cutting licenses while token volume explodes is exactly what that transition looks like from a CFO's expense report.

The single most load-bearing correction: blended token unit price fell 67% YoY, from $18.40 to $6.07 per million tokens (Q1'25→Q1'26). Chamath's claim (per the 07-15 ingest) that "token costs are doubling every ~45 days" conflates unit price with total bill. Unit price is deflating fast; volume is what's doubling, and it outruns the price decline. That is a materially different — and for asset-heavy compute, more bullish — mechanism than the one the bear case asserts.

Second correction: the "0–2% realized ROI" figure traces to MIT NANDA's "The GenAI Divide" (Aug 2025) — an ~11-month-old study, predating the agentic wave, whose methodology is contested and whose publisher sells $250k corporate memberships for the protocol it recommends as the fix.

Read for the book: this strengthens the consumption/asset-heavy chains (MU, TSM, compute) and — importantly — independently strengthens the seat-erosion short (agentic-ai-seat-erosion-to-saas-rerate, NOW), because the seat-to-meter transition is the same event viewed from either end.

Findings

The bear evidence is real, and it is specifically about seats and pilots

Forrester's prediction is that "many enterprises will delay a quarter of their planned AI spending until 2027 as they struggle to see a return on investment" (CIO). Brian Hopkins, VP of emerging technology at Forrester, noted roughly half of organizations in Forrester's financial-services and healthcare client base plan to postpone spending. Supporting metrics: only 15% of AI decision-makers reported AI-related earnings increases in the past year; fewer than one-third can link AI to income increases.

Concrete cost-control actions in 2026 (AI Business Weekly):

  • Microsoft terminated internal Claude Code licenses after "per-engineer bills hit $500-$2,000 per month," redirecting engineers to GitHub Copilot CLI.
  • Uber set a "$1,500 monthly spending cap per employee per agentic coding tool" after depleting its "entire 2026 AI budget in four months."
  • An unnamed enterprise reportedly "spent $500 million on Claude in a single month with no spending controls in place."
  • Gartner: "fewer than one-third of corporate decision-makers… could identify specific financial outcomes attributable to their AI investments."

Note what the Uber datapoint actually says. Burning a full-year AI budget in four months is demand exceeding budget, not demand failing. The $1,500 cap is a rationing response — evidence that consumption at the offered price outran the plan. Rationing is what a shortage looks like from the buyer's side.

The bull evidence: token volume is compounding, and it is measurable

  • Google: monthly tokens processed went 9.7T (May 2024) → ~480T (I/O 2025) → >3.2 quadrillion (May 2026) — ~330× over two years, 7× in the last year (TECHi, Crypto Briefing). 8.5M+ developers build monthly with Gemini; 375 Google Cloud customers each consumed >1 trillion tokens in the 12 months prior to I/O 2026.
  • OpenRouter: 25 trillion tokens/week (~100T/month), up 5× in six months, across 8M+ users and 400+ models; API throughput ~19B tokens/minute — a run-rate consistent with sustained inference, not batch (Tech Startups, Yahoo Finance). CEO Alex Atallah: "Running inference at scale is fundamentally a multi-model problem. The era of picking a single model is over."
  • Deloitte (2026): 67% of enterprises already consume >1 billion tokens/month.
  • Model-vendor run-rates: Anthropic reached $30B annualized (7 Apr 2026), passing OpenAI's $25B — from $1B in Jan 2025 ($1B→$4B Jun'25→$9B Dec'25→$30B Apr'26), i.e. 30× in 15 months, with ~85% of revenue from enterprise/developer customers (SQ Magazine, Trending Topics). OpenAI grew ~$20B→$25B over the same window (~25%), with ~85% tied to ChatGPT consumer subs of which ~95% pay nothing.
  • Hyperscaler capex is being raised, not cut. The big four are tracking ~$725B combined 2026 capex, +77% vs ~$410B in 2025, 75% AI-related ($450B) (Yahoo Finance, CFA analysis). Microsoft added $30.88B in fiscal-Q3 capex (+84% YoY) with AI revenue past a $37B annual run-rate; Meta raised FY guidance to $125–145B; Alphabet spent $35.67B in Q1 with Google Cloud backlog >$460B; Amazon $44.2B quarterly capex, AWS +28%, its chip business at a $20B run-rate.

The reconciliation: the seat is being replaced by the meter

The decisive structural fact — Anthropic restructured enterprise contracts starting late 2025, formally landing early 2026: fixed-seat bundles are gone, replaced by "a low base seat fee, plus full token consumption billed at API rates" (Optimum Partners). Salesforce CEO Marc Benioff said on the All-In podcast that Salesforce is on track to spend $300 million on Anthropic tokens in 2026.

This explains every apparently-contradictory datapoint at once: a CFO cancelling per-seat licenses while the token bill explodes is not a company losing faith in AI — it is a company migrating from a seat-priced contract to a metered one. Both facts print in the same quarter.

Corroborating the direction of travel: Morgan Stanley projects inference will be 70–80% of AI compute spending by 2027, shifting AI from a one-time capital project to a recurring, usage-driven operating cost (Optimum Partners).

Aggregate spend is rising, not falling: Gartner forecasts worldwide AI spending ~$2.59T in 2026, +47% YoY; enterprise gen-AI spend went ~$11.5B (2024) → $37B (2025); the average enterprise AI budget grew from $1.2M/yr (2024) to $7M (2026) (Optimum Partners, Vaasblock).

The unit-price correction — the sharpest finding

Blended AI cost fell 67% YoY, from $18.40 to $6.07 per million tokens between Q1 2025 and Q1 2026, yet 73% of enterprises reported AI costs exceeded original projections (Optimum Partners).

Both at once means one thing: volume growth is outrunning a steeply deflating unit price. This is the classic Jevons shape — falling price per unit, rising total expenditure.

Consequences for the book:

  1. It contradicts the specific mechanism in the ai-roi-reckoning bear case as ingested on 07-15 ("token costs doubling every ~45 days"). Token prices are collapsing; bills are rising. The bear's own observation (rising bills) is real; his stated cause (rising unit costs) is inverted. Calibration-worthy.
  2. It qualifies token-price-inflation-favors-asset-heavy-compute — the endpoint (asset-heavy compute captures the value) survives, but the stated driver (token price inflation) is not what the data shows. The driver is volume growth outpacing price deflation. The concept's chain should be re-argued on volume, not price, or it rests on a premise the evidence cuts against.
  3. It supports enterprise-token-budgeting and the consumption-exposed names (MU, TSM) over per-seat-exposed software.

The "0–2% ROI" stat is old and contested

The 95%-of-pilots-fail figure comes from MIT NANDA's "The GenAI Divide: State of AI in Business 2025," published July–August 2025 — 150 leader interviews, 350 employee surveys, 300 public deployments (Fortune, Healthcare IT News). Caveats that matter when it's cited as current evidence:

  • It is ~11 months old and predates the agentic-coding wave that produced the very consumption growth in dispute.
  • Narrow success definition — "deployment beyond pilot phase with measurable KPIs" and "ROI impact measured six months post pilot," excluding efficiency gains, churn reduction, conversion, pipeline velocity (Marketing AI Institute).
  • Disclosed conflict of interest — NANDA is an MIT Media Lab project reportedly charging ~$250,000 for corporate memberships, and the report promotes the NANDA protocol as the solution; critics call it "marketing disguised as science" (Fortune).
  • Even within it, 5% of integrated pilots are "extracting millions in value" — and ~80% of pilot-to-production work is data engineering/governance/integration, i.e. the failures are implementation, not model capability.

Corroborating-but-independent bearish reads that do NOT depend on MIT: S&P Global found 42% of companies abandoned most AI projects in 2025; IBM put initiatives delivering expected ROI at 25%; KPMG reports spending climbing while ROI stays elusive (Techstrong). Only 7% of respondents report an established measurable return (Vaasblock).

Which public companies each side of the evidence favors

Favored by the consumption-compounding evidence (volume grows, unit price falls, inference → 70–80% of compute spend by 2027):

  • MU — HBM is the physical bottleneck on inference volume; volume growth is the demand driver, price-agnostic to token deflation (hbm-cowos-as-binding-bottleneck).
  • TSM — CoWoS/advanced packaging, same logic.
  • NVDA / AI-infra broadly — but note the cluster is already 53% of the book (breadth report), so this adds little independent edge.
  • Power/materials cascade (CCJ, CEG) — token volume is ultimately a megawatt claim (ai-capex-to-power-and-materials-cascade).

Favored by the seat-erosion evidence (Anthropic killing fixed-seat bundles; Microsoft/Uber cutting licenses):

  • Short/derate per-seat softwareagentic-ai-seat-erosion-to-saas-rerate (NOW). This is the finding's most under-priced implication: the bear evidence for AI spending is simultaneously bull evidence for the seat-erosion short. The same Forrester/Microsoft/Uber facts that read as "AI disappointing" read as "the seat is dying" — and the book already has that thesis.

Genuinely damaged by the evidence:

  • Nothing cleanly. The bear case survives as a pilot-implementation critique and a rate-shock vulnerability (per Colas, 07-15 ingest), not as evidence of demand roll-over.

Contradictions and open questions

  • Is OpenRouter's 100T monthly tokens paid? The source "does not specify whether the 100 trillion monthly tokens represent paid inference spend or include free usage." Google's 3.2 quadrillion likewise blends free consumer surfaces with paid API. The volume series is not a clean revenue proxy — this is the weakest link in the bull leg, and the honest counter to it. Anthropic's $30B run-rate (85% enterprise/dev) is the cleaner paid signal.
  • Does the 67% unit-price decline hold going forward, or was it a one-off from a model-generation step? If price deflation stalls while volume growth also decelerates, the Jevons argument weakens on both legs at once.
  • Forrester's 25% postponement vs Gartner's +47% total spend — these are reconcilable (postponing a quarter of planned growth still leaves large growth) but no source reconciles them explicitly. Worth watching whether "postponed" becomes "cancelled."
  • Uber's cap: rationing or retreat? Read here as rationing (demand > budget). A genuine retreat would show as falling absolute token spend at Uber — not observed either way in these sources.
  • Chamath's "gains have asymptoted" — untested by this pass. The economic evidence here speaks to spend and volume, not to model-capability slope.

Provenance

Rounds run: 3 (full)

Sub-questions by round:

Round 1 (broad survey):

  1. Is there hard evidence enterprise AI spending is decelerating in 2026?
  2. What is the actual AI inference revenue run-rate and its growth?
  3. What is the provenance of the "0–2% realized ROI" claim?
  4. Are hyperscalers cutting or raising AI capex guidance?

Round 2 (drill-down):

  1. Is the Forrester 25%-postponement finding about licenses/seats or total spend? — targeting whether "deceleration" is category-specific
  2. Do the Microsoft/Uber cost-control actions distinguish seat spend from token spend? — targeting the same
  3. Is token volume independently measurable and still compounding? — targeting the bull leg's hard data

Round 3 (resolve remaining uncertainty):

  1. Is the compounding token volume paid consumption or free usage? — targeting the bull leg's weakest link
  2. Has the MIT 95% study been criticized on methodology / is it current? — targeting the bear leg's foundation
  3. Is enterprise paid API/consumption spend rising, and how are contracts priced? — targeting the reconciliation

Anchor source: no Grokipedia anchor attempted — fast-moving contemporary topic where encyclopedia coverage necessarily lags; web/primary sources are the appropriate anchor.

URLs fetched / searched (key load-bearing sources):

Round 1:

Round 2:

Round 3:

Source-quality caveat: several of the aggregator/analysis domains above (Optimum Partners, Vaasblock, AI Business Weekly, TECHi) are secondary commentary, not primary research. The load-bearing 67% price-decline figure rests on a single secondary source and should be treated as partial until corroborated against a primary price series (e.g. vendor API price sheets or an index like Epoch AI's) — flagged as the first gap for a follow-up pass.

Tools used: WebSearch, WebFetch. Generated: 2026-07-16

Referenced by