brain/
sourcestock-market

Autoresearch: 2026 AI inference / AI revenue run-rate — Jain $70-80B vs Baker $200B+

Reconciles Rajiv Jain's June 2026 ~$70-80B 'revenue on AI' and Colossus 11% utilization against Gavin Baker's 'well over $200B inference revenue' and Colossus 1 55% ARR. The gap is mostly definitional plus timing; 11% is MFU not occupancy; Colossus is xAI's. No hyperscaler has split inference from training billings.

Source

Autoresearch: 2026 AI inference / AI revenue run-rate — Jain $70-80B vs Baker $200B+

Generated by /autoresearch on 2026-08-17. Synthesized across 3 rounds from 15 successful web fetches, anchored by Grokipedia inference-economy (and a narrower Colossus_(supercomputer) fetch). See Provenance. Treat as raw material — review before promoting into a project or thread. Context: vault/projects/stock-market

Priors capture skipped — unattended weekday shift; no user present to interview.

Summary

The June 2026 Jain–Baker gap is mostly definitional, plus six weeks of compounding, not a 3× contradiction on one quantity. Jain's "revenue on AI talk about maybe 70, 80 billion" (2026-06-08) sits on top of then-current frontier-lab run-rates: OpenAI around $20B ARR at the start of 2026, Anthropic's May Series H print at $47B run-rate. Baker's "well over 200 billion in inference revenue" (2026-06-11) is an end-of-2026 forecast of a broader inference bucket, and he was already treating the $300B 2027 consensus as too low. By mid-August 2026 the two-lab floor has moved: OpenAI disclosed a $40B ARR and Anthropic printed >$11.5B of preliminary Q2 revenue (May run-rate $47B). That already clears Jain's June snapshot without proving Baker's year-end $200B+.

The Colossus half of the dispute is not a contradiction. Jain's own transcript attributes Colossus to xAI, not OpenAI — the wiki's "OpenAI's Colossus" line is a mis-record. The 11% figure, as reported by The Information on 2026-05-02 and restated as model FLOP utilization (MFU) in later write-ups, is a training-efficiency ratio against theoretical peak FLOPs, not an occupancy rate on sold capacity. Baker's 55% is a financial return on Colossus 1 as a rented asset. Both can be true.

Resolver 4 is still open. Microsoft's FY2026 10-K discloses $24.1B of revenue from commercial arrangements with OpenAI and does not split inference from training or Azure compute from revenue-share. Alphabet's Q2 Cloud print ($24.8B, +82%) names "enterprise AI Solutions and enterprise AI Infrastructure" as drivers and does not split them. No fetched hyperscaler filing isolates inference billings.

Conviction implication for the stock-market wiki: record the definitional narrowing; do not move mechanism conviction. Jain's "economics are really bad" and Baker's "the math maths" remain talking-their-book readings of different numerators.

Findings

Theme 1 — "Revenue on AI" and "inference revenue" are not the same quantity

Grokipedia's Inference Economy page already shows the definitional scatter. It cites one forecast that "spending on inference-focused applications" reaches $20.6 billion by 2026 from $9.2 billion in 2025, and a different forecast that the "global AI inference sector" was $97.24 billion in 2024 and is projected to $253.75 billion by 2030. Those cannot both be measuring Baker's "inference revenue." The page also states that inference now accounts for "up to 90 percent of a model's total lifetime cost" and that LLM inference cost "has dropped by a factor of 1,000 over three years" — i.e. unit cost is falling while aggregate spend is rising, which is exactly the setting in which two honest observers can quote numbers an order of magnitude apart.

Fetched market-research pages confirm the scatter rather than resolving it:

  • Fortune Business Insights values the "global AI inference market" at $103.73B in 2025 and $117.80B in 2026, growing to $312.64B by 2034 (CAGR 12.98%). That 2026 print sits between Jain and Baker and is a hardware-plus-software industry TAM, not lab ARR.
  • Precedence Research values AI inference-as-a-service at $18.60B in 2025 and $23.40B in 2026. That is Jain-adjacent and much closer to Grokipedia's $20.6B "inference-focused applications" line.
  • Intel Market Research values the "AI inference platform" market at $8.1B in 2025 and $8.9B in 2026.

A Gartner 2026-08-10 press release at https://www.gartner.com/en/newsroom/press-releases/2026-08-10-gartner-forecasts-worldwide-artificial-intelligence-optimized-iaas-spending-to-grow-96-percent-in-2026 appeared in search with a table putting 2026 AI-optimized IaaS at $42.3B and "global spending on inference" at $23.3B vs training $19.0B. Fetch timed out; treat those Gartner figures as unverified until a successful fetch. The URL is recorded so a later pass can retrieve it.

What this does to the dispute. Jain said "the revenue on AI." Baker said "inference revenue." Neither defined the term on-air. The fetched 2026 market-size range for things called inference is roughly $9B–$118B depending on whether the analyst means IaaS-for-inference, inference-as-a-service, or a hardware-inclusive TAM. Baker's "well over $200B" is above every fetched 2026 market-research print and only becomes plausible if the numerator is lab + hyperscaler AI-product run-rate, not a syndicated "inference market" TAM. Jain's $70–80B is a plausible June-2026 lab snapshot (see Theme 2) and a poor match for a hardware-inclusive TAM.

Theme 2 — Disclosed 2026 lab run-rates: Jain's June snapshot is stale; Baker's year-end $200B+ is still a forecast

OpenAI. TechTimes (2026-08-15), citing CNBC/Bloomberg reporting of a closed-door 2026-08-14 investor meeting, says CFO Sarah Friar disclosed $40 billion ARR, a 20% month-over-month jump in July, with business-customer count up 32% in the same month. The same piece says OpenAI's ARR was "roughly $20 billion" at the start of 2026 — the figure Friar had confirmed in a January review — and that enterprise has now overtaken consumer ("We entered the year at 60-40… those lines have now crossed"). TechTimes flags the load-bearing caveat: ARR is a run-rate (recent month × 12), not GAAP. Audited FY2025, as retold there, was $13.07B booked revenue against a $20.92B operating loss, with $17.2B paid to Microsoft for Azure compute in 2025.

Anthropic. Value Add Pulse (2026-08-15), citing CNBC/Bloomberg, says preliminary Q2 2026 revenue exceeded $11.5B, vs $787M in Q2 2025 and $4.73B in Q1 2026, with first positive adjusted operating income. The same piece restates Anthropic's May disclosure that run-rate revenue crossed $47B, vs ~$10B for all of 2025. A $11.5B quarter annualizes to ~$46B — still below the May $47B run-rate, which the write-up reads as intra-quarter acceleration rather than overstatement. Figures are preliminary and non-GAAP on the profit line.

The June arithmetic. Jain spoke on 2026-06-08. At that date the public lab floor was roughly OpenAI ~$20–25B ARR plus Anthropic's late-May $47B run-rate ≈ $67–72B, plus xAI/others. That is Jain's "maybe 70, 80 billion" if he meant frontier-lab revenue, not hyperscaler AI cloud. Baker, three days later, was not quoting a then-current print; he was forecasting year-end 2026 "well over $200B" and calling a $300B 2027 consensus low (DruckFin recap of the 2026-06-11 BG2 episode).

The August arithmetic. OpenAI $40B ARR + Anthropic Q2 annualized ~$46B (or May $47B run-rate) ≈ $86B+ for two labs. That clears Jain's June snapshot and is still less than half of Baker's year-end $200B+ unless one adds Google/Microsoft/Amazon/xAI/Cursor and accepts double-counting (Theme 3). TechTimes' own gloss — "together the two companies are approaching $100 billion in combined annualized revenue" — is the honest mid-August two-lab floor. It does not get you to $200B.

Baker had already used a similar stack in May. On the 2026-05-22 All-In episode (already in the stock-market wiki, not re-fetched here) he said OpenAI and Anthropic at "call it $100 billion of ARR now" with "80% ish gross margins on inference," and that adding Gemini, Cursor, xAI, and open source made "200, 300, $400 billion of ARR the end of this year" easy to see. The June "well over $200B inference revenue" line is the same forecast, restated. It is not a disclosed segment number.

Theme 3 — Hyperscaler AI numbers are large, mixed, and not an inference split

Microsoft AI run-rate (primary). Microsoft's own 2026-04-29 press release quotes Satya Nadella: "Our AI business surpassed an annual revenue run rate of $37 billion, up 123% year-over-year" (Microsoft Source). This is the first time Microsoft put a hard number on "AI business." It is a run-rate, not a 10-K line, and the release does not define the mix (Azure AI consumption vs Copilot seats vs GitHub Copilot vs other). TechTimes' FY2026 Q4 recap repeats the $37B Q3 figure, adds Azure >$100B for the full fiscal year (+41%), Intelligent Cloud Q4 $39.31B, Copilot >30 million paid seats, and commercial RPO $678B (+84%). It also notes Q4 capex $35.8B vs FCF ~$19.6B — the first quarter in the AI buildout where quarterly capex exceeded FCF — and FY2026 capex ~$115.95B.

Microsoft ↔ OpenAI related-party revenue (primary). Microsoft's FY2026 Form 10-K (period ended 2026-06-30) states, under related-party disclosures: "For fiscal year 2026, we recorded revenue from commercial arrangements with OpenAI, inclusive of revenue-sharing payments, of $24.1 billion, and accounts receivable from OpenAI as of June 30, 2026 was $6.0 billion" (msft-20260630.htm). Neowin correctly flags that the filing does not break the $24.1B into Azure compute vs revenue-share vs other, so it cannot be read as Microsoft's share of OpenAI revenue or as Azure inference. Adding this $24.1B to OpenAI's $40B ARR double-counts: much of it is OpenAI's cost of compute, not a second pile of "AI revenue."

Alphabet / Google Cloud (secondary extract of a primary PDF). Alphabet's Q2 2026 Exhibit 99.1, as extracted from this hosted copy, reports Google Cloud revenue $24.768B vs $13.624B a year earlier (+82%), "led by an increase in Google Cloud Platform (GCP) across enterprise AI Solutions and enterprise AI Infrastructure, as well as core GCP services." Pichai: Cloud growth "driven by demand for AI infrastructure and AI solutions"; Gemini models "process 22 billion API tokens per minute"; Gemini App 950 million MAU; "nearly 90% of the Fortune 100" on Gemini Enterprise. The release defines Google Cloud as infrastructure, platform, applications, Workspace, "and other products," and says Cloud now also generates product revenue from TPU system sales. There is no inference-vs-training split and no AI-only Cloud dollar figure. A $24.8B quarter annualizes to ~$99B of all Google Cloud, not of inference.

Implication for Baker's $200B. A naïve stack — two labs ~$86B+ plus Microsoft AI $37B plus some slice of Google Cloud plus AWS plus xAI/Cursor — can be narrated past $200B. It is not a disclosed number, and it double-counts lab-to-hyperscaler compute spend. Jain's $70–80B does not include that stack. This is resolver 1.

Theme 4 — Colossus is xAI's; 11% is MFU; 55% is a financial return

Whose Colossus (resolver 2) — resolved from the original transcript. Jain's 2026-06-08 Capital Allocators remarks, already in vault/projects/stock-market/sources/2026-06-08-podcast-capital-allocators-contrarian-quality-at-gqg-partners-rajiv-jain-ep.md, read: "How in the world OpenAI invested trillion dollars when your revenue is maybe 20 billion x AI cash losses were give and take double digit billions 12 to 15 and their capacity utilization on the Colossus was 11% now, now they're selling the capacity to anthropic." "Their" attaches to xAI, not OpenAI. The stock-market entity page that says "OpenAI's Colossus supercluster" is a mis-record. Grokipedia's Colossus (supercomputer) page is unambiguous: Colossus is xAI's Memphis cluster, initially 100,000 H100s in 122 days, later mixed H100/H200/GB200, with a February 2026 expansion cited at ~555,000 GPUs / ~2 GW and a Southaven "MACROHARDRR" site. That matches the wiki's existing Midha/Baker usage.

What 11% is (resolver 3). Wccftech reprints The Information's 2026-05-02 item: "xAI's GPU fleet is running at about 11% utilization," ~550,000 H100/H200s, Meta ~43% / Google ~46% as comparison points, xAI targeting 50%. Wccftech's own gloss ("only able to utilize 11% of the 550,000 GPUs… equivalent of 60,000 GPUs") over-reads the metric as occupancy. NewsGlobeNow's 2026-05-05 write-up, citing the same Information memo (xAI president Michael Nicolls), is more careful: the figure is model FLOP utilization (MFU); "The 11% figure does not mean 89% of xAI's GPUs are completely idle"; MFU is observed training throughput over theoretical peak FLOPs; production LLM training often lands ~35–45% MFU. Nicolls reportedly called 11% "embarrassingly low" and set a 50% target. The same piece says xAI is renting unused compute, including to Cursor.

This is the same 11% the stock-market wiki already has from anjney-midha on 2026-06-13 Odd Lots: Colossus 2 at "<60% node utilization and <11% MFU." Jain on 2026-06-08 said "capacity utilization… 11%." Three speakers, one number, three labels. The Information/Midha framing (MFU) is the technically specified one. Jain collapsed it to "capacity utilization," which is how a quality-growth PM would hear a leak, not how a training engineer would write it.

What 55% is. Baker on BG2 (2026-06-11), already in the wiki: "Your colleague at Altimeter, Freda, also, she calculated a 55% ARR on Colossus 1." The DruckFin recap paraphrases the same line as a 55% IRR ("If you can borrow money at six, seven, eight percent and invest in something with a 55% IRR, that math maths"). ARR vs IRR is a real ambiguity in the secondary recap; the vault transcript says ARR. Either way it is a return on the Colossus 1 asset, not a utilization rate. Engineered.at's thin NextBigFuture-derived note (2026-05-22) attributes to Duan a ~$6B annual rent on Colossus 1 and a cited $15B/year Anthropic rental, with an implied ~$9B/year on Colossus 2. That page is a one-minute aggregator; treat the $6B/$15B figures as unverified third-hand. A tradersunion.com page that appeared to carry the $6B rent estimate was Cloudflare-blocked.

Compatibility. An asset can post a high financial return on the sold tranche (Colossus 1 rented to Anthropic/Google at the $22–23B and ~$50B per-gigawatt deal prices Baker/Fox cited on the same podcast) while the training stack on the remaining fleet prints 11% MFU. Selling capacity to Anthropic — which Jain himself notes — is the mechanism that turns idle training FLOPs into rent. Resolver 3 is narrowed to "both can be true"; it is not a 11% vs 55% shoot-out.

Theme 5 — No fetched hyperscaler filing splits inference billings from training (resolver 4, still open)

  • Microsoft 10-K $24.1B OpenAI commercial revenue: no compute / revenue-share / inference / training split (10-K; Neowin).
  • Microsoft "AI business" $37B ARR: no product-line split in the April 29 release.
  • Alphabet Q2 Cloud $24.8B: AI infrastructure + AI solutions + core GCP + first TPU-system product sales, no inference/training split (Exhibit 99.1 copy).
  • Gartner IaaS inference-vs-training split ($23.3B vs $19.0B) would be the closest published cut if the press release fetches; it did not.

Until a 10-K / 10-Q / earnings exhibit names inference billings as a number, resolver 4 stays open. The closest primary we have is Microsoft putting any dollar figure on "AI business" ($37B ARR) and on OpenAI-related revenue ($24.1B) — progress on existence of an AI P&L, not on the inference/training cut.

Contradictions and open questions

  • Definitional, not resolved to a single 2026 number. Jain $70–80B ≈ June lab snapshot. Baker >$200B = year-end forecast of a broader inference stack. Mid-August two-lab floor ~$86B–$100B. Syndicated "inference market" prints range $9B–$118B. No one fetched source puts a single official 2026 "inference revenue" on a 10-K.
  • Double-counting. Microsoft $24.1B from OpenAI is largely OpenAI's compute bill. Adding it to OpenAI ARR inflates "AI revenue." Baker's stack (labs + Gemini + Cursor + xAI + open source) has the same problem if hyperscaler AI ARR is folded in.
  • ARR vs booked revenue. OpenAI $40B ARR vs FY2025 booked $13.07B. Anthropic $47B May run-rate vs Q2 print $11.5B. Jain may have been closer to booked thinking; Baker is a run-rate thinker. That alone can move a number by 2×.
  • 11% label drift. Information/Midha: MFU. Jain: "capacity utilization." Wccftech: occupancy (over-read). The number is the same leak; the ontology is not.
  • 55% ARR vs IRR. Vault transcript says ARR; DruckFin recap says IRR. Freda's worksheet was not fetched. The $6B Colossus 1 rent / $15B Anthropic rental figures are third-hand.
  • Resolver 4 still open. No hyperscaler isolate of inference vs training billings.
  • Gartner IaaS table unverified. Fetch timed out.
  • X pass produced no citable permalinks. site:x.com searches returned zero results; the X MCP server in this environment is needsAuth. Wccftech embeds a 2026-05-02 @theinformation tweet without a status URL.

Provenance

Rounds run: 3 of 3

Sub-questions by round:

Round 1 (broad survey):

  1. What do 2026 sources mean by "revenue on AI" vs "inference revenue"?
  2. What lab run-rates (OpenAI, Anthropic, others) were disclosed after Jain/Baker spoke?
  3. What is the primary source and metric for Colossus 11%, and can it coexist with a 55% return on Colossus 1?
  4. Has any hyperscaler disclosed inference billings separately from training?

Round 2 (drill-down):

  1. Confirm Microsoft's OpenAI $24.1B and "AI business" $37B in primary filings/releases — targeting resolver 4 and the hyperscaler stack.
  2. Retrieve Freda/Altimeter Colossus 1 55% worksheet or a close secondary — targeting resolver 3.
  3. Map 2026 "inference market" TAM prints against lab ARR — targeting resolver 1.

Round 3 (resolve remaining uncertainty):

  1. Microsoft official $37B quote and Alphabet Cloud AI wording — targeting whether a hyperscaler number can be added to labs without double-counting.
  2. MFU-vs-occupancy clarification on the 11% leak — targeting resolver 3.

Anchor source (Grokipedia, fetched before round 1):

  • Inference Economy — 20,017 chars extracted — defines the inference-vs-training cost shift and already contains a $20.6B (2026 applications) vs ~$97B (2024 sector) definitional split. Bonus fetch, not counted against the 15-URL cap.
  • Narrower later fetch: Colossus (supercomputer) — confirms Colossus is xAI's Memphis cluster (counted in the URL budget). First slug attempt Inference_Economy 404'd; live slug is inference-economy.

X sources (--include-x enabled):

  • 0 items surfaced with a citable permalink (0 direct fetch, 0 snippet-with-URL).
  • WebSearch queries with site:x.com / site:twitter.com returned no results.
  • X MCP server status in this environment: needsAuth.
  • Secondary embed only: Wccftech quotes a 2026-05-02 @theinformation tweet ("xAI's GPU fleet is running at about 11% utilization") without a status URL — not treated as an X primary.

URLs fetched (15 successful, 4 failed):

Anchor / bonus:

Round 1:

Round 2:

Round 3:

Existing vault sources consulted (not fetched; already ingested):

  • 2026-06-08-podcast-capital-allocators-contrarian-quality-at-gqg-partners-rajiv-jain-ep — Jain $70–80B and Colossus 11% (xAI).
  • 2026-06-11-podcast-bg2-pod-the-spacex-ipo-fable-5-ai-capex-update-market — Baker $200B+ and 55% ARR.
  • 2026-06-13-podcast-odd-lots-anjney-midha-s-plan-to-radically-lower-the-price — Midha Colossus 2 <11% MFU.

Tools used: WebSearch, WebFetch, grokipedia-fetch (_lib/grokipedia.py), x-fetch attempted (WebSearch site:x.com + X MCP unavailable). Generated: 2026-08-17 13:20 UTC

Referenced by