Autoresearch: Independent evidence that inference is the bulk of 2026 AI compute spend and that speed is the priced attribute
Non-Feldman check of Step 1 of inference-demand-to-wafer-scale-advantage: whether 2026 AI compute spend is inference-majority, and whether inference speed (not just tokens/$ or concurrency/MW) is the priced attribute. Verdict: keep-partial.
Autoresearch: Independent evidence that inference is the bulk of 2026 AI compute spend and that speed is the priced attribute
Generated by
/autoresearchon 2026-09-19. Synthesized across 3 rounds from 12 successful web pages (3 failed), anchored by a Grokipedia miss (query "AI inference" resolved to VRAM hardware, not spend mix). See Provenance. Treat as raw material — review before promoting into a project or thread. Context: vault/projects/stock-market Scope: research only. Do not write vault. No buy / sell / size. Feldman / Cerebras-CEO quotes are not independent and are not used as evidence here. Confirm Step 1 ofinference-demand-to-wafer-scale-advantageonly if ≥2 independent sources or one authoritative primary support both halves: (1) inference — not training — is the bulk of 2026 AI compute spend, and (2) inference speed (not just tokens/$ or concurrency/MW) is the priced attribute.
Summary
Independent (non-Cerebras) sources as of 2026-09-19 do not support confirming both halves of Step 1. On spend share, the best independent numbers still show training/R&D as equal or larger than inference at the lab level: Epoch puts OpenAI 2024 at ~$1.8B inference vs ~$5B R&D compute, OpenAI 2025 at a roughly even R&D/inference split, and Anthropic 2025 at $2.7B inference vs $4.1B training. Epoch’s allocation theory expects comparable training and inference spend, not inference dominating by an order of magnitude. The Information (via THE DECODER) says OpenAI inference quadrupled in 2025 and that 2026 training is planned at $32B, inside a ~$50B total compute envelope (Brockman / Epoch) — if those figures are commensurate, training is still a large or majority slice, not a leftover. Hyperscaler IR still does not print a training-vs-inference split. A third-hand “SemiAnalysis 5–7× inference vs training” line was not found in the fetched SemiAnalysis GTC 2026 piece. On speed as the priced attribute, two lab primaries independently sell a speed premium: OpenAI Fast mode (latency SLA in tokens/second; billed at a premium to Standard) and Anthropic Fast mode (same weights, up to 2.5× output tok/s, 2× list price vs standard Opus 5). Google Priority is a 75–100% premium for latency plus reliability, not speed alone. SemiAnalysis simultaneously says SRAM/LPU speed “can demand a large market premium” and that GPUs still win tokens-at-scale / tokens per dollar. Recommendation: keep-partial. Half 1 fails the bar. Half 2 is real as a priced attribute, not as the exclusive priced attribute.
Findings
Half 1 — Inference as the bulk of 2026 AI compute spend: not independently confirmed
Lab-level dollar splits that can be cited without Feldman cut against “inference is the bulk.” Epoch’s reconstruction of OpenAI’s 2024 cloud bill — from The Information / New York Times investor documents — is ~$5B R&D compute (training + research) and ~$1.8B inference (Epoch, OpenAI 2024 compute). That is inference at roughly 26–30% of OpenAI compute, not the majority. Epoch is explicit that these are cloud opex, not hyperscaler capex, and that Microsoft’s own Copilot/Azure inference is outside the $1.8B (Epoch, OpenAI 2024 compute).
The 2025 update does not flip the claim to inference-majority. Epoch’s AI Chip Users explorer (2026-09-09) says OpenAI’s 2025 compute was “divided roughly evenly between R&D and inference, while a higher share went to R&D in 2024” (Epoch, AI Chip Users explorer). Even at parity, inference is not “the bulk.” The same note says OpenAI has committed to spend approximately $50B on compute in 2026, “roughly triple its 2025 compute budget,” without printing a 2026 inference share (Epoch, AI Chip Users explorer).
Anthropic 2025, reconstructed by Epoch from The Information, is $4.1B training / R&D compute vs $2.7B inference (implied from 40% gross margin on $4.5B revenue), on $9.7B total expenses (Epoch, company spending breakdown). Inference is ~40% of Anthropic compute, not the bulk. MiniMax (2025 Q1–Q3) and Z.ai (2025 H1) IPO filings, which Epoch tabulates as primary financials, are even more R&D-heavy: MiniMax $142M R&D compute vs $38M inference; Z.ai $183M vs $9M (Epoch, company spending breakdown).
Epoch’s GPT-5 economics note (updated 2026-03-06) estimates OpenAI 2025 inference at $4B (paid + free) against $15B R&D for the year — again R&D larger than inference, even after Epoch raised the inference estimate (Epoch, Can AI companies become profitable?). That piece also reports The Information’s figure that OpenAI spent “around $4.5bn in serving paid users in 2025 of over $8bn spent in inference,” implying a large free-user inference bill — growth in inference cost, not a documented majority of 2026 industry spend (Epoch, Can AI companies become profitable?).
2026 OpenAI numbers still do not show inference as the bulk. THE DECODER’s recap of The Information’s internal-forecast leak (2026-02-21) says a “major driver” of the cash-burn revision is inference — day-to-day running costs quadrupled in 2025 — and that OpenAI “plans to spend $32 billion on model training in 2026 and around $65 billion in 2027,” with training “alone” projected near $440B through 2030 inside $665B “training and operating” (THE DECODER / The Information). If Epoch’s ~$50B 2026 compute envelope and The Information’s $32B 2026 training line are even roughly the same object, training is still a majority or large plurality of that envelope, not a residual. They may not be the same object (compute opex vs cash training vs Stargate rent). That ambiguity is exactly why this cannot confirm “bulk of 2026 spend.”
Greg Brockman’s courtroom line that OpenAI expects to spend $50B on computing in 2026 is a total envelope for “develop[ing] more advanced AI models and serv[ing] them,” not a split (Bloomberg, 2026-05-05; cited via search snippet only — Bloomberg HTML not fetched).
Independent theory expects parity, not inference dominance. Epoch (Erdil, 2024-03-29) argues that if labs can trade training compute for inference compute along the observed ~1 OOM : 1 OOM frontier, cost-minimizing labs should spend comparable amounts on each, and “we should not expect one of these categories of expenditure to dominate the other by an order of magnitude or more in the future” (Epoch, Optimally allocating compute). The 2023 companion paper does say “inference is the dominant cost for models deployed at scale” over a model’s lifetime (Epoch, Trading off compute in training and inference) — a lifetime-TCO claim for a deployed model, not a 2026 industry-wide spend census, and not a claim that 2026 capex is inference-majority.
SemiAnalysis, fetched, does not supply the 5–7× industry ratio. SemiAnalysis’s GTC 2026 recap (“The Inference Kingdom Expands,” 2026-03-24) is about Nvidia productizing Groq LPUs for high-interactivity decode, not about industry FLOP-hour or dollar shares. It does not, in the fetched text, state that inference consumes 5–7× training compute (SemiAnalysis, GTC 2026). A Nonce Media essay attributes “roughly five to seven times as much aggregate compute as training” to “SemiAnalysis … GTC 2026 coverage” (Nonce Media); that attribution was not reproduced in the fetched SemiAnalysis piece and is treated as unverified third-hand. Do not promote it as SemiAnalysis primary.
What SemiAnalysis does say, independently of Feldman: inference demand accelerated in late 2025 (open-weight adoption + “surge in inference demand”); Blackwell lead times into mid-2026; “all capacity coming online until August to September 2026 has already been booked”; and training and inference coexist as distinct demand — large-MoE inference wants GB300 NVL72, while “training workloads can have the best price performance on H100s” (SemiAnalysis, GPU shortage / rental index). That is a demand-growth claim, not a 2026 spend-majority claim.
Hyperscaler IR still does not print the split. A 2026 filing-scorecard recap that walked Microsoft Cloud / AWS / Google Cloud 10-Q language concludes: Azure’s closest wording is “cloud and AI consumption-based services”; AWS and Google Cloud “do not split training clusters from token serving”; “Until a company prints training versus inference, do not pretend the 10-Q did” (Stock Alarm Pro, 2026 filing scorecard). That is secondary commentary, but it matches the absence of a split in the lab/IR pages fetched this pass.
Nvidia’s most recent primary (Q2 FY2027, 2026-08-26) names both sides and gives no mix. Huang describes a four-phase cycle — data prep, pre-training, post-training, agentic inference — and sells one “fungible” NVL72 system across all of them. He says frontier labs have “extraordinary demand for training and inference compute.” On Groq LPX he is explicit that it is for “high-interactivity services” and that “the vast majority of data centers will use Vera Rubin NVLink72” (Yahoo recap of the call; PDF transcript not fetched — off *.gov whitelist) (Yahoo Finance, NVDA Q2 FY2027 highlights). That is consistent with inference mattering; it is not a statement that inference is the bulk of 2026 spend.
Bottom of half 1: Independent sources support “inference demand grew hard in 2025–26” and “inference is a large, rising opex line.” They do not support “inference is the bulk of 2026 AI compute spend.” The quantitative record is parity or training/R&D-heavy at the labs that disclose, no IR split at the hyperscalers, and a theoretical prior of comparability. Criterion for confirm is not met.
Half 2 — Speed as the priced attribute: a priced attribute, not the exclusive one
OpenAI Fast mode is an authoritative primary that speed is sold. Fast mode (renamed from Priority processing on 2026-07-30) is “priced at a premium relative to Standard processing rates,” billed per token, and for Enterprise carries a latency SLA in tokens per second — e.g. GPT-5.6 Sol 99% > 80 tok/s, Terra >70, Luna >100, with Fast Sol list prices of $8 / $40 per 1M short-context input/output (OpenAI, Fast mode). Official copy: “Predictably low latency: Fast mode generates tokens faster and at a more consistent speed than the Standard processing service, even during peak demand” (OpenAI, Fast mode). OpenAI also tells customers not to put “large data processing or asynchronous jobs on Fast mode” because those “often do not need the improved performance” (OpenAI, Fast mode) — i.e. the SKU exists because interactive speed is separately valuable. Official Standard list prices for the same models were not retrieved this pass (openai.com/api/pricing 403; developers.openai.com / platform.openai.com pricing timed out), so the exact Fast/Standard multiple is not independently recomputed here. The existence of a speed SLA and a premium tier is.
Anthropic Fast mode is a second, independent primary that prices output tokens/second. Official docs (research preview, fast-mode-2026-02-01): “Fast mode delivers up to 2.5x higher output tokens per second from Claude Opus 5 and Claude Opus 4.8 at premium pricing.” Same weights; the gain is OTPS, not TTFT (Anthropic, Fast mode). List: Fast $10 / $50 per MTok in/out vs standard Opus 5 / 4.8 $5 / $25 (Anthropic, Fast mode; Anthropic, Pricing) — a clean 2× premium for ~2.5× output speed. Fast mode is not available with Priority Tier, Batch, or Bedrock/Foundry (Anthropic, Fast mode). That separation matters: Anthropic also sells Priority Tier as capacity / uptime (99.5% target; “minimize server overloaded errors”), and new Priority commitments are no longer for sale (Anthropic, Service tiers). Capacity and speed are different SKUs.
Google Priority is a third lab primary, but it prices a bundle, not speed alone. Gemini Priority is “a premium inference tier … for business-critical workloads that require lower latency and the highest reliability at a premium price point,” priced 75–100% more than Standard, with Flex/Batch at a 50% discount for minutes-to-hours latency (Google, Priority inference; page last updated 2026-09-02). The comparison table’s latency column is coarse (Priority “Seconds” vs Standard “Seconds to minutes” vs Flex “1–15 min” vs Batch “up to 24 hours”) (Google, Priority inference). That is independent evidence that latency is priced. It is also evidence that reliability, shedability, and tokens/$ are priced on the same card. Google is not selling “tok/s vs concurrency/MW” as a single axis.
SemiAnalysis prices both speed and tokens/$. On Groq/LPU (independent of Feldman; this is Nvidia’s licensed SRAM decode chip): “the standalone Groq LPU system is not economical for serving tokens at scale, but it can serve tokens very quickly which can demand a large market premium” (SemiAnalysis, GTC 2026). SRAM wins “very fast time to first token and tokens per second per user” at the expense of batch throughput; “GPUs win for throughput and cost” (SemiAnalysis, GTC 2026). InferenceX v2 reports Blackwell vs Hopper improvements as tokens per dollar (e.g. “9.7× … up to 65× … improvement in tokens per dollar compared to Hopper” at stated tok/s/user points) (SemiAnalysis, InferenceX v2). The Tokenomics model itself forecasts hardware demand as “inference demand from adoption, training demand from architecture development,” and unit costs as “training and inference cost per token” (SemiAnalysis Tokenomics; SemiAnalysis Models). That is the opposite of “speed, not tokens/$, is the priced attribute.” It is “speed has a premium and the volume market is tokens/$ / concurrency.”
Nvidia’s own framing matches SemiAnalysis, not an exclusive speed market. Huang: Groq 3 LPX is for “high-interactivity services,” but “the vast majority of data centers will use Vera Rubin NVLink72” (Yahoo Finance, NVDA Q2 FY2027 highlights). That is a primary (CEO, earnings Q&A, via recap) that interactive speed is a niche SKU, not the industry-priced attribute.
Bottom of half 2: ≥2 independent primaries (OpenAI Fast, Anthropic Fast) show that interactive tokens-per-second is a priced SKU at a ~2× list premium. Google Priority and SemiAnalysis LPU-premium corroborate that latency has a market. The same sources — SemiAnalysis tokens/$, Google Flex/Batch discounts, Anthropic Batch + retired Priority capacity tier, Nvidia “vast majority NVL72” — show that tokens/$, throughput, and concurrency remain co-equal priced attributes. The Step 1 wording is “rewards speed” / “speed is the priced attribute,” not “speed is one SKU among several.” Exclusive reading: not met. Narrow reading (“speed is a priced attribute”): met, with two lab primaries.
Both halves together (the confirm test)
The gate was: do not confirm Step 1 unless ≥2 independent sources or one authoritative primary support both halves.
| Half | Independent support? | Notes |
|---|---|---|
| Inference is the bulk of 2026 AI compute spend | No | Epoch + IPO filings + 2025 even-split + $32B 2026 training line all cut against or fail to show majority. No hyperscaler IR split. SemiAnalysis 5–7× unverified. |
| Speed (not just tokens/$ / concurrency/MW) is the priced attribute | Partial | Two lab primaries price tok/s. Same market still prices tokens/$, batch discounts, and capacity. SemiAnalysis: speed premium is real and niche vs GPU TCO. |
No single authoritative primary supports both. The two strongest independent clusters (Epoch/The Information on spend; OpenAI/Anthropic docs on speed) support different halves, and the spend cluster contradicts half 1. Keep Step 1 partial.
Contradictions and open questions
- Lifetime-TCO vs annual-spend. Epoch 2023: inference often dominates a deployed model’s lifetime cost (Epoch, Trading off compute). Epoch 2024–26 lab accounts: annual R&D/training still ≥ inference. Both can be true. Step 1 is about near-term spend, so the annual figures govern.
- $50B compute vs $32B training (2026). Epoch/Brockman total vs The Information training line may not share a definition (opex vs cash vs Stargate). Unresolved; do not subtract to invent an $18B inference residual.
- Microsoft / Azure inference is missing from OpenAI’s bill. Epoch flags this explicitly (Epoch, OpenAI 2024 compute). A hyperscaler-inclusive 2026 census could look more inference-heavy. No IR print exists.
- Nonce Media 5–7× vs fetched SemiAnalysis GTC 2026. Contradiction of attribution. Treat as unverified until the paywalled remainder of the SemiAnalysis piece (or another SA note) is fetched and quoted.
- Etched / SRAM-commoditizes-speed. Not independently re-validated this pass. SemiAnalysis’s LPU analysis is the closest independent analog: speed premium exists and is not the scale market (SemiAnalysis, GTC 2026). That leans toward the rival-ASIC tension already on the mechanism (contest moves to concurrency/MW / tokens/$), not toward confirming exclusive speed-pricing.
- OpenAI Fast vs Standard multiple. Fast prices and tok/s SLAs are official; Standard list for GPT-5.6 Sol was not fetched (403/timeout). Do not cite a 2× OpenAI multiple from secondary blogs.
- Industry-wide (labs + hyperscalers + neoclouds + China labs) 2026 mix. Still unpublished in any fetched primary.
Disposition
Keep-partial. Do not flip Step 1 of inference-demand-to-wafer-scale-advantage to confirmed. Independent evidence supports a 2025–26 inference-demand surge and the existence of priced speed SKUs. It does not support “inference is the bulk of 2026 AI compute spend,” and it does not support “speed, not tokens/$ or concurrency/MW, is the priced attribute.” No buy / sell / size. No new AI-infra chain.
Provenance
Rounds run: 3 (full)
Sub-questions by round:
Round 1 (broad survey):
- What do independent analysts (SemiAnalysis, Epoch, Dell’Oro) estimate for the 2025–26 training vs inference share of AI compute spend?
- Do OpenAI, Anthropic, Google, or Microsoft disclose an inference-vs-training spend, GPU-hour, or capex split?
- Is inference speed (latency / interactive tokens-per-second) the priced attribute in 2026 markets, versus tokens/$ or concurrency/MW?
- What do Epoch or academic sources say about inference compute overtaking training?
Round 2 (drill-down):
- Did 2025–26 lab compute mix flip to inference-majority, or do Epoch / The Information still show parity or training-heavy? — targeting the 2024-stale-split gap
- Do Google and Anthropic independently sell a speed/latency premium, or is OpenAI Fast mode unique? — targeting whether speed is the priced attribute industry-wide
- Does any 2026 OpenAI cash-burn readout print training vs inference dollars? — targeting the missing 2026 lab split
Round 3 (resolve remaining uncertainty):
- Does Anthropic’s official Fast mode page price output tokens/second, not just Priority capacity? — targeting whether speed is independently a priced SKU
- Can OpenAI’s Fast premium be measured against Standard list prices? — targeting exclusivity of speed vs tokens/$
- Is Brockman’s $50B 2026 compute envelope split in any primary or near-primary account? — targeting the 2026 spend-share hole
Anchor source (Grokipedia, fetched before round 1):
- Command as specified:
python3 .claude/skills/_lib/grokipedia.py fetch "AI inference" --max-chars 3000 - Resolved to VRAM in AI inference — 3,017 chars extracted (93 citations). Wrong topic: local GPU memory for running models, not 2026 spend mix or API pricing. Not used as a load-bearing source. A second helper query (
AI inference compute) hit the same slug. Large language model was fetched as a vocabulary check (1,517 chars) and does not address spend share. Encyclopedia was not used as the sole source.
X sources: --include-x not passed; skipped.
URLs fetched (12 successful, 3 failed):
Round 1:
- Most of OpenAI’s 2024 compute went to experiments (Epoch) — academic / data insight — $1.8B inference vs $5B R&D (2024)
- Compute accounts for the majority of expenses of AI companies (Epoch) — academic / data insight — Anthropic 2025 $4.1B training vs $2.7B inference; MiniMax / Z.ai IPO splits
- Fast mode for API Customers (OpenAI) — official primary — speed SKU, tok/s SLA, premium vs Standard
- Optimally allocating compute between inference and training (Epoch) — academic — theoretical parity, not OOM dominance
- The Great GPU Shortage – Rental Capacity (SemiAnalysis) — specialist analyst — 2025–26 inference-demand surge; no spend-share %
- GTC 2026 – The Inference Kingdom Expands (SemiAnalysis) — specialist analyst — LPU speed premium + GPU wins tokens-at-scale; no 5–7× ratio in fetched text
- Trading off compute in training and inference (Epoch) — academic — lifetime inference-dominant TCO for deployed models (2023)
Round 2:
- Introducing the AI Chip Users Explorer (Epoch, 2026-09-09) — academic — OpenAI 2025 roughly even R&D/inference; $50B 2026 compute commitment
- Can AI companies become profitable? (Epoch) — academic — 2025 inference ~$4B (or >$8B per The Information note) vs $15B R&D
- OpenAI adds $111 billion to its cash burn forecast (THE DECODER / The Information) — news recap of leak — inference quadrupled 2025; $32B training in 2026
- Priority inference (Gemini API) — official primary — 75–100% premium for latency + reliability; Flex/Batch 50% off
- Service tiers (Claude API) — official primary — Priority = capacity/uptime, not tok/s; new commitments closed
Round 3:
- Fast mode (Claude Platform Docs) — official primary — 2.5× OTPS at $10/$50 vs standard $5/$25
- Pricing (Claude Platform Docs) — official primary — standard vs Fast list tables
[Failed: https://openai.com/api/pricing/]— 403[Failed: https://developers.openai.com/api/docs/pricing]— timeout[Failed: https://platform.openai.com/docs/guides/priority-processing]— timeout
Search-dumped, cited with qualification (not counted as clean WebFetch primaries for the 5–7× claim):
- InferenceX v2 (SemiAnalysis) — tokens-per-dollar bench
- AI Value Capture (SemiAnalysis) — inference TCO vs rental $/FLOP
- Nonce Media 2026 infra essay — unverified 5–7× attribution
- AI CapEx, Inference, and Who Gets Paid (Stock Alarm Pro) — 10-Q “no split” recap
- NVDA Q2 FY2027 highlights (Yahoo) — Huang LPX vs “vast majority NVL72”
- Bloomberg / Brockman $50B — snippet only; paywall
Not used as independent evidence (known, excluded): Andrew Feldman Odd Lots / All-In / CBRS quotes; Etched (Rob Wachen) SRAM-commoditizes-speed claim (interested rival; not re-fetched as confirmation).
Tools used: WebSearch, WebFetch, _lib/grokipedia.py fetch (as specified). No X pass. Consensus MCP not used (web/IR/Epoch path). SOURCE_RELIABILITY consulted: preferred Epoch, official lab docs, SemiAnalysis newsletter, Microsoft/OpenAI-class official pages already marked Reliable; skipped Reuters (hard-blocked) and Bloomberg HTML (paywall); skipped NVDA Q4CDN PDF (off *.gov whitelist).
Generated: 2026-09-19 ~20:50 UTC