Enterprise token budgeting
Enterprise token budgeting
One-line summary: After a 1H26 "tokenmaxxing" episode — Meta employees burning 60T tokens in 30 days, Uber exhausting an annual Claude Code budget in four months — enterprises have imposed per-employee token budgets, downgraded default models, and switched off premium tiers. The question this raises for every AI-infrastructure chain in this wiki is whether inference demand growth is a physical trend or a budget-cycle artifact. SemiAnalysis, having surveyed 50+ enterprises, argues it is the former.
The insight
- The flip is real and dated. From 2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s: "Companies are now shifting focus from tokenmaxxing to token budgeting." Budgets run "starting at $250 and going up to tens of thousands a month"; companies are "downgrading default models and turning off premium tiers."
- But the whales, who are the revenue, are not cutting. "90th+ percentile customers make up most of the revenue and are at very little risk to API revenue cuts." Ramp data: "99th percentile customers spend almost $90,000/yr per employee while 90th percentile customers spend ~$7,300… median Ramp customer spending just $136."
- The headline anecdotes are idiosyncratic, per the author. "the headline Uber, Meta, and other Fortune 500 tokenmaxxing stories were a result of poor incentives and lax oversight rather than an absence of high ROI activities." Meta "is only a 3-5% customer for Anthropic per our estimates."
- Conclusion, with a dated window: "there is not a material risk present to 2H26 AI budgets and we expect Anthropic and OpenAI's API business to continue to grow at their current net new rates m/m."
- The vertical-replication claim is the growth engine: "over 70% of ARR today across OpenAI and Anthropic can be attributed to coding use cases," and "What the coding market has done to AI Lab ARR will be repeated with Cyber… and again with white-collar knowledge work."
Why it matters to stock-market
- This is the demand-side load test for the entire AI-capex complex. nvidia-gpu-backstop-to-neocloud-financeability projects a >$7T AI debt market underwritten, ultimately, by token demand. If enterprise token spend is budget-capped, that debt is collateralized by an assumption. This concept is where that assumption gets checked.
- Named beneficiaries: "Our estimates for AWS Bedrock this quarter drive our total AWS growth rate number well above street" (AMZN — testable at the next AWS print); TaaS providers "Together, Fireworks, Baseten, and others who make up over $4B of ARR today" (private).
- A seat-vs-consumption read-through. Microsoft's M365 Copilot is described as gameable: "employees can game the system by using Copilot's 365 chat… before spending metered tokens on Claude or Codex." That is the same bolt-on-seat pattern flagged in agentic-ai-seat-erosion-to-saas-rerate.
Contradictions / tensions
- Same-publisher dependency — the most important caveat here. This source and nvidia-gpu-backstop-to-neocloud-financeability's source are both SemiAnalysis, four days apart. One argues the debt market is financeable; the other supplies the demand evidence that underwrites it. They are two halves of one house view, not independent corroboration. Any
confirmedgrade on enterprise-token-demand must come from a different publisher or a primary print (AWS/Azure segment disclosure). - The author is arguing against the tape. "our work suggests that headlines are overblown, enterprises continue to spend" is a contrarian claim, and the anecdotes it dismisses (Uber's $1,500/month/employee cap) are concrete while the rebuttal is survey-based.
- The conditional in the growth claim is easy to miss: the cyber vertical is "(Mythos re-release dependent)."
Evidence
- From 2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s (all quotes above).
- Corroborating the demand side from an operator: henry-he in 2026-06-29-podcast-odd-lots-baidu-s-cfo-on-how-it-became-a-full-stack-ai: "80% of the incremental demand today on... token inference related." Different geography, different publisher — the closest thing to an independent check currently in the wiki.
- The budgeting-discipline behavior, from a compute vendor's seat (independent publisher): andrew-feldman (Cerebras CEO) in 2026-07-10-podcast-all-in-podcast-open-source-wins-agi-is-here-and-scorsese-s-ai — the Costco analogy for enterprise maturation: "at first people opened up and said, everybody as much tokens as you want... And now we're jumping on and saying, whoa... these guys should have as much as they need, they're enormously productive. Over here, we can use maybe an open source model, maybe a cheaper model... And now we're sort of running it like a business." Corroborates the tokenmaxxing→budgeting flip and the tiered-model routing (premium for high-ROI users, open-source/cheap for the rest) that feeds open-source-share-shift-bullish-for-compute.
Open questions
- Does AWS's next print confirm the above-street Bedrock number? That is the single cheapest falsification test available.
- Is the 2021→2026 enterprise-spend s-curve genuinely early ("The media fortune 500 is well below $100 per employee still"), or is per-employee spend a poor denominator once agents replace employees?
Valuation snapshot
2026-08-28 mark (2026-08-27 closes,
twelvedata): AMZN $256.26 (-1.54%, -10.8% off 52w high) · MSFT $505.06 (+1.75%, -8.8% off 52w high)(Table below is the last full snapshot — market-cap / fwd-P/E columns are not refreshed by the daily price pass and are dated as shown.)
Last refreshed 2026-07-20 (pre-open; marks are the Friday 2026-07-17 close, markets closed over the weekend). Every price fill tagged twelvedata. Mkt cap / Fwd P/E are not in the Twelve Data free tier and were not re-sourced this run — neither is a name where the fundamental moves this concept's read.
| Ticker | Price | 52w range | Mkt cap | Fwd P/E | Day / vs 52w hi | What's priced in (one line) |
|---|---|---|---|---|---|---|
| AMZN | $247.23 | $196.00–$278.56 | — | — | −1.06% day; −11.2% from hi | The named, testable expression. SemiAnalysis: "Our estimates for AWS Bedrock this quarter drive our total AWS growth rate number well above street" (2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s). Priced: AWS at consensus. Not priced: an above-street Bedrock number — which is the single cheapest falsification test available for this whole concept |
| MSFT | $393.82 | $349.20–$555.45 | — | — | −1.82% day; −29.1% from hi | Green on a −2.2% XLK day, yet 27.8% off its high — the widest gap between day-strength and high-distance in the book. The seat-vs-consumption read is the relevant one here and it is unflattering: M365 Copilot is described as gameable — "employees can game the system by using Copilot's 365 chat… before spending metered tokens on Claude or Codex" — the same bolt-on-seat pattern flagged in agentic-ai-seat-erosion-to-saas-rerate |
Read-across from the pull: IGV $93.70 (−0.26%) — software held up far better than semis (SOXX −4.5%) on 2026-07-16. If the market were pricing a token-demand collapse, the application layer would not be the resilient one. Weak evidence, but it points against the budget-cliff read.
Forward-looking outcomes (12-month)
Bull case — token budgeting is a maturation, not a ceiling, and the whales keep paying: the concept's strongest fact is that the flip is real and irrelevant to revenue — "90th+ percentile customers make up most of the revenue and are at very little risk to API revenue cuts," with Ramp data showing "99th percentile customers spend almost $90,000/yr per employee while 90th percentile customers spend ~$7,300… median Ramp customer spending just $136." The headline anecdotes are idiosyncratic on the author's own account — "the headline Uber, Meta, and other Fortune 500 tokenmaxxing stories were a result of poor incentives and lax oversight rather than an absence of high ROI activities," and Meta "is only a 3-5% customer for Anthropic per our estimates." The dated conclusion: "there is not a material risk present to 2H26 AI budgets." Then the growth engine: "over 70% of ARR today across OpenAI and Anthropic can be attributed to coding use cases," and "What the coding market has done to AI Lab ARR will be repeated with Cyber… and again with white-collar knowledge work." Implied price: AMZN +15–25% on an above-street AWS print. Cited: 2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s.
Base case — demand grows, discipline sticks, and the mix shifts down-market: budgeting is real and permanent, and it routes work to cheaper models rather than eliminating it. andrew-feldman in 2026-07-10-podcast-all-in-podcast-open-source-wins-agi-is-here-and-scorsese-s-ai — the Costco analogy, from an independent publisher and a compute vendor's seat: "at first people opened up and said, everybody as much tokens as you want... And now we're jumping on and saying, whoa... these guys should have as much as they need, they're enormously productive. Over here, we can use maybe an open source model, maybe a cheaper model... And now we're sort of running it like a business." Total tokens keep growing; revenue per token falls. That is bullish compute and ambiguous for the labs — see open-source-share-shift-bullish-for-compute. Implied price: AMZN +8–15%. Cited: 2026-07-10-podcast-all-in-podcast-open-source-wins-agi-is-here-and-scorsese-s-ai, 2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s.
Bear case — budget caps bite, and the debt stack underneath is collateralized by an assumption: the concrete anecdotes are concrete (Uber's $1,500/month/employee cap) while the rebuttal is survey-based, and the author is explicitly "arguing against the tape." If enterprise token spend is genuinely budget-capped, then the >$7T AI debt market projected by nvidia-gpu-backstop-to-neocloud-financeability is "collateralized by an assumption" — this concept is where that assumption gets checked, and a failure here propagates to the whole AI-capex complex, not just to AMZN. Implied price: AMZN −15–20%; the systemic read is worse than the single-name one. Cited: 2026-06-30-feed-semianalysis-tokenbudgeting-our-conversations-with-enterprises-on-token-s.
Currently undervalued vs base case? AMZN: Marginal — and the constraint on saying more is source independence, not price. At −10.3% from its high, AMZN is not obviously mispriced, and the case for it rests almost entirely on one publisher's above-street Bedrock estimate. The disqualifying caveat is this page's own, and it is the most important thing on it:
"Same-publisher dependency — the most important caveat here. This source and nvidia-gpu-backstop-to-neocloud-financeability's source are both SemiAnalysis, four days apart. One argues the debt market is financeable; the other supplies the demand evidence that underwrites it. They are two halves of one house view, not independent corroboration. Any
confirmedgrade on enterprise-token-demand must come from a different publisher or a primary print (AWS/Azure segment disclosure)."
A circular evidence base cannot support a position on either page. The one genuinely independent check currently in the wiki is thin but real — henry-he in 2026-06-29-podcast-odd-lots-baidu-s-cfo-on-how-it-became-a-full-stack-ai: "80% of the incremental demand today on... token inference related" — different geography, different publisher, and directionally supportive. andrew-feldman adds a second independent voice on the behaviour (the flip is real) though not on the magnitude.
MSFT: No. −27.8% from its high looks like the compression a buyer wants, but this page's evidence cuts against MSFT rather than for it: the M365 Copilot seat is described as the thing enterprises game to avoid metered spend. That is the seat-erosion pattern in agentic-ai-seat-erosion-to-saas-rerate, and it makes MSFT the funding source of the token economy as much as a beneficiary of it.
Catalyst path:
- AMZN Q2 2026 earnings (late July/early August) — "the single cheapest falsification test available." Does AWS print above street on Bedrock strength? One number resolves the central claim.
- MSFT Q4 FY2026 (late July) — Azure AI segment disclosure; the second primary print that could break the same-publisher circularity.
- Any non-SemiAnalysis survey of enterprise token spend — the evidence gap that currently caps conviction at low-medium regardless of what the prints say.
Related
- nvidia-gpu-backstop-to-neocloud-financeability — the debt stack this demand underwrites.
- inference-taas-mix-to-aws-margin-expansion — the AWS margin chain this feeds.
- agentic-ai-seat-erosion-to-saas-rerate — the seat-vs-consumption sibling.
- compute-utilization-overhang-as-latent-supply — the bear case on inference demand.
Evidence — the measured shape of an agentic workload (added 2026-08-12)
From 2026-08-03-feed-semianalysis-kimi-k3-the-manos-the-mythos (⚠ paywalled partial), benchmarking on replayed agentic traces: "a median of 142k input tokens and a median of 444 output tokens per turn with a median of 65 turns per session. The short output tokens per turn is typical for workloads on agentic harnesses, where the agent calls tools frequently, even edits are tool uses."
A ~320:1 input-to-output ratio per turn, across 65 turns, is a fundamentally different cost object from the chat-shaped workloads most token-budget reasoning assumes: the spend is dominated by prefill and KV-cache residency, not decode. That is why prefix caching is the load-bearing economic lever — and why cache thrash above concurrency 8 (see hbm-cowos-as-binding-bottleneck) translates directly into a serving-cost cliff rather than a gradual degradation. Also dated: an OpenRouter price floor of $3/M input, $15/M output for the model as of 2026-07-30.
Update (2026-09-16) — allocator offices are metering tokens. Independent of SemiAnalysis, not a NOW re-rate.
From 2026-09-14-podcast-capital-allocators-ai-in-the-investment-office-abby-barlow-laura (Matt Bank, University of Virginia): "We've had to meter some of our token usage because of things that were way out of bounds with respect to what the benefits were." Agentic end-to-end workflows "haven't really seen that in a cost effective way work." That is a non-SemiAnalysis instance of the tokenmaxxing → budgeting flip, from an LP office rather than a hyperscaler. It does not size 2H26 lab ARR. Do not re-date NOW. Do not mint a seat-SaaS chain from allocator workflow chatter.
Update (2026-09-17) — operator token-burn: 4 billion tokens / ~$1,300 in a day. Independent of SemiAnalysis. Do not re-date NOW.
matt-barrie in 2026-09-17-podcast-macro-voices-macrovoices-549-matt-barrie-ai-gent-provocateur (Freelancer.com CEO; recorded ~10 September): a ~40-agent fleet burned "4 billion tokens in one day" on OpenRouter, "about 1300 US dollars," of which "$900 of it was Sonnet." He then asks how to cut cost: DeepSeek V4 Flash on a pair of DGXs at ~45–60 tokens/second vs that day's ~46,000 tokens/second on frontier APIs. "the amount of tokens that on a per person basis or per enterprise basis that are going to be consumed are going to go through the absolute roof" — "I'm not going to pay the prices that what the frontier models want to charge."
That is a second non-SemiAnalysis operator instance of the same flip: spend spikes, then routing to cheaper / local models. It corroborates the shape (tokenmaxxing → budgeting / mix-shift) and does not size 2H26 lab ARR or NOW seats. Do not re-date NOW. Do not mint a net-new AI-infra chain from one CEO's OpenRouter bill.