Open vs closed source model economics
Open vs closed source model economics
One-line summary: Open-weight models are ~3–5% behind the best closed models on quality but dramatically cheaper per unit of intelligence — because the user pays only for serving (power + compute), not for training. The "battle" between the two is, per andrew-feldman (May 2026), undecided; closed is "strictly better by a little bit," and the durable question is how big a premium that little bit commands.
The insight
The open-vs-closed split in AI is not the open-vs-closed split in traditional software, and conflating them is the core error this concept exists to flag.
- "Free open source" doesn't transfer. joe-weisenthal's framing: in traditional software, open source is free; in AI there is "no real such thing as free open source AI software" because even a free-to-license model still costs chip depreciation + electricity to run. The relevant cost axis is serving cost, not license cost.
- The quality gap is small; the cost gap is large. andrew-feldman (running an inference cloud that serves both) puts the closed-vs-open quality difference at "3, 4%, 5%" with closed "strictly better." But on cost per unit of intelligence, open is cheaper "by a lot" — because the user "what you're not paying for was the cost to train it." A ~1T-parameter open model (Kimi K2) runs on Cerebras "10 or 15 times faster than others" at just power + compute cost.
- A levelized-cost-of-intelligence metric is missing. Joe proposes "cost per IQ point" / "levelized cost of intelligence" as the unit that would let buyers compare honestly — and notes the industry doesn't have it yet. Without it, the premium for closed quality is hard to price.
- The market is bifurcating, not consolidating. Feldman expects no single winner — "I don't think there's going to be one," analogizing to x86 (Intel/AMD) + ARM + custom silicon coexisting. Closed frontier labs (OpenAI, Anthropic), open-model serving (Cursor, Cognition on open weights), and specialists all persist.
- Quiet enterprise migration to open is already happening. tracy-alloway reports "a lot of big companies in the US ... very quietly shifting from some of the closed source models to the open source models like the Chinese ones, like Kimi" and Qwen.
This dovetails with the thread's llm-as-commodity-thesis (Ghodsi: models are interchangeable at the unit level; durable value is above the model layer) and with cuda-moat-erosion-at-inference (runtime portability removes lock-in at the inference layer). Where the commodity thesis says "models commoditize," this concept adds the open-vs-closed pricing structure underneath that commoditization: open weights commoditize fastest because their cost floor is just serving.
Evidence
-
andrew-feldman in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "you can jump up right now and run Kimi K2. It's a 1 trillion parameter model. It's an open source model on cerebras where 10 or 15 times faster than others. And what you're paying for is the cost of our power and some cost of the compute that took to calculate it. What you're not paying for was the cost to train it."
-
andrew-feldman in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "The open source models, there are no open source models that are as good as the closed source models. Think of it as 3, 4%, 5% different... What is clear is that the closed source is strictly better by a little bit, by how much varies and it's more expensive."
-
andrew-feldman in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "You have OpenAI with their coding software, you have Anthropic with their coding software. And you've got companies like Cursor and Cognition that are using open source. We power OpenAI and we power Cognition. You have a battle underway between closed source and open source. And I think that the winners of that battle is yet to be determined."
-
joe-weisenthal in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "there's no really such thing as like free AI software. Because even if it's like free, you still have to pay for the depreciation of the chips and you have to pay for the electricity to run them. So there is no real such thing as like free open source AI software."
-
joe-weisenthal in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "are the open source models cheaper on a per unit of intelligence basis? If we had some way of saying levelized cost of intelligence, which I don't know if the industry has yet, are open source models cheaper per IQ point?"
-
tracy-alloway in 2026-05-21-odd-lots-why-cerebras-ceo-andrew-feldman-built-the-world-s: "I've heard of a lot of big companies in the US who have been very quietly shifting from some of the closed source models to the open source models like the Chinese ones, like Kimi... And Qwen."
-
June 2026 update — the 3-5% gap closed on some benchmarks. alexander-wissner-gross in 2026-06-26-podcast-moonshots-the-10b-satellite-empire-putting-ai-in-orbit-why: "the gestalt with GLM 5.2 is that it takes roughly double the number of tokens to get to the same capability output as the best western frontier models, but at half the total price. So the Chinese are evidently figuring out how to more efficiently or at least more cheaply reason." This directly answers this page's own open question ("does the closed-quality premium widen or collapse") in the collapse direction — but by a cost route, not a capability-per-token route. See chinese-open-weight-frontier-parity.
-
salim-ismail in 2026-06-26-podcast-moonshots-the-10b-satellite-empire-putting-ai-in-orbit-why: "frontier intelligence cannot be monopolized anymore"
-
clark-tang in 2026-06-11-podcast-bg2-pod-the-spacex-ipo-fable-5-ai-capex-update-market, on why closed retained the value even as open took the volume: "the reason why closed source models have captured so much of the value is because the models actually get the intention and actually carry through the work"
-
gavin-baker in 2026-06-11-podcast-bg2-pod-the-spacex-ipo-fable-5-ai-capex-update-market, on the volume/value split: "Open source might be 80% of tokens." — and the inversion that follows: "It's actually really bullish for compute and hardware because if the frontier models are capturing less of the margin then you're going to spend more on compute." See open-source-share-shift-bullish-for-compute.
Design implications
- The right comparison unit is serving cost per unit of intelligence, not license cost. Whoever can serve a near-frontier open model fastest/cheapest (the wafer-scale bet of cerebras) captures the open-serving market.
- The closed premium is bounded by the (small) quality gap times each workload's sensitivity to that gap. Joe's prediction — companies getting "more skilled at allocating from different forms of inference" — implies the closed premium is a task-routing question, not a blanket one: premium closed model for the hard 5% of tasks, cheap open model for the rest.
- Chinese open-weight models (Kimi, Qwen) are a live part of the US enterprise stack, which carries an export-control / data-governance dimension this source only gestures at.
Contradictions / tensions
-
The headline cost claims come from a CEO whose inference cloud monetizes serving open models — clear incentive to talk up open-model economics. Treat "cheaper by a lot" as directional.
-
Closed labs are not standing still: a 3–5% quality gap measured in May 2026 is a vintage snapshot (see the thread's capability-tracking discipline); the gap and the premium could widen or narrow with each frontier release.
-
"No single winner" is a forecast, not an observation — Feldman's x86/ARM analogy is plausible but the AI model market could still consolidate around 1–2 closed labs if the quality gap proves to compound rather than stay constant.
-
The May 2026 vintage snapshot has been overtaken. This page was written when "closed is strictly better by a little bit" was the state of play. By late June 2026, GLM 5.2 matched or exceeded top Western models on coding, long-range-agency and design benchmarks, and alexander-wissner-gross reports the "six to eight months behind" assumption starting to "creak." The quality framing may be the wrong axis entirely: GLM 5.2 is worse per token and better per dollar. Whether that counts as closing the gap depends on which denominator you price. Worth a
/calibrateentry — the "closed is durably better" prior moved on evidence. -
Distillation may explain the convergence rather than refute the gap. will-marshall: "They did distillation almost certainly on the best models." If parity is borrowed, the premium is a lead measured in trace-generation cycles, not a moat. See distillation-and-iterated-amplification.
HN / FT color (August 2026, not a rate card)
- From 2026-08-27-significant-ai-developments-last-30-days: HN top story "Beating GPT-5.6 Sol on retrieval with 100x cheaper open models" (Neon/Castform, 437 points, 5 Aug) is discourse on the cheap-open vs frontier-closed premium. Not a fetched Neon price table and not a rewrite of gpt-5-6-sol-pricing (official over-20% cut, already ingested this week).
- From the same source: FT headline "OpenAI and Anthropic in price war as Chinese AI rivals gain ground" is an FT headline on HN, not a rate card. Do not invent prices or dates from it.
2026-09-03 — closed Chinese API at the Code Arena price frontier; IFM openness claim; Mostik lock-in framing
Chronological color — does not rewrite the May Feldman 3–5% snapshot or the June GLM cost-route update.
- From 2026-09-03-x-ai-pass-3-sep-2026-open-model-fleets-agentic-coding-price (@Alibaba_Qwen + @arena): qwen-3-8-max-0902 is a closed API at $2 / $6 per 1M tokens (cache $0.17 / $0.25), listed against claude-fable-5-1's $10 / $50, and placed by the arena operator at the blended $5/MToken Pareto frontier. The 3-point WebDev lead over Claude Opus 5 (Max) is a statistical tie, not a quality-gap collapse. Durable fact is the price axis. Full treatment: qwen-3-8-max-0902.
- From the same source (@IFM_AI): k2-horizon leads with "fully open code, training data and recipes." Not a license grant. No bench table.
- From the same source (@aimalysheva): mostik-ai frames lock-in as the wrong question ("why does a frontier model have to generate your answer at all") and names a commercial interest in "preventing frontier model lock-in." Performance figures disagree across two same-day posts — do not carry them. See latent-space-model-handoff.
2026-09-07 — MiniCPM5-2B Apache-2.0 + open data (not a quality-gap rewrite)
Chronological color — does not rewrite the May Feldman 3–5% snapshot or the June GLM cost-route update.
- From 2026-09-07-openbmb-minicpm5-2b-open-release-aa-index-15: openbmb released minicpm5-2b under Apache 2.0 with accompanying UltraData / UltraX / JustRL II stack pieces. Independent AA Intelligence Index v4.2 = 15 (under-4B open lead). Launch Index 23 is a superseded issuer claim — keep both; prefer 15. Issuer-table 53.9 is not an AA composite. No serving-price table in this pass — do not invent a levelized-cost figure. Full treatment: minicpm5-2b.
2026-09-07 evening — AA cost-per-Index-task (v4.3; not a Feldman rewrite)
Chronological color — does not rewrite the May Feldman 3–5% snapshot, the June GLM cost-route update, or MiniCPM’s v4.2 15.
- From 2026-09-07-artificial-analysis-intelligence-index-v4-3 (cost post): at Index 53, gpt-6-astra (max) $3.26/task vs claude-fable-5-1 (max with fallback) $7.63/task (57% lower for Astra). At Index 42, glm-5-3-flash $0.25 vs GPT-5.6 Terra (max) $1.40. GPT-5.6 Luna (max) Index 38 at $0.18/task. Root long-post: OpenAI occupies most of the intelligence-vs-cost frontier across Astra reasoning efforts; Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42), and MiMo-V2.5-Pro (26) also on that frontier. This is AA’s cost-per-Index-task, not Joe’s still-missing levelized-cost-of-intelligence unit. Full treatment: artificial-analysis-intelligence-index.
2026-09-08 — Rumik OSS 1 CC-BY-NC Indic TTS (not a quality-gap rewrite)
Chronological color — open weights with a research/non-commercial license, not Apache-2.0. Does not rewrite Feldman 3–5% or AA cost-per-Index-task.
- From 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit (introducing rumik oss 1 + HF card): rumik-oss-1 is 3B open weights under tiny Aya Fire CC-BY-NC 4.0 + AUP. Issuer IndicEmo 2.92/5 vs Gemini 3.1 Flash TTS 4.58 on their table — Gemini still higher; do not invent “beats closed models.” No serving-price table. Full treatment: rumik-oss-1.
2026-09-10 — DeepSeek-V4.1-Flash official USD table (not a Feldman rewrite)
Chronological color — issuer rate card, not AA cost-per-Index-task and not Joe’s still-missing LCOI unit. Does not rewrite Feldman 3–5% or the June GLM cost-route update.
- From 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch (Models & Pricing): deepseek-v4-1-flash (
deepseek-flash) lists USD per 1M tokens — cache-hit input off-peak $0.003 / peak $0.006; cache-miss input off-peak $0.15 / peak $0.30; output off-peak $0.60 / peak $1.20. Off-peak = 50% of peak; peak hours 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri. Context 1M, max output 384K, vision ✓. Cite the official USD table; do not invent FX from RMB recaps. Open-weights pointer on Hugging Face; tech-report PDF not parsed. Full treatment: deepseek-v4-1-flash.
Open questions
- Does the closed-quality premium widen (gap compounds → closed pulls away) or collapse (open catches up → premium → zero) over the next several frontier-release cycles?
- Will a standardized "levelized cost of intelligence" metric emerge, and who defines it? Without it, the premium stays hard to price.
- How much of the quiet enterprise migration to Chinese open models survives export-control / data-governance scrutiny?
Related
- andrew-feldman
- joe-weisenthal
- tracy-alloway
- cerebras
- nvidia
- llm-as-commodity-thesis
- cuda-moat-erosion-at-inference
- inference-speed-as-a-pricing-premium
- chinese-open-weight-frontier-parity
- distillation-and-iterated-amplification
- proprietary-coding-data-as-moat
- open-source-share-shift-bullish-for-compute
- gpt-5-6-sol-pricing
- k2-horizon
- qwen-3-8-max-0902
- mostik-ai
- latent-space-model-handoff
- claude-fable-5-1
- minicpm5-2b
- openbmb
- artificial-analysis-intelligence-index
- gpt-6-astra
- glm-5-3-flash
- rumik-oss-1
- rumik-ai
- deepseek
- deepseek-v4-1-flash