X 10am: DeepSeek-V4.1-Flash open multimodal MoE launch
Sep 10 issuer X + docs: DeepSeek launches V4.1-Flash (552B MoE, 8B/16B active Causal Encoder–Decoder, native vision), API deepseek-flash, V4-Pro routed to Flash from Sep 14 UTC, peak/off-peak USD pricing cut.
X 10am: DeepSeek-V4.1-Flash open multimodal MoE launch
Generated by Grok Bot research on 2026-09-10. WebSearch + fetch ladder + native X. Treat as raw material — review before promoting into a project or thread.
Dedup: Same-day AI-thread source
2026-09-10-overnight-x-openai-defense-factory-uk-aisi-mythos-51-access.mdalready covers OpenAI Defense Factory, UK AISI Mythos 5.1 access discourse, Unity Claude Code plugin (video), and DeepMind RT of Paglieri. Do not re-file those permalinks. This 10am pass is new issuer grain only: @deepseek_ai V4.1-Flash launch thread (~06:10 UTC).
Summary
On 2026-09-10 (~06:10 UTC), @deepseek_ai introduced DeepSeek-V4.1-Flash as the smallest model in a new architecture family with native visual understanding, a 552B-parameter MoE, and a Causal Encoder–Decoder that activates 8B parameters on input and 16B on output. Issuer docs put the model live on the API as deepseek-flash, retire prior Flash / Flash-Vision-Exp aliases onto it, claim V4.1-Flash beats V4-Pro on performance/cost/speed/runtime, and schedule deepseek-v4-pro traffic to route to Flash at Flash rates from 04:00 UTC on 2026-09-14 until V4.1-Pro. Contested: third-party “matches Opus 5” framing from news clusters; stick to issuer benchmarks and pricing tables. Unconfirmed Amodei/Pentagon rumor clusters around the drop are not grain.
Findings
DeepSeek issuer X: V4.1-Flash architecture + open weights
- @deepseek_ai (2026-09-10 06:10:09Z; note_tweet; photo): introduces DeepSeek-V4.1-Flash — “smallest model in our new architecture family, with native visual understanding”; designed for greater capability, faster inference, higher throughput, and scaling to larger models (1/6).
- @deepseek_ai (2026-09-10 06:10:10Z; note_tweet; photo): 552B-parameter MoE; new Causal Encoder–Decoder with 8B active on input and 16B on output; claims new pre-training + larger-scale RL post-training deliver benchmarks ahead of flagship models including DeepSeek-V4-Pro (2/6). Treat “ahead of flagships” as issuer claim pending independent evals.
- @deepseek_ai (2026-09-10 06:10:11Z; note_tweet; photo): vs prior generation, V4.1-Flash KV cache needs 1/4 the HBM and 1/8 the SSD storage; frames cache-hit charges as a large share of agent costs (3/6).
- @deepseek_ai (2026-09-10 06:10:12Z; note_tweet): live on DeepSeek API with native multimodal; set model to
deepseek-flash; V4-Flash & V4-Flash-Vision-Exp retired with temporary route aliases; V4-Pro phasing out — from 04:00 UTC Sept 14, 2026 alldeepseek-v4-prorequests route to V4.1-Flash at Flash rates until V4.1-Pro; partners @WorkBuddy_AI (incl. CodeBuddy) and @opencode support (4/6). - @deepseek_ai (2026-09-10 06:10:13Z; note_tweet; photo): peak/off-peak pricing continues; off-peak = 50% of peak; new pricing effective 04:00 UTC on Sept 10, 2026 (5/6).
- @deepseek_ai (2026-09-10 06:10:14Z; note_tweet): open-source inference support; invites large-scale deploy talk (2,000 GPUs + storage cluster); links Hugging Face model and tech report PDF (6/6). Attachments across the thread are photos, not video.
Issuer pages (hydrated)
- DeepSeek news: Introducing DeepSeek-V4.1-Flash (WebFetch 200, 2026-09-10): restates the X claims — 552B MoE, 8B/16B Causal Encoder–Decoder, KV-cache HBM/SSD ratios, API
deepseek-flash, V4-Pro route from 04:00 UTC Sept 14, off-peak half of peak, pricing effective 04:00 UTC Sept 10, WorkBuddy/OpenCode partners. - DeepSeek API Change Log — 2026-09-10 (curl 200; WebFetch 409): official release note with issuer benchmark table for V4.1-Flash, including GPQA Diamond 90.9, HLE 36.8 (39.1* pure-text subset), Codeforces rating 3471, MathArena Apex 65.6, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, Terminal-Bench 4.0 31.2, DeepSWE v1.1 74.2, ProgramBench 20.3, NL2Repo-Bench 65.4, CyberGym 88.1, SEC-Bench Pro 62.8, ExploitGym 15.3, HLE w/tools 63.9, Automation-Bench 54.8, Agents' Last Exam 31.8, Chartography w/tools 78.9, BabyVision w/tools 89.6, ZeroBench-main w/tools 49.0. Same API retirement / V4-Pro routing language as X.
- DeepSeek Models & Pricing (curl 200; WebFetch 409):
deepseek-flash→ DeepSeek-V4.1-Flash; context 1M, max output 384K, vision ✓. USD per 1M tokens for Flash: cache-hit input off-peak $0.003 / peak $0.006; cache-miss input off-peak $0.15 / peak $0.30; output off-peak $0.60 / peak $1.20. Peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. Concurrency limit listed 2500 for Flash vs 500 for Pro. Footnotes match the V4-Pro → Flash routing date.
Skipped / discovery-only this pass
- OpenAI Defense Factory, Byrne/AISI Mythos 5.1, Unity Claude plugin video, DeepMind Paglieri RT — already in today’s overnight AI-thread source.
search_newsclusters on Claude Marketplace expansion, AutoResearchExam (Bespoke Labs), Opus-button satire, agent-harness discourse, and unconfirmed Amodei/Pentagon DeepSeek briefing — do not cite news-cluster summaries as posts; Marketplace needs Anthropic issuer primary (no new@AnthropicAIposts in window); AutoResearchExam is vendor-benchmark secondary unless a later pass hydrates it; Pentagon claim stays unconfirmed rumor.@karpathy— zero original posts in window (real silence).@OpenAI/@GoogleDeepMind/@sama/@ArtificialAnlysin window — already filed or non-new grain for this pass.
Contradictions and open questions
- Issuer claim that V4.1-Flash is “ahead of flagship models, including DeepSeek-V4-Pro” should stay tagged as DeepSeek self-report until independent third-party boards (e.g. Artificial Analysis) re-score the drop.
- Routing V4-Pro → Flash at Flash prices from Sep 14 is a product retirement, not a free upgrade narrative — customers lose the Pro SKU until V4.1-Pro.
- USD pricing on
api-docs.deepseek.comvs RMB figures in secondary TechNode/IT Home recaps — cite the official USD table retrieved here; do not invent FX conversions. - News-cluster “matches or beats Claude Opus 5” is not in the issuer X/docs pull; do not promote that comparison from Grok news summaries.
- Open-weight cyber-risk discourse (already live in today’s AISI clip) will re-fire around a 552B multimodal Flash release — keep that policy thread separate from this product grain unless AISI/Anthropic issue new primaries.
Provenance
Method: Grok Bot / WebSearch / fetch ladder (WebFetch news page; curl for api-docs) / native X (search_news, get_users_by_username(s), get_users_posts, get_posts_by_ids; no search_posts_all)
Generated: 2026-09-10
Rounds: 1 of 3 — early-exit after issuer X thread + news + changelog + pricing hydration; further rounds would chase Marketplace/AutoResearchExam without new lab posts.
URLs fetched: deepseek.com news success (WebFetch); api-docs updates + pricing success via curl (WebFetch 409 both); Hugging Face links cited from X entities (not PDF-parsed this pass).
X spend note: conservative lab timelines + one search_news; karpathy silent; DeepSeek thread is the only new high-signal issuer drop vs today’s overnight clip.
Web sources:
- Introducing DeepSeek-V4.1-Flash (DeepSeek news) — architecture, API, retirement, pricing effective date
- DeepSeek API Change Log 2026-09-10 — official benchmarks + API routing
- DeepSeek Models & Pricing — USD peak/off-peak Flash table, context/vision limits
- DeepSeek-V4.1-Flash on Hugging Face — open weights pointer from issuer X
- DeepSeek_V41_Tech_Report.pdf — paper pointer from issuer X (not PDF-parsed)
X sources:
- X post by @deepseek_ai (2026-09-10) — launch 1/6
- X post by @deepseek_ai (2026-09-10) — 552B MoE / 8B–16B active 2/6
- X post by @deepseek_ai (2026-09-10) — KV cache 1/4 HBM 1/8 SSD 3/6
- X post by @deepseek_ai (2026-09-10) — API deepseek-flash + V4-Pro route Sep 14 4/6
- X post by @deepseek_ai (2026-09-10) — off-peak 50% + pricing effective 5/6
- X post by @deepseek_ai (2026-09-10) — HF model + paper 6/6
Grokipedia:
- not used