brain/
sourceartificial-intelligence

X 10am: DeepSeek-V4.1-Flash open multimodal MoE launch

Sep 10 issuer X + docs: DeepSeek launches V4.1-Flash (552B MoE, 8B/16B active Causal Encoder–Decoder, native vision), API deepseek-flash, V4-Pro routed to Flash from Sep 14 UTC, peak/off-peak USD pricing cut.

Source

X 10am: DeepSeek-V4.1-Flash open multimodal MoE launch

Generated by Grok Bot research on 2026-09-10. WebSearch + fetch ladder + native X. Treat as raw material — review before promoting into a project or thread.

Dedup: Same-day AI-thread source 2026-09-10-overnight-x-openai-defense-factory-uk-aisi-mythos-51-access.md already covers OpenAI Defense Factory, UK AISI Mythos 5.1 access discourse, Unity Claude Code plugin (video), and DeepMind RT of Paglieri. Do not re-file those permalinks. This 10am pass is new issuer grain only: @deepseek_ai V4.1-Flash launch thread (~06:10 UTC).

Summary

On 2026-09-10 (~06:10 UTC), @deepseek_ai introduced DeepSeek-V4.1-Flash as the smallest model in a new architecture family with native visual understanding, a 552B-parameter MoE, and a Causal Encoder–Decoder that activates 8B parameters on input and 16B on output. Issuer docs put the model live on the API as deepseek-flash, retire prior Flash / Flash-Vision-Exp aliases onto it, claim V4.1-Flash beats V4-Pro on performance/cost/speed/runtime, and schedule deepseek-v4-pro traffic to route to Flash at Flash rates from 04:00 UTC on 2026-09-14 until V4.1-Pro. Contested: third-party “matches Opus 5” framing from news clusters; stick to issuer benchmarks and pricing tables. Unconfirmed Amodei/Pentagon rumor clusters around the drop are not grain.

Findings

DeepSeek issuer X: V4.1-Flash architecture + open weights

  • @deepseek_ai (2026-09-10 06:10:09Z; note_tweet; photo): introduces DeepSeek-V4.1-Flash — “smallest model in our new architecture family, with native visual understanding”; designed for greater capability, faster inference, higher throughput, and scaling to larger models (1/6).
  • @deepseek_ai (2026-09-10 06:10:10Z; note_tweet; photo): 552B-parameter MoE; new Causal Encoder–Decoder with 8B active on input and 16B on output; claims new pre-training + larger-scale RL post-training deliver benchmarks ahead of flagship models including DeepSeek-V4-Pro (2/6). Treat “ahead of flagships” as issuer claim pending independent evals.
  • @deepseek_ai (2026-09-10 06:10:11Z; note_tweet; photo): vs prior generation, V4.1-Flash KV cache needs 1/4 the HBM and 1/8 the SSD storage; frames cache-hit charges as a large share of agent costs (3/6).
  • @deepseek_ai (2026-09-10 06:10:12Z; note_tweet): live on DeepSeek API with native multimodal; set model to deepseek-flash; V4-Flash & V4-Flash-Vision-Exp retired with temporary route aliases; V4-Pro phasing out — from 04:00 UTC Sept 14, 2026 all deepseek-v4-pro requests route to V4.1-Flash at Flash rates until V4.1-Pro; partners @WorkBuddy_AI (incl. CodeBuddy) and @opencode support (4/6).
  • @deepseek_ai (2026-09-10 06:10:13Z; note_tweet; photo): peak/off-peak pricing continues; off-peak = 50% of peak; new pricing effective 04:00 UTC on Sept 10, 2026 (5/6).
  • @deepseek_ai (2026-09-10 06:10:14Z; note_tweet): open-source inference support; invites large-scale deploy talk (2,000 GPUs + storage cluster); links Hugging Face model and tech report PDF (6/6). Attachments across the thread are photos, not video.

Issuer pages (hydrated)

  • DeepSeek news: Introducing DeepSeek-V4.1-Flash (WebFetch 200, 2026-09-10): restates the X claims — 552B MoE, 8B/16B Causal Encoder–Decoder, KV-cache HBM/SSD ratios, API deepseek-flash, V4-Pro route from 04:00 UTC Sept 14, off-peak half of peak, pricing effective 04:00 UTC Sept 10, WorkBuddy/OpenCode partners.
  • DeepSeek API Change Log — 2026-09-10 (curl 200; WebFetch 409): official release note with issuer benchmark table for V4.1-Flash, including GPQA Diamond 90.9, HLE 36.8 (39.1* pure-text subset), Codeforces rating 3471, MathArena Apex 65.6, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, Terminal-Bench 4.0 31.2, DeepSWE v1.1 74.2, ProgramBench 20.3, NL2Repo-Bench 65.4, CyberGym 88.1, SEC-Bench Pro 62.8, ExploitGym 15.3, HLE w/tools 63.9, Automation-Bench 54.8, Agents' Last Exam 31.8, Chartography w/tools 78.9, BabyVision w/tools 89.6, ZeroBench-main w/tools 49.0. Same API retirement / V4-Pro routing language as X.
  • DeepSeek Models & Pricing (curl 200; WebFetch 409): deepseek-flash → DeepSeek-V4.1-Flash; context 1M, max output 384K, vision ✓. USD per 1M tokens for Flash: cache-hit input off-peak $0.003 / peak $0.006; cache-miss input off-peak $0.15 / peak $0.30; output off-peak $0.60 / peak $1.20. Peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday. Concurrency limit listed 2500 for Flash vs 500 for Pro. Footnotes match the V4-Pro → Flash routing date.

Skipped / discovery-only this pass

  • OpenAI Defense Factory, Byrne/AISI Mythos 5.1, Unity Claude plugin video, DeepMind Paglieri RT — already in today’s overnight AI-thread source.
  • search_news clusters on Claude Marketplace expansion, AutoResearchExam (Bespoke Labs), Opus-button satire, agent-harness discourse, and unconfirmed Amodei/Pentagon DeepSeek briefing — do not cite news-cluster summaries as posts; Marketplace needs Anthropic issuer primary (no new @AnthropicAI posts in window); AutoResearchExam is vendor-benchmark secondary unless a later pass hydrates it; Pentagon claim stays unconfirmed rumor.
  • @karpathy — zero original posts in window (real silence).
  • @OpenAI / @GoogleDeepMind / @sama / @ArtificialAnlys in window — already filed or non-new grain for this pass.

Contradictions and open questions

  • Issuer claim that V4.1-Flash is “ahead of flagship models, including DeepSeek-V4-Pro” should stay tagged as DeepSeek self-report until independent third-party boards (e.g. Artificial Analysis) re-score the drop.
  • Routing V4-Pro → Flash at Flash prices from Sep 14 is a product retirement, not a free upgrade narrative — customers lose the Pro SKU until V4.1-Pro.
  • USD pricing on api-docs.deepseek.com vs RMB figures in secondary TechNode/IT Home recaps — cite the official USD table retrieved here; do not invent FX conversions.
  • News-cluster “matches or beats Claude Opus 5” is not in the issuer X/docs pull; do not promote that comparison from Grok news summaries.
  • Open-weight cyber-risk discourse (already live in today’s AISI clip) will re-fire around a 552B multimodal Flash release — keep that policy thread separate from this product grain unless AISI/Anthropic issue new primaries.

Provenance

Method: Grok Bot / WebSearch / fetch ladder (WebFetch news page; curl for api-docs) / native X (search_news, get_users_by_username(s), get_users_posts, get_posts_by_ids; no search_posts_all) Generated: 2026-09-10 Rounds: 1 of 3 — early-exit after issuer X thread + news + changelog + pricing hydration; further rounds would chase Marketplace/AutoResearchExam without new lab posts. URLs fetched: deepseek.com news success (WebFetch); api-docs updates + pricing success via curl (WebFetch 409 both); Hugging Face links cited from X entities (not PDF-parsed this pass). X spend note: conservative lab timelines + one search_news; karpathy silent; DeepSeek thread is the only new high-signal issuer drop vs today’s overnight clip.

Web sources:

X sources:

Grokipedia:

  • not used
Referenced by