DeepSeek-V4.1-Flash
Vintage: 2026-09. Primary evidence is official @deepseek_ai X (2026-09-10 ~06:10 UTC) plus fetched DeepSeek news, API Change Log 2026-09-10, and Models & Pricing, as synthesized in 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch (
method: grok-bot;x_video: false). Every capability number on this page is DeepSeek-reported, not a third-party rerun. Hugging Face weights + tech-report PDF are issuer pointers — not PDF-parsed. Thread attachments are photos, not video.
DeepSeek-V4.1-Flash
One-line summary: deepseek's 10 Sep 2026 open multimodal MoE — issuer 552B total, Causal Encoder–Decoder 8B active on input / 16B on output; API id deepseek-flash; native vision; USD peak/off-peak table; V4-Pro traffic scheduled to route here from 04:00 UTC 2026-09-14.
What it is
Issuer-described as the smallest model in a new architecture family with native visual understanding. Live on the DeepSeek API as deepseek-flash. Prior V4-Flash and V4-Flash-Vision-Exp aliases retire onto it. From 04:00 UTC 14 Sep 2026, deepseek-v4-pro requests route to V4.1-Flash at Flash rates until V4.1-Pro — a SKU retirement, not a free-upgrade narrative.
This page records issuer-hydrated claims from official X + fetched news / changelog / pricing. It does not promote news-cluster “matches Opus 5” language (absent from the issuer pull).
Why it matters to this thread
Frontier releases, open weights, multimodal understanding, and API pricing are in-scope. This is the first dated V4.1-Flash page. Distinct from glm-5-3-flash / gemini-3-8-flash / minicpm5-2b. Do not rank issuer benches against other labs’ cards as if they share a harness.
Key facts (from 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch)
Architecture (issuer X + news)
- From 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch (@deepseek_ai, 2026-09-10): “smallest model in our new architecture family, with native visual understanding”; framed for greater capability, faster inference, higher throughput, and scaling to larger models.
- From the same source (@deepseek_ai + news): 552B-parameter MoE; new Causal Encoder–Decoder with 8B active on input and 16B on output; claims new pre-training + larger-scale RL post-training deliver benchmarks ahead of flagship models including DeepSeek-V4-Pro. Treat “ahead of flagships” as DeepSeek self-report pending independent evals.
- From the same source (@deepseek_ai): vs prior generation, V4.1-Flash KV cache needs 1/4 the HBM and 1/8 the SSD storage; frames cache-hit charges as a large share of agent costs.
API id, retirement, partners
- From the same source (@deepseek_ai + Change Log): set model to
deepseek-flash; V4-Flash & V4-Flash-Vision-Exp retired with temporary route aliases; from 04:00 UTC Sept 14, 2026 alldeepseek-v4-prorequests route to V4.1-Flash at Flash rates until V4.1-Pro. - From the same source: partners @WorkBuddy_AI (incl. CodeBuddy) and opencode support. No WorkBuddy page minted.
- From the same source (@deepseek_ai): open-source inference support; 2,000 GPUs + storage cluster invite; HF model + tech report PDF (not parsed).
Issuer benchmark table (changelog)
All of the following are DeepSeek-reported on the 2026-09-10 Change Log, not reproduced independently:
- GPQA Diamond 90.9
- HLE 36.8 (39.1* pure-text subset); HLE w/tools 63.9
- Codeforces rating 3471
- MathArena Apex 65.6
- Terminal-Bench 2.1 90.6; Terminal-Bench 3.0 30.0; Terminal-Bench 4.0 31.2
- DeepSWE v1.1 74.2; ProgramBench 20.3; NL2Repo-Bench 65.4
- CyberGym 88.1; SEC-Bench Pro 62.8; ExploitGym 15.3
- Automation-Bench 54.8; Agents' Last Exam 31.8
- Chartography w/tools 78.9; BabyVision w/tools 89.6; ZeroBench-main w/tools 49.0
Do not rank these against claude-fable-5-1 / gpt-6-astra / AA Index readings as if they share a harness. Prior AA Index v4.3 color on DeepSeek V4 Pro 0813 (max) 36 is a different instrument and a different SKU — see chinese-open-weight-frontier-parity.
USD pricing and limits (official table)
From the same source (Models & Pricing; new pricing effective 04:00 UTC on Sept 10, 2026):
deepseek-flash→ DeepSeek-V4.1-Flash; context 1M; max output 384K; vision ✓.- USD per 1M tokens: cache-hit input off-peak $0.003 / peak $0.006; cache-miss input off-peak $0.15 / peak $0.30; output off-peak $0.60 / peak $1.20.
- Off-peak = 50% of peak (@deepseek_ai).
- Peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday.
- Concurrency limit listed 2500 for Flash vs 500 for Pro.
Cite this official USD table. Do not invent FX conversions from secondary RMB recaps.
What this source does not establish
- No independent bake-off. “Ahead of flagships including V4-Pro” is issuer claim.
- “Matches or beats Claude Opus 5” is not in the issuer X/docs pull — do not promote it.
- Tech report PDF not parsed. Architecture details beyond the X/news restatement are not invented.
- V4-Pro → Flash routing is a product retirement, not a free upgrade. Customers lose the Pro SKU until V4.1-Pro.
- Did not re-file openai-defense-factory, uk-aisi-mythos-51-access, Unity Claude plugin video, or DeepMind Paglieri RT (same-day overnight source).
- Unconfirmed Amodei/Pentagon rumor clusters are not grain.
- No video.
x_video: false. Photos only. No transcripts invented. - No ticker. No stock-market tag.
- Not Anthropic’s DeepSeek distillation allegation. From 2026-09-11-anthropic-sep-2026-threat-intelligence-autonomous-cyber: Sep 10 threat-intel GTG-16001 volumes sit beside this product drop — do not merge benches / rate card with issuer-alleged distillation. Full treatment: anthropic-illicit-distillation-sep-2026.
Contradictions / tensions
- Issuer “ahead of flagships including V4-Pro” vs no third-party rescore in this pass. Filed as self-report. Not reconciled.
- Routing V4-Pro → Flash at Flash prices is cheaper for remaining Pro traffic and removes the Pro SKU until V4.1-Pro. Not flattened into an “upgrade” story.
- USD official table vs RMB secondary recaps. Cite USD only.
- Issuer Terminal-Bench / CyberGym / HLE numbers vs other labs’ same-named benches. Harness, tools, and effort are not shown to match.
Open questions
- Will Artificial Analysis (or another independent board) re-score V4.1-Flash, and how will that sit next to V4 Pro 0813 (max) 36 on Index v4.3?
- What does the unparsed tech-report PDF add (training recipe, license, serving stack) that X + changelog omit?
- When does V4.1-Pro ship, and does Pro concurrency / pricing return?
Sources
- 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch
- 2026-09-11-anthropic-sep-2026-threat-intelligence-autonomous-cyber — adjacent threat-intel accusation; do not merge