brain/
← all entities
entitygenericartificial-intelligence

DeepSeek-V4.1-Flash

Notes

Vintage: 2026-09. Primary evidence is official @deepseek_ai X (2026-09-10 ~06:10 UTC) plus fetched DeepSeek news, API Change Log 2026-09-10, and Models & Pricing, as synthesized in 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch (method: grok-bot; x_video: false). Every capability number on this page is DeepSeek-reported, not a third-party rerun. Hugging Face weights + tech-report PDF are issuer pointers — not PDF-parsed. Thread attachments are photos, not video.

DeepSeek-V4.1-Flash

One-line summary: deepseek's 10 Sep 2026 open multimodal MoE — issuer 552B total, Causal Encoder–Decoder 8B active on input / 16B on output; API id deepseek-flash; native vision; USD peak/off-peak table; V4-Pro traffic scheduled to route here from 04:00 UTC 2026-09-14.

What it is

Issuer-described as the smallest model in a new architecture family with native visual understanding. Live on the DeepSeek API as deepseek-flash. Prior V4-Flash and V4-Flash-Vision-Exp aliases retire onto it. From 04:00 UTC 14 Sep 2026, deepseek-v4-pro requests route to V4.1-Flash at Flash rates until V4.1-Pro — a SKU retirement, not a free-upgrade narrative.

This page records issuer-hydrated claims from official X + fetched news / changelog / pricing. It does not promote news-cluster “matches Opus 5” language (absent from the issuer pull).

Why it matters to this thread

Frontier releases, open weights, multimodal understanding, and API pricing are in-scope. This is the first dated V4.1-Flash page. Distinct from glm-5-3-flash / gemini-3-8-flash / minicpm5-2b. Do not rank issuer benches against other labs’ cards as if they share a harness.

Key facts (from 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch)

Architecture (issuer X + news)

  • From 2026-09-10-x-10am-deepseek-v41-flash-open-multimodal-moe-launch (@deepseek_ai, 2026-09-10): “smallest model in our new architecture family, with native visual understanding”; framed for greater capability, faster inference, higher throughput, and scaling to larger models.
  • From the same source (@deepseek_ai + news): 552B-parameter MoE; new Causal Encoder–Decoder with 8B active on input and 16B on output; claims new pre-training + larger-scale RL post-training deliver benchmarks ahead of flagship models including DeepSeek-V4-Pro. Treat “ahead of flagships” as DeepSeek self-report pending independent evals.
  • From the same source (@deepseek_ai): vs prior generation, V4.1-Flash KV cache needs 1/4 the HBM and 1/8 the SSD storage; frames cache-hit charges as a large share of agent costs.

API id, retirement, partners

  • From the same source (@deepseek_ai + Change Log): set model to deepseek-flash; V4-Flash & V4-Flash-Vision-Exp retired with temporary route aliases; from 04:00 UTC Sept 14, 2026 all deepseek-v4-pro requests route to V4.1-Flash at Flash rates until V4.1-Pro.
  • From the same source: partners @WorkBuddy_AI (incl. CodeBuddy) and opencode support. No WorkBuddy page minted.
  • From the same source (@deepseek_ai): open-source inference support; 2,000 GPUs + storage cluster invite; HF model + tech report PDF (not parsed).

Issuer benchmark table (changelog)

All of the following are DeepSeek-reported on the 2026-09-10 Change Log, not reproduced independently:

  • GPQA Diamond 90.9
  • HLE 36.8 (39.1* pure-text subset); HLE w/tools 63.9
  • Codeforces rating 3471
  • MathArena Apex 65.6
  • Terminal-Bench 2.1 90.6; Terminal-Bench 3.0 30.0; Terminal-Bench 4.0 31.2
  • DeepSWE v1.1 74.2; ProgramBench 20.3; NL2Repo-Bench 65.4
  • CyberGym 88.1; SEC-Bench Pro 62.8; ExploitGym 15.3
  • Automation-Bench 54.8; Agents' Last Exam 31.8
  • Chartography w/tools 78.9; BabyVision w/tools 89.6; ZeroBench-main w/tools 49.0

Do not rank these against claude-fable-5-1 / gpt-6-astra / AA Index readings as if they share a harness. Prior AA Index v4.3 color on DeepSeek V4 Pro 0813 (max) 36 is a different instrument and a different SKU — see chinese-open-weight-frontier-parity.

USD pricing and limits (official table)

From the same source (Models & Pricing; new pricing effective 04:00 UTC on Sept 10, 2026):

  • deepseek-flash → DeepSeek-V4.1-Flash; context 1M; max output 384K; vision ✓.
  • USD per 1M tokens: cache-hit input off-peak $0.003 / peak $0.006; cache-miss input off-peak $0.15 / peak $0.30; output off-peak $0.60 / peak $1.20.
  • Off-peak = 50% of peak (@deepseek_ai).
  • Peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday.
  • Concurrency limit listed 2500 for Flash vs 500 for Pro.

Cite this official USD table. Do not invent FX conversions from secondary RMB recaps.

What this source does not establish

  • No independent bake-off. “Ahead of flagships including V4-Pro” is issuer claim.
  • “Matches or beats Claude Opus 5” is not in the issuer X/docs pull — do not promote it.
  • Tech report PDF not parsed. Architecture details beyond the X/news restatement are not invented.
  • V4-Pro → Flash routing is a product retirement, not a free upgrade. Customers lose the Pro SKU until V4.1-Pro.
  • Did not re-file openai-defense-factory, uk-aisi-mythos-51-access, Unity Claude plugin video, or DeepMind Paglieri RT (same-day overnight source).
  • Unconfirmed Amodei/Pentagon rumor clusters are not grain.
  • No video. x_video: false. Photos only. No transcripts invented.
  • No ticker. No stock-market tag.
  • Not Anthropic’s DeepSeek distillation allegation. From 2026-09-11-anthropic-sep-2026-threat-intelligence-autonomous-cyber: Sep 10 threat-intel GTG-16001 volumes sit beside this product drop — do not merge benches / rate card with issuer-alleged distillation. Full treatment: anthropic-illicit-distillation-sep-2026.

Contradictions / tensions

  • Issuer “ahead of flagships including V4-Pro” vs no third-party rescore in this pass. Filed as self-report. Not reconciled.
  • Routing V4-Pro → Flash at Flash prices is cheaper for remaining Pro traffic and removes the Pro SKU until V4.1-Pro. Not flattened into an “upgrade” story.
  • USD official table vs RMB secondary recaps. Cite USD only.
  • Issuer Terminal-Bench / CyberGym / HLE numbers vs other labs’ same-named benches. Harness, tools, and effort are not shown to match.

Open questions

  • Will Artificial Analysis (or another independent board) re-score V4.1-Flash, and how will that sit next to V4 Pro 0813 (max) 36 on Index v4.3?
  • What does the unparsed tech-report PDF add (training recipe, license, serving stack) that X + changelog omit?
  • When does V4.1-Pro ship, and does Pro concurrency / pricing return?

Sources

Related

Referenced by