DeepSeek V4.1 Flash: 552B open-weight MoE goes live under MIT
DeepSeek released the weights of V4.1 Flash on October 6, 2026 via ModelScope: a 552B-parameter multimodal MoE (1 shared + 384 routed experts, top-6 routing) with ~8B/16B active per token, 1,048,576-token context, and ~890 bytes of KV cache per token.
Why it matters
MIT-licensed weights for a model independent evaluator Vals AI ranks #1 among open-weight models, with a KV cache about 1/437th of DeepSeek's first-generation model.
What you could build with it
A startup could self-host V4.1 Flash on its own GPUs to offer a cheap long-context coding or document-analysis API, undercutting closed-API prices with MIT-licensed weights and the model's tiny per-token active footprint.
Does it hold up?
Too early to judge: weights only landed on October 6. The independent record so far is one Vals AI index ranking and a caution that DeepSeek's own harness inflated its headline coding score.
Built with DeepSeek V4.1 Flash
- abhs.in deep dive: DeepSeek V4.1 Flash specs, pricing and benchmarksarticle · Independent walkthrough of the architecture (40-layer causal encoder-decoder, FP4 KV cache), the MIT license, API pricing ($0.30/$1.20 per M tokens), and Vals AI's #1 open-weight ranking alongside a note that DeepSeek's own coding harness inflated its headline score.
- dev.to: the great sparsification of 2026 open-weight MoE modelsarticle · Analysis of the active-vs-total parameter split across recent open releases, placing V4.1 Flash at ~1.4% active (8B prefill / 16B decode of 552B) and explaining why that ratio drives inference budgets.
Learn more
First spotted on article: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71JetBrains Mellum 2.1: 12B MoE coding model trained for agentsJetBrains released Mellum 2.1, a 12B MoE model (2.5B active parameters, 128K context) under Apache 2.0, trained mainly…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.