DeepSeek-V4.1-Flash: 552B MoE with 890-byte-per-token KV cache
DeepSeek's MIT-licensed multimodal MoE uses a causal encoder-decoder (CED) design with cross-layer KV sharing, shrinking the KV cache to 890 bytes per token for million-token agent workloads.
Why it matters
First widely-available model with a causal encoder-decoder architecture: roughly 1/4 the KV cache of V4-Flash and 1/437 of V1, so a 1M-token agent context costs well under 1 GB of cache while leading V4-Pro on Terminal-Bench 2.1 (90.6) and DeepSWE v1.1 (74.2).
What you could build with it
An indie developer could run a document-intelligence service that keeps entire corporate document repositories resident in a million-token cache and answers cross-document questions, with cache costs a fraction of what a vanilla transformer architecture would require.
Does it hold up?
Strong evidence: independent analyses verify the architecture and reported benchmarks (Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2); self-hosting needs serious infra (552B weights), so most users consume it via API or hosted weights.
Built with DeepSeek-V4.1-Flash
- From 390 KB to 890 Bytes — V4.1-Flash: The 890-Byte Cachelocal-ai-zone.github.io · Independent deep-dive chapter on V4.1-Flash's CED split, CSA2 attention and 4-bit MXFP4 KV design, explaining how it halves prefill compute for input-heavy agent workloads.
- PocketLLM: DeepSeek V4.1 Flash design notesGitHub · Independent notes transcribing the model config from released tensors, measuring per-step expert-execution timings, and documenting self-hosting constraints.
- DeepSeek V4.1 Flash pricing, specs and API routingyottalabs.ai · Specs and pricing table comparing V4.1-Flash against V4-Flash and V4-Pro, with API pricing starting at $0.003 per million cached input tokens off-peak.
- deepseek-ai/deepseek-recipegithub · 373 ★ · Official companion toolkit (Rust libraries with Python bindings) that encodes Messages/Chat Completions/Responses API requests into DeepSeek V4/V4.1 prompt formats; referenced in the model card.
Learn more
First spotted on hf: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.