Naive-N0.5-Flash: 309B open-weights MoE with 1M context and no full-attention layers
Beijing startup NaiveAI released Naive-N0.5-Flash (Sep 27/28) under MIT: a 309B MoE (15.5B active) with native 1M-token context built from hybrid sliding-window + DeepSeek sparse attention, no full-attention layers anywhere, trained on Xiaomi MiMo-V2.5 base with AI-assisted research loops.
Why it matters
The first large open-weights model with zero full-attention layers reaching 1M context — a real architectural bet on sparse attention at frontier scale, plus a live demo of AI-assisted model R&D (151 optimization rounds in six days).
What you could build with it
An indie hacker could serve Naive-N0.5-Flash as a cheap long-context API for whole-repo code Q&A or multi-day research dossiers, undercutting frontier pricing ($0.10/$0.40 per 1M tokens) for context-heavy workloads.
Does it hold up?
Too early to judge — only vendor-reported benchmarks exist; CellCog reports an unresolved IP dispute between NaiveAI and MiroMind over the technology claims, and the 315GB weights keep third parties at arm's length.
Built with Naive-N0.5-Flash
- NaiveAI Open-Weights 309B Naive-N0.5-Flash With No Full Attentionaiweekly · Independent technical breakdown of the hybrid SWA-DSA architecture; flags that CellCog reports an unresolved IP dispute between NaiveAI and MiroMind over technology claims underlying the release.
- NaiveAI releases a 309B model built with AI-assisted researchruntimewire · Journalist breakdown of the AI-assisted-research claim, stressing all performance figures are company-reported with no independent verification published.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.