Mooney: 2.39-bit Qwen3.8-Flash-Next quant that fits a DGX Spark
Qwen Dev Ambassador Daniel Lougen released Mooney, a 2.39-bit-per-weight quant of the ~180B MoE Qwen3.8-Flash-Next that runs on a single NVIDIA DGX Spark at 95.5% of full-precision quality.
Why it matters
A 180B-class frontier MoE compressed to run on one consumer workstation with near-full quality — independently rebuilt and verified the same day, proving frontier local inference doesn't need a data center.
What you could build with it
A hardware startup or homelab community could package the Mooney recipe — a 180B-class MoE squeezed onto a single workstation — into a one-script installer, giving small teams frontier-adjacent local inference without cloud bills.
Does it hold up?
Evidence-backed: an independent team rebuilt it the same afternoon, hash-verified all 90 GB, and ran it on their own DGX Spark — 45.8 tok/s decode, 95.5% of BF16 quality on a 536-task suite. Requires a DGX Spark and comfort with the cuda.fast ds4 engine; not a one-click install.
Built with Mooney
- Mooney: squeezing a 180B model onto one DGX Spark at 2.39 bits/weight (independent rebuild)al-engr.com · Independent engineers cloned the repo, hash-verified all 90 GB, and A/B tested it on their own DGX Spark the same afternoon: 45.8 tok/s decode with the speculative MTP head.
- llama.cpp merges Qwen multi-token decoding with a reported 55% speed gaingroundtruth.day · Related runtime work: merged llama.cpp support for Qwen3.8-Flash-Next's multi-token prediction head for speculative decoding.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.