llama.cpp merges Qwen3.8-Flash-Next MTP support
llama.cpp merged PR #29761 adding multi-token prediction (MTP) speculative-decoding support for the Qwen3.8-Flash-Next checkpoint, with a reported ~55% decode-throughput gain on DGX Spark.
Why it matters
Free speedup for an already-shipped 125B open-weight checkpoint — better use of the same model, not a new model release.
What you could build with it
An indie developer could ship a one-click local serving bundle that downloads a Qwen3.8-Flash-Next GGUF with MTP tensors, launches llama-server with speculative-decoding flags, and serves a chat UI — turning a multi-GPU workstation into a fast private coding assistant.
Does it hold up?
Evidence-backed but setup-dependent: the PR author measured 28.36 to 43.88 tok/s (+55%) on DGX Spark with an IQ4_XS quant; community numbers (24.2 to 29.3 t/s on a Radeon 8060S, 43.8 to 73.5 t/s on an M5 Max) confirm real gains, while follow-up issues (MTP draft GGUFs failing to load on build b11058) show the path is still maturing.
Built with llama.cpp merges Qwen3.8-Flash-Next MTP support
- ik_llama.cpp PR #2369: qwen4exp MTP (NextN) self-speculative decoding supportgithub · Community fork PR enabling MTP self-speculative decoding for the qwen4exp architecture, building on earlier MTP work.
- How to Double Qwen 3.8 27B Text Generation Speedgeeky-gadgets.com · Third-party writeup showing MTP in llama.cpp lifting Qwen 3.8 from 7.9 to 17.1 tokens/sec on compatible hardware.
- Running Qwen3.8-Flash-Next 125B on Three RTX 3090s at 80 Tokens per Secondn1n.ai · Builder post on custom CUDA/C++ patches to llama.cpp reaching 80 tok/s for the 125B checkpoint on three RTX 3090s.
- ikawrakow/ik_llama.cppgithub · Community llama.cpp fork with qwen4exp MTP (NextN) self-speculative decoding support.
- ucicelos/flashnext-hybridgithub · 2 ★ · Build combining upstream llama.cpp qwen4exp support with an MTP draft-head port and tuned kernels.
- SergiioB/intel-arc-pro-b70-inference-cookbookgithub · Inference cookbook with a Qwen3.8-Flash-Next llama.cpp recipe reporting +42% decode from MTP.
Learn more
First spotted on github: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.