Victoria + Maple — 44%-pruned Qwen3.8-Flash-Next and a Canada-first fine-tune
Community release by rmonsurate: Victoria is Qwen3.8-Flash-Next with 44% of experts cut (512→288/layer via REAP) and retrained at 4-bit (NVFP4) with a trained draft head — 70.0% Terminal-Bench 2.1 at 280 tok/s on a single NVIDIA B300; Maple is a Canadian-locale fine-tune on top.
Why it matters
A rare community compression of a frontier MoE that keeps ~80% of agentic coding ability while fitting on one GPU with 2.08x speculative speedup — with reproducible evals and full training curves published.
What you could build with it
An indie hacker or research lab with a single B300 can run a real agentic-coding fleet on one box — pair Victoria's GGUF build with an agent harness for offline, no-API-cost coding work.
Does it hold up?
Too early to judge — evals are self-reported on the author's own B300 setup with no independent reproduction yet.
Built with Victoria + Maple
- r/LocalLLaMA release thread (70 upvotes)reddit · Subreddit activity tracker confirms the New Model release thread for Victoria + Maple in the latest cycle; author answers setup questions in comments.
- Maple (Canada-first fine-tune)huggingface · Companion model: Victoria fine-tuned to default to Canadian sources/answers (62.9% official Canadian-source citation vs 6.0% before).
- rmonsurate/llama.cppgithub · llama.cpp fork (branch qwen4exp-draft-mtp) to build draft-head support from source.
Learn more
First spotted on hf: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.