MAI-Code-1.1-Flash downloadable for on-device coding
Microsoft released its 137B-total / 6.8B-active coding model as a downloadable ~3-bit quantized build with full 256K context and zero inference charges for local calls.
Why it matters
Brings a frontier-class Copilot coding model to on-device hardware (>120GB RAM recommended) at ~80% memory reduction with SWE-Bench Verified / Terminal-Bench parity claims — a free local option for eligible coding work.
What you could build with it
An indie developer could run the quantized MAI-Code-1.1-Flash on a local workstation as a free offline pair-programmer for boilerplate, refactors and test generation, sending only tricky architectural decisions to a cloud model.
Does it hold up?
Too early to judge the downloadable on-device build: it was announced Oct 7 and Copilot experimental access only rolls out 'by end of the month'. The 3-bit quant keeps full 256K context and matches the full model on SWE-Bench Verified/Terminal-Bench 2.1 per Microsoft — but all evidence is vendor-reported and >120GB RAM is the stated floor. No independent hands-on runs exist yet.
Built with MAI-Code-1.1-Flash downloadable for on-device…
- VS Code demo: shipping a real feature end-to-end with MAI-Code-1-Flashyoutube · Walkthrough of MAI-Code-1-Flash inside VS Code Copilot Chat: codebase exploration, building a feature, running and testing it, plus a cost-benefit breakdown. (Note: published for the 1.0 launch, but shows the on-device/Copilot router concept now extended to 1.1.)
- Neowin: Microsoft releases MAI-Code-1.1-Flash to compete with Chinese modelsneowin.net · 1.1 pricing detail: 0.25x premium-request multiplier for annual Copilot subscribers ($0.20/$0.02/$1.20 per 1M input/cached/output), 22% better on Terminal-Bench 2.1 in Copilot CLI, 15% better .NET tasks, native vision support — framed as Microsoft's answer to GLM-5.2 and Kimi K3.
Learn more
First spotted on article: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.