Colibri: pure-C engine that runs 744B-2.8T MoE models on consumer hardware
A zero-dependency pure-C inference engine that treats storage, RAM and VRAM as one hierarchy, streaming MoE experts from disk so frontier models run on ordinary PCs.
Why it matters
Runs GLM-5.2 (744B) on a 25GB-RAM consumer machine and Kimi K3 (2.8T)-class checkpoints, with a web dashboard showing per-expert routing heat: no GPU required.
What you could build with it
An indie research lab could serve a 744B open-weight MoE from a single workstation with a fast NVMe drive as an internal API for experiments, avoiding cloud GPU rental for low-throughput eval workloads.
Does it hold up?
Demonstrably installable: honestly reported numbers are slow (0.05-1 tok/s depending on hardware) but real; nine model families supported, Apache-2.0 licensed. Best viewed as a research platform for multitier inference, not a production server.
Built with Colibri
- 5 GitHub Repos Exploding Right Now (October 2026)YouTube (Nex/Archie/Orbi) · Counts down Colibri at #5: runs huge MoE models (up to GLM-5.2 744B) on an ordinary PC by streaming experts from disk, pure C, no GPU.
- Colibri, Pure-C Inference Engine Running 2.8T MoE on Consumer HardwareSINGULISM · Writeup of the 'AI memory multitiering' approach, the per-layer LRU expert cache, and the multitier memory hierarchy.
- GitHub - JustVugg/colibri on daily.devdaily.dev · Developer discussion post with honest numbers: 0.05-0.1 tok/s cold on WSL2, ~1 tok/s on Apple M5 Max.
Learn more
First spotted on github: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71JetBrains Mellum 2.1: 12B MoE coding model trained for agentsJetBrains released Mellum 2.1, a 12B MoE model (2.5B active parameters, 128K context) under Apache 2.0, trained mainly…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.