Aleph Alpha Kolibri-1: 78B open-weight bilingual MoE for sovereign on-prem AI
Aleph Alpha released Kolibri-1, a 78.1B-total / 3.46B-active-parameter English-German mixture-of-experts model with up to 1M-token context and Apache 2.0 open weights on Hugging Face.
Why it matters
A European 'sovereign' open-weight model aimed at regulated sectors: trained in Germany and Finland under EU law, it offers frontier-adjacent bilingual quality at 3.5B-active compute cost, so ministries and industrial firms can run serious AI on their own hardware instead of sending sensitive documents to a foreign cloud.
What you could build with it
A European legal-tech startup could self-host Kolibri-1 to power an on-prem contract-review assistant for GDPR-bound law firms: client documents never leave the firm's own servers, while the 262K-token native context holds entire case files and the per-request reasoning-effort control trades response speed against analysis depth.
Does it hold up?
Servable today via `vllm serve Aleph-Alpha/Kolibri-1` plus the aleph-alpha-inference plugin, exposing an OpenAI-compatible server. Needs ~78GB GPU memory (2x A100-80GB, 2x H100, or 1x H200/B200/B300); independent quality verdicts are too early since all published benchmarks are vendor-run and no hosted provider serves it yet.
Built with Aleph Alpha Kolibri-1
- Hacker News front-page discussion of Kolibri (400+ points)Hacker News (via LAXIMA signal) · HN thread hit ~648 points with hundreds of comments debating the sovereignty claims, the honest benchmark reporting, and serving costs of the 78B/3.5B-active design.
- Aleph Alpha puts Kolibri's full weights on Hugging Face under Apache 2.0Startup Fortune · Developer-focused writeup of the release, including an HN reporter's measurement of ~170 tokens/sec running the FP8 checkpoint on a single Nvidia RTX Pro 6000.
- The Model Is Open, the Memory Bill Is NotDenny Sentinel · Independent analysis arguing Kolibri computes like a 3.5B model but still demands a 78GB-resident footprint, with concrete deployment guidance for agent builders.
- Community GGUF quantization for llama.cppHugging Face community · Day-one community GGUF quant (Q3_K_S) of Kolibri-1 for local inference with llama.cpp, one of several community quants (GGUF, MLX, NVFP4) published within 48h of release.
- Aleph-Alpha/aleph-alpha-inferencegithub · 19 ★ · Official inference package providing the Kolibri-specific reasoning and tool-call parsers required to serve Kolibri-1 through vLLM.
- mARTin-B78/aleph-alpha-inference-kolibri-1-dgx-sparkgithub · Community fork packaging the aleph-alpha-inference setup for running Kolibri-1 on Nvidia DGX Spark.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.