Radar / AI infrastructure / DwarfStar 4 (ds4)
DwarfStar 4 (ds4): Redis creator's C inference engine with asymmetric 2-bit quantization
antirez's DwarfStar 4, a narrow C inference engine (Metal/CUDA/ROCm) with asymmetric quantization that compresses only routed MoE experts to 2-bit while keeping shared paths precise; trending Oct 2-3 for running DeepSeek V4.1 Flash at 2-bit with ~790 tok/s prefill on a 128GB Mac.
Why it matters
A purpose-built engine beats generic GGUF runners for its target models: near-frontier open weights running at usable speed on a single high-memory consumer machine.
What you could build with it
A privacy-conscious consultancy could run a 2-bit quantized DeepSeek V4.1 Flash on a single high-memory workstation to process client documents offline without any cloud API costs.
Does it hold up?
Usable today but narrow: verified as the only Apple Silicon route for V4.1 Flash, with hard published numbers (790 tok/s prefill). It is deliberately not a generic runner — only a handful of DeepSeek/GLM models are supported, and it needs a 128GB machine.
Built with DwarfStar 4 (ds4)
- Antirez demos DeepSeek V4.1 Flash at 2-bit on DwarfStaraiweekly · Walkthrough of 2-bit V4.1 Flash on ds4: 790.2 tok/s prefill and 39.4 tok/s generation on a 128GB M5 Max Mac.
- DwarfStar compresses frontier models to run them on local machinesruntimewire · Reports the Oct 2 attention wave and Sanfilippo's thesis that AI is too critical to be only a provided service; flags the high hardware bar.
- DeepSeek V4.1 Flash Guide: Pricing, Benchmarks, Local Setupcodersera · Independent guide verifying ds4 (ds4.1flash branch) as the only confirmed Apple Silicon route for running V4.1 Flash locally today.
- JordiPosthumus/dwarf-star-gategithub · Community-built local DS4 gateway for Macs and mixed fleets with session-affinity routing and telemetry.
Learn more
First spotted on article: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71DeepSeek open-sources Ascend infrastructure stackDeepSeek published Ascend-optimized versions of its NVIDIA-proven infra components — TileLang Ascend, DeepGEMM-Ascend…infra · JEV 0.7Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.