Radar / AI infrastructure / Strata
Strata
Open-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single consumer GPU with 12GB+ VRAM, via expert caching on GPU/RAM/CPU plus an n-gram lookup table on SSD.
Why it matters
Puts a 125B frontier-scale open model on a gaming rig at up to ~140 tokens/s on an RTX 3090, no cloud, zero per-token cost, with OpenAI/Anthropic-compatible local APIs.
What you could build with it
An indie developer could run a private frontier-class assistant entirely offline by pairing Strata's local OpenAI-compatible server with their existing coding agent, getting cloud-scale model quality with zero API costs and complete data privacy for sensitive projects.
Does it hold up?
Genuinely usable on modern hardware: community benchmarks and a byte-identical output guarantee across v0.1.39 show real engineering discipline. But it needs 32GB+ RAM, a 12GB+ GPU, and a 70GB download; older hardware paths are compile-checked, not tested.
Built with Strata
- Strata lets a 125 billion parameter model run on a gaming GPUarticle · Startup Fortune analysis of Strata's MoE expert-caching engineering, with an honest caveat that 3-bit quantization is meaningfully lossy and much of the online praise may be AI-amplified.
- r/LocalLLaMA thread: Strata vs llama.cpp technical comparisonforum · Community thread with a detailed side-by-side breakdown of Strata's vertical MoE offloading vs llama.cpp's horizontal layer offloading, plus skepticism on claimed benchmark speedups.
- Community benchmark: 2x RTX 5060 Ti 16 GB (PR #483)github · User-submitted benchmarks (56-63 tok/s decode, ~1108 tok/s prefill) merged into the repo's bench/results, with an agentic coding task comparing Strata's Flash-Next against Qwen3.8-27B on llama.cpp.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71claude-memCaptures tool-use observations across agent sessions, compresses them with AI into SQLite + Chroma hybrid search, and…infra · JEV 0.7DeepSeek open-sources Ascend infrastructure stackDeepSeek published Ascend-optimized versions of its NVIDIA-proven infra components — TileLang Ascend, DeepGEMM-Ascend…infra · JEV 0.7
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.