Radar / AI infrastructure / sparse-memory-lm
sparse-memory-lm: a 21M LM with a 16.8M-row product-key memory on an SSD
Hobby research project (Apache 2.0) giving a 21M-parameter LM a 6.4B-parameter lookup table that matches a 114M dense model and runs with the table memory-mapped off an NVMe SSD at ~140 tok/s on a consumer RX 9070.
Why it matters
Shows memory layers can trade VRAM for cheap SSD lookups with honest lab-notebook evidence — plus Triton kernels that run unchanged on RDNA4, MI350X, and H100/H200.
What you could build with it
A researcher could adapt the product-key memory pattern to build small specialist models that read a large lookup table off disk instead of loading more weights into VRAM. A student on consumer hardware could reproduce the full experiment for a thesis on memory-augmented architectures, since the $70 training budget and kernels are all public.
Does it hold up?
Reproducible but research-grade: the author's lab notebook is unusually honest — from-scratch training gains hold, but retrofitting the table to Qwen3.5-0.8B failed to beat a dense add-on; runs on a consumer RX 9070, but it's a tiny-model experiment, not a product.
Built with sparse-memory-lm
- Interactive table explorerGitHub · The author's own live demo: click any word of a Wikipedia text to see which table entries the model reads, with a browsable map of the whole table.
- B-16M checkpoint on Hugging FaceHuggingFace · The trained 16.8M-row B-16M model with the 4-bit table (3.2 GB) published for download and local demo generation.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74StrataOpen-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single consumer GPU with 12GB+…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71claude-memCaptures tool-use observations across agent sessions, compresses them with AI into SQLite + Chroma hybrid search, and…infra · JEV 0.7
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.