Radar / AI infrastructure / Magnitude
Magnitude: self-optimizing open-source inference engine for local agents
Open-source inference engine (YC S25) that profiles your machine, recommends the best open models for it, and tunes kernels on-device; claims up to 2x faster than llama.cpp with dynamic memory allocation for parallel agent sessions.
Why it matters
The first inference engine explicitly designed for local agents: long concurrent sessions get dynamic memory, model switching preserves tool use, and per-device kernel tuning reaches hardware-specific performance without hardware lock-in. 119 HN points / 54 comments on launch day.
What you could build with it
A product team or indie hacker can ship a zero-API-cost local agent runtime for privacy-sensitive customers (law firms, clinics, defense contractors): Magnitude picks and tunes the best model for each machine, so the product sells as a one-click install instead of a cloud subscription.
Does it hold up?
Too early to judge: launched today on HN with desktop apps (macOS/Windows/Linux) plus CLI and one-click agent connections (Pi, OpenCode, Hermes). No independent benchmarks beyond the founders' claims yet.
Built with Magnitude
- Launch HN thread: Magnitude (YC S25)hn · Founders' launch post with 119 points and 54 comments detailing on-device kernel tuning, dynamic memory, and the gap vs vLLM/llama.cpp for single-session agent workloads.
- agents-radar AI digest (2026-09-07) flagged +604 first-day starsgithub · Independent AI-trending digest noted magnitudedev/magnitude gaining 604 first-day stars as developers look to cut API costs by running small local models alongside agent tools.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69ai-memory: long-term memory for agent coding CLIsRust solution for long-term memory for agent coding CLIs, facilitating handoff between different agents and sessions.infra · JEV 0.68DeepSeek open-sources Ascend infrastructure stackDeepSeek published Ascend-optimized versions of its NVIDIA-proven infra components — TileLang Ascend, DeepGEMM-Ascend…infra · JEV 0.65Hindsight: agent memory that learnsAgent memory system built for learning over time: retain/recall/reflect operations with SOTA scores on the LongMemEval…infra · JEV 0.63
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.