Radar / AI infrastructure / Rapid-MLX
Rapid-MLX: Apple Silicon inference server built for coding agents
Open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX; claims up to 4x decode speed over mlx-lm with a focus on reliable tool calling for coding agents; ~3,962 GitHub stars, v0.6.14 shipped this week, tested against 15+ agent frameworks.
Why it matters
A coding agent built on either the OpenAI or Anthropic SDK can point at a local Mac server with no client rewrite; it targets the part of local inference that breaks first under agent workloads — tool calling.
What you could build with it
An indie developer could run a fully local coding-agent setup on a Mac Studio with Rapid-MLX as the backend, keeping proprietary code off cloud APIs while using familiar OpenAI/Anthropic SDK agents at zero per-token cost.
Does it hold up?
Early but promising: an honest agent-support matrix with PASS/XFAIL results, a community fork adding resumable downloads, and independent coverage cautioning that the 4x speed claims still need third-party reproduction.
Built with Rapid-MLX
- SmartChunks: Rapid-MLX brings fast OpenAI and Anthropic APIs to Apple Siliconarticle · Hands-on-oriented writeup covering the dual API surface, Apache 2.0 license, and the honest 1.5x-vs-4x benchmark framing; flags that the speed claims still await third-party reproduction on independent Macs.
- Community PR adds Rapid-MLX to awesome-generative-ai local LLM sectiongithub · Community list PR noting ~1,462 stars at filing (now ~3,962), v0.6.14 shipped this week, 15+ agent frameworks tested (LangChain, LiteLLM, Aider, Codex, OpenHands), and Day-0 support for DeepSeek V4 and Qwen 3.5/3.6.
- raullenchai/rapid-mlx-dsh-providergithub · Native DeepSeek-Harness LLM adapter for Rapid-MLX, verified against dsh 0.1.0-rc.8; reads model facts from the server's /v1/models so switching models needs no re-setup.
Learn more
First spotted on github: source.
More AI infrastructure
DeepGEMM: DeepSeek's open GPU BLAS kernel libraryDeepSeek's DeepGEMM, a clean and efficient open-source BLAS kernel library for NVIDIA GPUs (up to 1550 TFLOPS on…infra · JEV 0.77Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74StrataOpen-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single consumer GPU with 12GB+…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.