Radar / AI infrastructure / Rapid-MLX

Rapid-MLX: Apple Silicon inference server built for coding agents

Open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX; claims up to 4x decode speed over mlx-lm with a focus on reliable tool calling for coding agents; ~3,962 GitHub stars, v0.6.14 shipped this week, tested against 15+ agent frameworks.

trendinginfra · routerJEV traction 0.59raullenchai/rapid-mlx · 3,962 ★added 2026-10-11
Open Rapid-MLX →View on the radar

Why it matters

A coding agent built on either the OpenAI or Anthropic SDK can point at a local Mac server with no client rewrite; it targets the part of local inference that breaks first under agent workloads — tool calling.

What you could build with it

An indie developer could run a fully local coding-agent setup on a Mac Studio with Rapid-MLX as the backend, keeping proprietary code off cloud APIs while using familiar OpenAI/Anthropic SDK agents at zero per-token cost.

Does it hold up?

Early but promising: an honest agent-support matrix with PASS/XFAIL results, a community fork adding resumable downloads, and independent coverage cautioning that the 4x speed claims still need third-party reproduction.

Built with Rapid-MLX

Learn more

technical deep dive →

First spotted on github: source.

More AI infrastructure

DeepGEMM: DeepSeek's open GPU BLAS kernel libraryDeepSeek's DeepGEMM, a clean and efficient open-source BLAS kernel library for NVIDIA GPUs (up to 1550 TFLOPS on…infra · JEV 0.77Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74StrataOpen-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single consumer GPU with 12GB+…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.