Radar / AI infrastructure / Janus
Janus
Single Go binary that runs GGUF models locally via llama.cpp with Vulkan (AMD/Intel/NVIDIA) or CPU fallback, exposing an OpenAI-compatible API.
Why it matters
Zero-dependency local LLM server (no Python, Docker, or Ollama) with hot-swappable models, thinking-model support, and chat-template auto-detection — built for wiring local models into Cursor/Cline.
What you could build with it
An indie developer could ship a privacy-first desktop app that bundles Janus as its inference sidecar, giving users a local AI assistant that works offline with any downloaded GGUF model and no cloud bill.
Does it hold up?
Too early to judge — Show HN from Oct 1 with thin evidence; commenters flag Vulkan overhead on Intel and missing benchmarks, so real-world performance is unverified.
Built with Janus
- Show HN discussion: Janus – Go binary that runs GGUF models via Vulkanhn · Launch thread (50+ points) where early readers probe Vulkan overhead on Intel hardware and ask for benchmarks against vLLM, SGLang and ExLlama.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69ai-memory: long-term memory for agent coding CLIsRust solution for long-term memory for agent coding CLIs, facilitating handoff between different agents and sessions.infra · JEV 0.68Ai2 Olmo-Core 3: open MoE training stackThe Allen Institute for AI released Olmo-Core 3, a redesigned open training framework for mixture-of-experts models…infra · JEV 0.67
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.