Radar / AI infrastructure / nestrilabs/virtio-nvgpu
nestrilabs/virtio-nvgpu
Experimental virtio device giving KVM virtual machines near-native NVIDIA GPU access; a March 2026 repo that hit the HN front page today with renewed buzz.
Why it matters
Near-native GPU performance inside VMs without PCI passthrough complexity; directly relevant to anyone self-hosting GPU workloads (inference servers, training rigs) in virtualized environments.
What you could build with it
Run vLLM inference servers in KVM VMs with near-native GPU access via virtio-nvgpu: isolate multi-tenant model serving without giving up performance.
Does it hold up?
Experimental but promising: one community benchmark supports the near-native claim, though the project is explicitly early-stage and Nvidia support in nesbox still needs funding.
Built with nestrilabs/virtio-nvgpu
- virtio-nvgpu LLM performance benchmark: 95% bare-metal speed — dev.toarticle · Independent benchmark: Ollama Llama 3 8B inference in a KVM guest hits 115.2 tok/s on RTX 4090 — 95.6% of bare metal — with a 30% operational-latency win over PCI passthrough for multi-tenant serving.
- Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest — WeSearch/HNarticle · Summary of the HN front-page story: forwards NVIDIA kernel-driver ioctls at the driver ABI level so the guest runs NVIDIA's own unmodified user-mode drivers.
- nestrilabs/nesboxgithub · 33 ★ · nestrilabs' lightweight microVM for cloud streaming whose NVIDIA guest path runs through virtio-nvgpu; README documents measured driver ABI profiles and per-card sharing limits.
- reindertpelsma/nvkvm-pvgithub · 75 ★ · Independent paravirtual NVIDIA GPU project for KVM guests (unmodified CUDA/PyTorch/Vulkan at host parity, no passthrough or vGPU license) solving the same problem class.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69ai-memory: long-term memory for agent coding CLIsRust solution for long-term memory for agent coding CLIs, facilitating handoff between different agents and sessions.infra · JEV 0.68DeepSeek open-sources Ascend infrastructure stackDeepSeek published Ascend-optimized versions of its NVIDIA-proven infra components — TileLang Ascend, DeepGEMM-Ascend…infra · JEV 0.65Hindsight: agent memory that learnsAgent memory system built for learning over time: retain/recall/reflect operations with SOTA scores on the LongMemEval…infra · JEV 0.63
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.