Radar / AI infrastructure / tpu-megakernels (Inferact)
tpu-megakernels (Inferact): open fused kernels that beat GB200 on Kimi K3
Open-source fused 'megakernels' for Google TPU v7 that serve Moonshot's Kimi K3 at 709 tokens/sec (and Qwen3.8-27B at 1,515 tok/s) — ~57% faster than the same model on 16x NVIDIA GB200 GPUs, with unchanged accuracy.
Why it matters
First credible public evidence that TPU inference software can beat NVIDIA's flagship GB200 on open Chinese MoE models; one hand-fused Pallas kernel collapses hundreds of scheduled kernels, making TPU a first-class, cheaper serving target for Kimi K3/Qwen3.8-27B. Built by vLLM co-founders in partnership with Google Cloud.
What you could build with it
An inference team or cloud-cost engineering shop can port the megakernel-style fused decode approach to other open MoEs (DeepSeek V4.x, GLM-5.3) on TPU v7 and sell 'GB200-class throughput at TPU pricing' — Inferact's kernels are currently Kimi K3 / Qwen3.8-27B-specific, so the generalization work is open.
Does it hold up?
Too early to judge. The code and engineering blog are real and detailed; accuracy parity (94.4% GPQA-Diamond on both platforms) was verified in Inferact's own runs, but all numbers are vendor-run with no independent production reproduction yet.
Built with tpu-megakernels (Inferact)
- RuntimeWire: Inferact's TPU megakernel hits 709 tokens per second on Kimi K3article · Independent coverage of the Sep 23 release; notes vLLM co-creator Woosuk Kwon's role and that production-scale concurrency/cost validation is still pending.
- vLLM project blog: serving Kimi K3 with DSpark speculative decodinggithub · vLLM team's Kimi K3 serving writeup (370 tok/s on GB300 with DSpark) — the NVIDIA-side baseline Inferact's TPU result is measured against.
- AI Market Watch: TPU v7 beats GB200 by 57% on Kimi K3article · Coverage stressing the software-not-silicon angle and Google Cloud's joint engineering partnership behind the open release.
- amp-pbc/kimi-k3-mi355xgithub · Community recipes for serving Kimi K3 on AMD MI355X GPUs — third-party serving work around the same model.
- xysheng-amd/amd-rl-runbookgithub · RL runbook including Kimi-K3-DSpark speculator training on AMD MI350; shows the community adapting Kimi K3 serving across hardware.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69ai-memory: long-term memory for agent coding CLIsRust solution for long-term memory for agent coding CLIs, facilitating handoff between different agents and sessions.infra · JEV 0.68DeepSeek open-sources Ascend infrastructure stackDeepSeek published Ascend-optimized versions of its NVIDIA-proven infra components — TileLang Ascend, DeepGEMM-Ascend…infra · JEV 0.65Hindsight: agent memory that learnsAgent memory system built for learning over time: retain/recall/reflect operations with SOTA scores on the LongMemEval…infra · JEV 0.63
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.