Radar / AI infrastructure / DeepGEMM

DeepGEMM: DeepSeek's open GPU BLAS kernel library

DeepSeek's DeepGEMM, a clean and efficient open-source BLAS kernel library for NVIDIA GPUs (up to 1550 TFLOPS on H800), is trending again as Chinese labs race to ship hardware-portable compute toolkits - the same week DeepSeek and Huawei released an Ascend-ported DeepGEMM variant with 1:1 API parity.

trendinginfra · routerJEV traction 0.77deepseek-ai/DeepGEMM · 8,803 ★added 2026-10-07
Open DeepGEMM →View on the radar

Why it matters

DeepGEMM became the de-facto reference for FP8 matrix multiplication in open LLM inference stacks; its API-parity port to Huawei Ascend chips shows kernel-level portability is becoming the battleground for AI hardware.

What you could build with it

An ML infrastructure team could adopt DeepGEMM's FP8 grouped GEMM kernels to cut inference cost on Hopper GPUs, then evaluate the Ascend-ported variant to run the same serving stack on Huawei hardware without rewriting kernel code.

Does it hold up?

Production-grade in FP8 LLM inference, used across open inference stacks for grouped GEMM; requires SM90/SM100 GPUs and a CUDA 12.9+ C++20 toolchain, so it is for kernel-level engineers rather than application developers.

Built with DeepGEMM

Learn more

official docs →

First spotted on github: source.

More AI infrastructure

Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Context Mode: context-window optimization for AI coding agentsMCP server for AI coding agents that sandboxes verbose tool output and persists session memory, cutting context use by…infra · JEV 0.74StrataOpen-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single consumer GPU with 12GB+…infra · JEV 0.74Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71claude-memCaptures tool-use observations across agent sessions, compresses them with AI into SQLite + Chroma hybrid search, and…infra · JEV 0.7

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.