Radar / AI infrastructure / DeepEP-Ascend
DeepEP-Ascend
DeepSeek's high-performance expert-parallel communication library for MoE training and inference on Huawei Ascend NPUs.
Why it matters
It ports DeepSeek's battle-tested NVIDIA communication stack to Huawei Ascend NPUs with an API aligned to the original DeepEP, giving labs a credible CUDA-free path for large-scale MoE training.
What you could build with it
An ML infrastructure startup could adopt DeepEP-Ascend alongside the companion DeepGEMM-Ascend kernels to offer managed MoE training and inference clusters on Huawei Ascend hardware, giving enterprises a domestic, API-compatible alternative to NVIDIA-based stacks.
Does it hold up?
Early but promising. The README publishes real benchmark tables (dispatch 373-375 GB/s at EP8 on Ascend 950DT, about 90-95% of the payload bandwidth limit), but the repo is 2 days old with only 21 forks, and PP/Engram/Bucket features are still experimental with no reported production deployments yet.
Built with DeepEP-Ascend
- DeepSeek 开源昇腾基础元件,API 与辉达平台版本对齐article · Chinese-language breakdown of the Ascend toolchain release (citing 量子位), covering DeepEP-Ascend's MoE expert-parallel all-to-all dispatch/combine with FP8 dispatch and its API alignment with the NVIDIA DeepEP.
- DeepSeek partners with Huawei to develop chip programming toolsarticle · Reuters wire coverage of the Sept 30 announcement: DeepSeek open-sourcing Ascend compute and communication libraries with Huawei's support, plus the joint 128-chip Ascend 950 supernode.
- DeepSeek open-sources AI tools for Huawei chips to reduce reliance on Nvidia's softwarearticle · News writeup of the six-component release (DeepGEMM, DeepEP, FlashMLA, TileKernels, DeepSelect, TileLang), noting the familiar Python interfaces for developers moving from Nvidia to Ascend.
- deepseek-ai/DeepGEMM-Ascendgithub · 464 ★ · Companion Ascend port of DeepSeek's DeepGEMM matrix-multiplication kernels (BF16/FP8/FP4, MQA logits, MegaMoE), released Sept 29 with up to 99.8% of hardware limit on dense GEMM.
Learn more
First spotted on github: source.
More AI infrastructure
Modular open-sources MAX inference server, Mojo stdlib and accelerator kernelsThe core of Modular's unified AI deployment stack is now open: the MAX inference server (OpenAI-compatible endpoints)…infra · JEV 0.76OmniRouteFree MIT AI gateway: one endpoint, 290+ providers and 500+ models with quota-aware auto-fallback.infra · JEV 0.75Microsoft releases 301,000 Copilot coding-agent tracesMicrosoft open-sourced 301,026 GitHub Copilot coding-agent sessions (9.3M model calls, 8.7M tool calls) with timings…infra · JEV 0.71Cloudflare Agents Week: Sandboxes GA, 50K concurrent Workflows, Managed OAuth for agentsA dozen agent-infrastructure launches in one week: persistent Linux Sandboxes (GA) with real shell/filesystem/state…infra · JEV 0.69ai-memory: long-term memory for agent coding CLIsRust solution for long-term memory for agent coding CLIs, facilitating handoff between different agents and sessions.infra · JEV 0.68Ai2 Olmo-Core 3: open MoE training stackThe Allen Institute for AI released Olmo-Core 3, a redesigned open training framework for mixture-of-experts models…infra · JEV 0.67
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.