D2K-Bench: can LLM agents turn expert designs into efficient GPU kernels?
Qwen-team diagnostic benchmark (26 tasks, 85 workloads) measuring how well LLM agents turn expert design guidance into efficient GPU kernels, with code released alongside the paper.
Why it matters
First benchmark to separate design discovery from implementation in agent-written GPU kernels, showing expert guidance lifts combined scores from 57 to 70/100 across frontier models.
What you could build with it
A GPU cloud provider could adopt D2K-Bench's task suite to score how well different agent harnesses write efficient CUDA kernels on its hardware, publishing a leaderboard that steers customers toward the best-performing stacks.
Does it hold up?
Too early to judge — paper and code released this week; no independent runs of the benchmark are public yet.
Learn more
First spotted on github: source.
More AI models
OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.