Radar / AI models / D2K-Bench

D2K-Bench: can LLM agents turn expert designs into efficient GPU kernels?

Qwen-team diagnostic benchmark (26 tasks, 85 workloads) measuring how well LLM agents turn expert design guidance into efficient GPU kernels, with code released alongside the paper.

new launchmodels · compositeJEV traction 0.38QwenLM/D2K-Bench · 0 ★added 2026-10-08
Open D2K-Bench →View on the radar

Why it matters

First benchmark to separate design discovery from implementation in agent-written GPU kernels, showing expert guidance lifts combined scores from 57 to 70/100 across frontier models.

What you could build with it

A GPU cloud provider could adopt D2K-Bench's task suite to score how well different agent harnesses write efficient CUDA kernels on its hardware, publishing a leaderboard that steers customers toward the best-performing stacks.

Does it hold up?

Too early to judge — paper and code released this week; no independent runs of the benchmark are public yet.

Learn more

technical deep dive →

First spotted on github: source.

More AI models

OpenAI publishes 722 AI-generated math manuscripts from an unreleased modelOpenAI released 722 mathematical manuscripts in 372 result families to GitHub under Apache-2.0, produced by an…models · JEV 0.74Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.