Radar / AI models / Dust

Dust: pretraining transformers without backpropagation

The first zeroth-order optimizer competitive with backprop at transformer pretraining — it perturbs activations per-token as a 'virtual population' instead of computing gradients.

new launchmodels · frontierJEV traction 0.28added 2026-10-07
Open Dust →View on the radar

Why it matters

If it scales, training could drop the differentiability requirement entirely, unlocking exotic architectures (external program calls in the loop) that backprop cannot train.

What you could build with it

A research lab or startup building custom neural hardware could use Dust to train architectures that are not differentiable — for example networks with an external program call in the loop — without writing custom autodiff, getting a workable gradient estimate from forward passes alone.

Does it hold up?

Too early to judge: the paper demonstrates parity with backprop up to 1B tokens on research-scale models, but the method explicitly needs substantially more compute than backprop today and has no production deployments yet.

Built with Dust

Learn more

technical deep dive →

First spotted on hn: source.

More AI models

Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.