Darwin-180B-RSI: VIDRAFT's open 180B reasoning model now #1 on 10 HF leaderboards
Korean startup VIDRAFT's 180B MoE (built on Qwen3.8-Flash-Next) now leads 10 official Hugging Face leaderboards via recursive self-improvement — the model re-trains on its own verified solutions.
Why it matters
First model with perfect scores on Hugging Face's certified AIME 2026 and HMMT leaderboards, and the 'zero-token confidence' readout that scores an agent action's likelihood of success in 0.06s with no generated tokens.
What you could build with it
A research team could build a self-improving training pipeline inspired by Darwin's verify-then-retrain RSI loop to harden a domain model on checkable tasks. An indie developer could run the 4-bit GGUF locally to test its ZTC confidence readout as a cheap action-gating layer for agent tool calls.
Does it hold up?
Benchmark-strong, deployment-heavy: self-reported #1s across 10 official HF leaderboards, but a 336 GB bf16 footprint; the 4-bit GGUF (111 GB) with reported CPU-only 18-21 tok/s lowers the bar for serious tinkerers — independent verification still pending.
Built with Darwin-180B-RSI
- DEV: Open 180B model leads 10 official HF leaderboardsblog · Deep dive into the 10-leaderboard sweep and the ZTC zero-token judge demo that made HF Spaces of the Week.
- DEV: Darwin-180B-RSI tops legal benchmarks without legal trainingblog · Coverage of the LEXam/LEXam-hard #1 results (68.94, beating GPT-5's 62.65) and the recursive self-improvement training recipe.
- A self-taught AI never trained on law just topped a Swiss law-exam benchmarkblog · Deep dive into the LEXam/LEXam-hard results, the 4-sample majority-vote eval protocol, and how the scores compare to GPT-5 and Claude.
- Open 180B model tops five Hugging Face official leaderboardsMedium · Breakdown of the AIME/HMMT perfect scores, GPQA Diamond 94.44, and the MoE architecture behind the sweep.
Learn more
First spotted on hf: source.
More AI models
Mistral Large 4 (Le Chonk): 1.05T-param open-weight MoEMistral's new flagship: a 1.05T-parameter mixture-of-experts multimodal model with 49B active params and 1M-token…models · JEV 0.72ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.