Radar / AI models / SWE-Sweep

SWE-Sweep: benchmark for whether AI finds bugs without being told what's wrong

New open benchmark (MIT) from Meta, Stanford, Harvard and UW researchers testing whether models find and fix real bugs on their own across 100 repos and 22 languages — 4,000 real GitHub bugs.

new launchmodels · compositeJEV traction 0.64facebookresearch/swe-sweep · 10 ★added 2026-10-01
Open SWE-Sweep →View on the radar

Why it matters

Measures proactive bug-finding rather than fix-the-pointed-issue coding, and the first results are sobering: every model tested fixes under 5% of bugs, a hard baseline for autonomous coding claims.

What you could build with it

A research team building coding agents could adopt SWE-Sweep as the pre-merge gate for their agent's autonomous repair skills, turning 'finds bugs on its own' from a demo claim into a measured number.

Does it hold up?

Too early to judge — launched today; the headline numbers (best setup fixes under 5% of bugs, best model costing $7,230 per run) are author-reported and the harness has no independent reproductions yet.

Learn more

official docs →

First spotted on article: source.

More AI models

ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.