DoubtBench – does Jev know when humans disagree?
First System One benchmark scored against human uncertainty rather than a single right answer: ~7,500 questions built from NVIDIA HelpSteer2's multi-annotator ratings, scoring both accuracy and whether a decision model's probabilities track human disagreement.
Why it matters
Moves decision-model evaluation from 'right answer' to calibrated uncertainty — exactly what you need before trusting a probability threshold to automate real decisions.
What you could build with it
A team building decision-model features could run DoubtBench to check whether their model's confidence actually tracks human disagreement before wiring thresholds to auto-approve flows; the concrete result is uncertainty estimates trustworthy enough to route borderline cases to humans.
Does it hold up?
Too early to judge broadly — repo created 2026-10-04, 2 stars, and only four models scored so far (Claude Haiku 4.5 verbalized leads at 69.2, Jev 1.13 at 67.8, vs 90.4 human ceiling on the validation split of 7,455 questions). The methodology and figures are published in the README.
Learn more
README with leaderboard and methodology →
First spotted on github: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.