kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8
A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision architecture: yes/no, multiple-choice and rating questions answered as calibrated probabilities in one forward pass, with a TypeSafe System One-compatible API that works with the official Python SDK unchanged.
Why it matters
The first open, trainable reproduction of a typed-decision architecture with published frozen evals: Kev-9B scores 0.837 out-of-domain accuracy on the locked test, and the Qwen3.5 port reportedly cost about $95 of H100 time.
What you could build with it
A product team running LLM-as-a-judge loops can replace per-decision generator calls with a self-hosted kev server: one document plus N typed questions in a single forward pass, with calibrated confidence to set thresholds on.
Does it hold up?
Strong early evidence: the repo ships trained weights, training recipes, frozen eval suites and a live Hugging Face Space; the 0.8B runs on a laptop and the TypeSafe SDK works against a kev server unchanged, the closest to plug-and-play any Jev reproduction has come.
Built with kev
- Kev: An Open-Source Decision Model That Answers, Not Generatesblog · Deep dive arguing kev shows small calibrated decision models beating 400B chat models on latency and cost for production classification.
- Kev-9B model card and evalsgithub · Independent model-card repo with kev-9B's frozen-item results: 0.837 locked-test accuracy, 0.243 Brier, full trial hashes.
- awesome-jev reproduction roundupgithub · Community-curated list of Jev reproductions comparing kev against nimble, litjev, ruling and others.
Learn more
First spotted on github: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66Claude Sonnet 5.5: 30% faster, up to 30% cheaper per taskAnthropic's new mid-tier: 70.6% on Terminal-Bench 4.0 (vs 10.3% for Sonnet 5), 46.2% FrontierCode 1.1 max-effort…models · JEV 0.65
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.