CLM-8B: Open System One Model for Agent Action Scoring
First open Contrastive Language Model: it does not generate text, it scores candidate agent actions against a state and returns probabilities, up to 9x faster than Jev.
Why it matters
Apache-2.0 weights; a tiny 75 MB trainable head on a frozen Qwen3-8B encoder runs on a single GPU via vLLM and exposes a TypeSafe-compatible API, so Jev-written requests can be replayed; vendor claims SOTA verifier results on agentic coding benchmarks (87.6% Terminal-Bench 2.1, 81.6% DeepSWE on held-out subsets).
What you could build with it
A team whose agents choose between actions can score each candidate with a local CLM-8B server, cutting decision latency and per-call cost, and fine-tune the small head on its own approve/reject labels.
Does it hold up?
Too early to judge: independent hands-on reports do not exist yet; benchmark claims are vendor-reported. The TypeSafe-compatible API means existing Jev users can trial it with minimal changes.
Built with CLM-8B
- CLM-8B: a 9x faster System-1 decision model than Jevmedium · Engineering writeup comparing CLM-8B's contrastive scoring approach against Jev-style generative judges on decision latency.
- CLM-8B contrastive verifier to speed best-of-n ranking and reduce latencyblog · Technical writeup on using CLM-8B as a contrastive verifier for best-of-N ranking in agent loops.
- Contrastive-LM releases CLM-8B open model scoring actions up to 9x faster than Jevnews digest · Two-source digest of the CLM-8B release and its TypeSafe-compatible decision API.
Learn more
First spotted on github: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.