vLLM Decision 2.0
Open decision-model family (0.6B to 27B) for the vLLM Semantic Router: answers up to 64 questions about a single request in one forward pass, 63 ms on a single GPU, Apache-2.0 licensed, loadable with Transformers.
Why it matters
First open decision-model suite that claims #1 at its size on the Jev Decision Index (0.6B, 0.8B, 2B, 4B) and #3 overall at 27B — 2.5x the Decision 1.0 score at the same size, 20x faster than the Jev API for a single decision, sized down to sub-billion for request-level intelligence in the serving stack.
What you could build with it
An infrastructure team could place Decision 2.0 models at the front of an LLM gateway: a 0.6B model classifies each incoming request (domain, safety, modality) in milliseconds and routes it to the cheapest adequate model, cutting inference spend while keeping policy decisions auditable in one pass.
Does it hold up?
Too early to judge — announced today; vendor benchmarks claim top decision-index scores, but no independent evaluations exist yet.
Built with vLLM Decision 2.0
- Decision 2.0 announcement (vLLM LinkedIn)x · Official vLLM announcement of Decision 2.0: 0.6B–27B open models, 64 questions per request in one pass, Decision Index rankings, HF collection link.
- Introducing Decision (Decision 1.0 deep dive)github · Technical introduction to the Decision model family from the vLLM team: the Signal-Decision architecture, Decision Studio, fine-tuning tools, and the open Apache-2.0 licensing model.
- Decision Studio playgroundhf · Interactive playground for trying the Decision models (hosted on Hugging Face Spaces by the vLLM semantic-router org).
- vllm-project/semantic-routergithub · 6,008 ★ · The open semantic-router project the Decision models are built for: signal-driven routing with the MoM model family, Decision Studio, and fine-tuning tools.
Learn more
First spotted on article: source.
More AI models
ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28: a new text-to-speech architecture with inline…models · JEV 0.72VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesVoiceStudio is an open-source, fully-local voice AI studio: voice cloning from a 3-second sample, voice design, video…models · JEV 0.71DeepSeek V4-Flash official API: public beta with upgraded agent capabilities, Responses API and Codex supportDeepSeek's official V4-Flash is now a public beta API with massively upgraded agent capabilities - benchmark scores…models · JEV 0.69Mistral Forge: enterprise platform for training and continuously improving proprietary modelsMistral AI launched Forge (Sep 28) — a full-lifecycle model training platform (pre-training, SFT, DPO/ODPO, RL…models · JEV 0.69kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8A family of small decision models (0.8B to 27B) built on Qwen3.5/3.8 that reproduces Jev's prefill-only typed-decision…models · JEV 0.68GPT-Synopsys: OpenAI and Synopsys multi-year deal to build an AI model that operates EDA toolsOpenAI and Synopsys signed a multi-year agreement on Sep 30 to jointly develop GPT-Synopsys, a specialized model that…models · JEV 0.66
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.