DecideForge: one-command eval and distillation for decision models
The problem
Dozens of open decision models now claim to beat Jev — Liquid d1, Cloudflare Clef, Strands Decider 2B, the kev family, ARC-1, Gutsy — each measured on a different yardstick (Decision Index vs JevBench vs vendor-run suites), and practitioners have no shared, cheap way to verify a claim or distill frontier judgment into a small model trained on their own data. Teams either pay per-call API prices for every judgment or trust self-reported scores; one survey of the wave concluded no independent shared leaderboard settles the question.
The idea
Why now
kev (jaredpalmer/kev) gives the open trainable Jev-like target; Hugging Face started hosting RL environments on the Hub on Oct 7; DoubtBench (Oct 5) adds disagreement-aware calibration eval; the Decision Index 0.2.1 suite is frozen and versioned; and the flood of 'beats Jev' claims from Liquid, Cloudflare, and AWS creates the demand for an independent verifier.
What it combines
Four radar capabilities: kev-open-trainable-jev-like-family-of-small-decision-models-on-qwen3-5-3-8 (the trainable target), hugging-face-hosts-rl-environments-on-the-hub (shared task environments), doubtbench (calibration under human disagreement), and litellm-lens (trace-level failure attribution). The mix is a closed train → test → trust loop; separately they are a repo, a hosting feature, a benchmark, and a tracing tool — together they end with a deployable fine-tuned checkpoint no single piece produces.
MVP
Build first: a CLI wrapping the decision-index rerun kit, JevBench, and DoubtBench against a local GGUF/kev checkpoint, emitting a markdown scorecard, plus a LoRA fine-tune of kev-0.8B on the user's labeled CSV. Deliberately skip: the hosted runner, HF RL-environment integration, and the badge program.
Distribution
B2B2C: sell hosted evaluation to inference providers (OpenRouter, Baseten) and agent platforms as 'verified decision model' badges; freemium CLI for builders with paid per-run scorecards; a marketplace of certified community models with revenue share.
Why it wins
EleutherAI's lm-evaluation-harness and UK AISI Inspect evaluate chat/capability, not calibrated typed decisions; JevBench and the Decision Index score models but do not distill or train. DecideForge is the only loop whose output is a deployable fine-tuned decision model, not just a number.
Risks
Benchmark gaming and vendor disputes over methodology. De-risk by copying the Decision Index's own model: frozen versioned suites, maintainer-style reruns, and signed public scorecards so every claim is reproducible by anyone.
Build it with
- kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8The open, trainable Jev-like decision-model family — the distillation target.
- Hugging Face hosts RL environments on the HubShared, versioned task environments for agentic-decision fine-tuning and eval.
- DoubtBench – does Jev know when humans disagree?Disagreement-aware calibration eval — catches overconfident small models the accuracy suites miss.
- LiteLLM Lens — AI agents that analyze agent traces inside the gatewayTrace-level analysis to attribute exactly where a distilled model diverges from the frontier judge.
Repo to start from
decideforge/decideforge — one-command decision-model workbench: frozen suite runner, DoubtBench calibration checks, signed scorecard publisher, and kev LoRA distillation on user labels.
Evidence
- The Decision Models Wave: Perplexity, Cloudflare, AWS vs Jev
- Jev Decision Index — Hugging Face Space
- JevBench — decision-model benchmark (GitHub)
- lm-evaluation-harness — EleutherAI (GitHub)
- decision-models-experiments — community comparisons (GitHub)
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.