Radar / Ideas / DecideForge: one-command eval and…

DecideForge: one-command eval and distillation for decision models

daily ideamoderateJEV confidence 0.622026-10-08
OutcomeA team ships a domain-calibrated 2B decision model matching frontier judgment accuracy at roughly 1/50th the inference cost, in a single day.

The problem

Dozens of open decision models now claim to beat Jev — Liquid d1, Cloudflare Clef, Strands Decider 2B, the kev family, ARC-1, Gutsy — each measured on a different yardstick (Decision Index vs JevBench vs vendor-run suites), and practitioners have no shared, cheap way to verify a claim or distill frontier judgment into a small model trained on their own data. Teams either pay per-call API prices for every judgment or trust self-reported scores; one survey of the wave concluded no independent shared leaderboard settles the question.

The idea

A CLI plus hosted runner: point it at any Jev-compatible endpoint or Hugging Face checkpoint and it runs a frozen suite (the Decision Index kit's 43-benchmark panel) plus DoubtBench-style disagreement calibration checks plus the team's own labeled CSV, then produces (a) a signed, reproducible scorecard and (b) a distilled kev-family 0.8–4B checkpoint fine-tuned on the team's labels. Hugging Face's RL environments on the Hub supply task environments for agentic-decision fine-tuning; LiteLLM Lens-style trace analysis pinpoints exactly where the small model diverges from the frontier judge.

Why now

kev (jaredpalmer/kev) gives the open trainable Jev-like target; Hugging Face started hosting RL environments on the Hub on Oct 7; DoubtBench (Oct 5) adds disagreement-aware calibration eval; the Decision Index 0.2.1 suite is frozen and versioned; and the flood of 'beats Jev' claims from Liquid, Cloudflare, and AWS creates the demand for an independent verifier.

What it combines

Four radar capabilities: kev-open-trainable-jev-like-family-of-small-decision-models-on-qwen3-5-3-8 (the trainable target), hugging-face-hosts-rl-environments-on-the-hub (shared task environments), doubtbench (calibration under human disagreement), and litellm-lens (trace-level failure attribution). The mix is a closed train → test → trust loop; separately they are a repo, a hosting feature, a benchmark, and a tracing tool — together they end with a deployable fine-tuned checkpoint no single piece produces.

MVP

Build first: a CLI wrapping the decision-index rerun kit, JevBench, and DoubtBench against a local GGUF/kev checkpoint, emitting a markdown scorecard, plus a LoRA fine-tune of kev-0.8B on the user's labeled CSV. Deliberately skip: the hosted runner, HF RL-environment integration, and the badge program.

Distribution

B2B2C: sell hosted evaluation to inference providers (OpenRouter, Baseten) and agent platforms as 'verified decision model' badges; freemium CLI for builders with paid per-run scorecards; a marketplace of certified community models with revenue share.

Why it wins

EleutherAI's lm-evaluation-harness and UK AISI Inspect evaluate chat/capability, not calibrated typed decisions; JevBench and the Decision Index score models but do not distill or train. DecideForge is the only loop whose output is a deployable fine-tuned decision model, not just a number.

Risks

Benchmark gaming and vendor disputes over methodology. De-risk by copying the Decision Index's own model: frozen versioned suites, maintainer-style reruns, and signed public scorecards so every claim is reproducible by anyone.

Build it with

Repo to start from

decideforge/decideforge — one-command decision-model workbench: frozen suite runner, DoubtBench calibration checks, signed scorecard publisher, and kev LoRA distillation on user labels.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.