DriftGuard: pager alerts when your pinned model silently gets worse
The problem
Teams pin model versions, but vendors change behavior under the same version string: Anthropic has published two 'nerf' postmortems (an infra bug, a default-effort change) and the 'was it nerfed?' debate never dies because nobody keeps a launch-day baseline. Existing observability watches your app's outputs, not the underlying model's capability.
The idea
Why now
LiveNerf just proved the methodology works — pre-registered panels, day-0 baselines, clustered standard errors — and its HN front-page run (694 points) shows the demand is visceral. Meanwhile kev open-sourced trainable Jev-like decision models, making daily grading 100x cheaper than frontier-judge evals, so continuous monitoring is now affordable.
What it combines
LiveNerf (pre-registered drift panels + Inspect harness: what to measure and how to prove it) + kev (open, trainable small decision models: cheap daily graders). The mix matters because LiveNerf without cheap graders is a one-off science project, and cheap graders without pre-registered panels are noise — together they make continuous, statistically honest drift detection a commodity.
MVP
One model, one panel (50 frozen coding prompts), daily re-runs against OpenAI + Anthropic endpoints, Slack alert on statistically significant drop. Deliberately skip: multi-model support, custom panels, dashboard — alerts first.
Distribution
B2B2C via AI gateways that already sit between teams and models (Portkey, Helicone, Maxim-style gateway platforms) as an add-on integration; paid per model-version monitored. Direct motion at platform teams that got burned by a silent downgrade.
Why it wins
Arize, WhyLabs, LangSmith and Langfuse monitor your application's outputs and prompt quality; none of them gives vendor-neutral, statistically pre-registered drift detection on the underlying model itself. DriftGuard measures the model, not your app.
Risks
Vendors could block synthetic probing or Goodhart the known panels. De-risk: rotate panel subsets per customer, keep panels private, and blend in natural-traffic sampling so probes are indistinguishable from real use.
Build it with
- LiveNerf: a pre-registered benchmark for post-release model driftPre-registered panel methodology and Inspect harness for statistically honest drift detection.
- kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8Open, trainable small decision models as cheap, self-hostable daily graders.
- Jev: TypeSafe AI's System One models for structured decisionsTyped judgments for grading with structured reasons instead of raw scores.
Repo to start from
driftguard-core — open-source frozen-panel runner + grader harness; the hosted scheduling, private panels and alerting are the paid layer.
Evidence
- HN front-page discussion: 'Livenerf: Has Opus 5.5 been nerfed yet?'
- production-genai-skills: quality monitoring playbook (no baseline = can't detect drift)
- LiveNerf repo: pre-registered 30-day model drift benchmark
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.