Radar / AI security / LiveNerf

LiveNerf: a pre-registered benchmark for post-release model drift

Open-source, pre-registered 30-day benchmark that detects whether Claude Opus 5.5 quietly gets worse after launch.

trendingsecurity · complianceJEV traction 0.32ninjahawk/livenerf · 903 ★added 2026-09-30
Open LiveNerf →View on the radar

Why it matters

The first day-0-baseline instrument for the 'was it nerfed?' question: frozen prompts, pinned CLI, exact graders, and a locked 78-question panel measured daily against launch-week baseline with clustered standard errors, on the UK AISI Inspect framework. Hit the HN front page today (694 points) because nobody has ever had a clean launch-day baseline before.

What you could build with it

A researcher or AI-audit startup can run livenerf-style pre-registered drift benchmarks on any frontier model to produce launch-day baselines, turning model-degradation complaints into evidence; a platform team can adopt its frozen-panel method to monitor their own hosted model quality.

Does it hold up?

Too early to judge: the 30-day measurement series is on day 6 of baseline collection and the first results row lands after day 20 (~Oct 14). The methodology is pre-registered and validation-passed, but no verdict on Opus 5.5 drift exists yet.

Built with LiveNerf

Learn more

technical deep dive →

First spotted on github: source.

More AI security

Cloudflare security-audit-skill: multi-phase security audits for coding agentsA coding-agent skill that runs multi-phase security audits with independently verified, machine-checkable findings.security · JEV 0.63Sandlock 0.8.8: deferred commit for agent sandboxesThe process-based Linux AI sandbox (no container, no VM) ships deferred commit: every run returns a changeset of what…security · JEV 0.56ClawSecure: free security scanner for OpenClaw AI agent skillsFree scanner that audits OpenClaw agent skills for vulnerabilities; the maker's audit of 2,890+ OpenClaw skills found…security · JEV 0.5Google DeepMind SynthID BioWatermarking technology that embeds an imperceptible, verifiable signature into AI-designed protein sequences and…security · JEV 0.49GitHub Security Lab Taskflow Agent: autonomous LLM fuzzing for C/C++Open-source agent that automates the full fuzzing lifecycle - entry-point discovery, harness generation, AFL++ runs…security · JEV 0.47LLM Agents Can Easily Tamper With Their Own TracesAn empirical study (arXiv 2026-09-24) showing that all tested local coding-agent harnesses except Muse Code allowed…security · JEV 0.4

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.