Radar / AI security / SafeActBench

SafeActBench: a 656-case benchmark for evidence-grounded tool-using agents (NUS)

NUS-led benchmark testing whether agents' consequential actions are backed by evidence established before acting, across 6 operational domains and 5 protocols, with deterministic scoring and no LLM judge.

new launchsecurity · complianceJEV traction 0.43caoshidong66/safeact · 6 ★added 2026-10-07
Open SafeActBench →View on the radar

Why it matters

Exposes a blind spot that endpoint checks miss — agents reaching the right outcome via unsupported actions — and finds ≥95% static judgment accuracy collapsing to ≤52% in interactive execution.

What you could build with it

A platform team shipping customer-facing agents could add SafeActBench to its CI gate, blocking agents that act on unverified evidence; concrete result: fewer wrongful refunds, config changes, or device actions in production.

Does it hold up?

Too early to judge: benchmark and harness just released; no third-party validations of the reported model scores yet.

Built with SafeActBench

Learn more

technical deep dive →

First spotted on hf: source.

More AI security

OpenAI, Anthropic and Google DeepMind jointly unveil cyber-focused safety models and safeguardsThe three rival labs disclosed weeks of behind-the-scenes coordination and jointly released a set of cyber-focused AI…security · JEV 0.73OpenAI textGrain watermarking for ChatGPT and Codex in the EUOpenAI will roll out its invisible textGrain watermarking system to ChatGPT/Codex users in the EU over the coming…security · JEV 0.65Cloudflare security-audit-skill: multi-phase security audits for coding agentsA coding-agent skill that runs multi-phase security audits with independently verified, machine-checkable findings.security · JEV 0.63LiveNerf: a pre-registered benchmark for post-release model driftOpen-source, pre-registered 30-day benchmark that detects whether Claude Opus 5.5 quietly gets worse after launch.security · JEV 0.61Sandlock 0.8.8: deferred commit for agent sandboxesThe process-based Linux AI sandbox (no container, no VM) ships deferred commit: every run returns a changeset of what…security · JEV 0.58OpenAI disrupts Moonshot-linked 'adversarial distillation' campaignOpenAI disclosed it shut down a coordinated campaign by 15,000+ users to extract hidden model reasoning, attributing a…security · JEV 0.58

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.