Radar / AI security / SafeActBench
SafeActBench: a 656-case benchmark for evidence-grounded tool-using agents (NUS)
NUS-led benchmark testing whether agents' consequential actions are backed by evidence established before acting, across 6 operational domains and 5 protocols, with deterministic scoring and no LLM judge.
Why it matters
Exposes a blind spot that endpoint checks miss — agents reaching the right outcome via unsupported actions — and finds ≥95% static judgment accuracy collapsing to ≤52% in interactive execution.
What you could build with it
A platform team shipping customer-facing agents could add SafeActBench to its CI gate, blocking agents that act on unverified evidence; concrete result: fewer wrongful refunds, config changes, or device actions in production.
Does it hold up?
Too early to judge: benchmark and harness just released; no third-party validations of the reported model scores yet.
Built with SafeActBench
- caoshidong66/safeactgithub · 6 ★ · All 656 benchmark cases, 6 domain environments, deterministic evaluator and runner scripts; MIT.
Learn more
First spotted on hf: source.
More AI security
OpenAI, Anthropic and Google DeepMind jointly unveil cyber-focused safety models and safeguardsThe three rival labs disclosed weeks of behind-the-scenes coordination and jointly released a set of cyber-focused AI…security · JEV 0.73OpenAI textGrain watermarking for ChatGPT and Codex in the EUOpenAI will roll out its invisible textGrain watermarking system to ChatGPT/Codex users in the EU over the coming…security · JEV 0.65Cloudflare security-audit-skill: multi-phase security audits for coding agentsA coding-agent skill that runs multi-phase security audits with independently verified, machine-checkable findings.security · JEV 0.63LiveNerf: a pre-registered benchmark for post-release model driftOpen-source, pre-registered 30-day benchmark that detects whether Claude Opus 5.5 quietly gets worse after launch.security · JEV 0.61Sandlock 0.8.8: deferred commit for agent sandboxesThe process-based Linux AI sandbox (no container, no VM) ships deferred commit: every run returns a changeset of what…security · JEV 0.58OpenAI disrupts Moonshot-linked 'adversarial distillation' campaignOpenAI disclosed it shut down a coordinated campaign by 15,000+ users to extract hidden model reasoning, attributing a…security · JEV 0.58
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.