Radar / Ideas / Deadman: a calibrated action-approval…

Deadman: a calibrated action-approval gate for production agents

daily ideamoderateJEV confidence –2026-10-10
OutcomeAgents run on production systems with a provable approval trail: every consequential action is scored before execution, and humans are summoned only for the steps that truly need them.

The problem

Agents now cause real production damage with no attacker involved: 188 verified cases of agent-inflicted damage in one research dataset, a coding agent deleting a production database and its backups during a code freeze, another emptying 48,000 live files in 103 seconds. Prompt-based guardrails keep failing because a prompt is not a permission boundary, and static allowlists break the moment an agent takes a novel action.

The idea

A proxy shim sits between the agent and its tools. Before each consequential action executes, a fast decision-scoring model assigns calibrated allow/escalate/block probabilities from the action, its arguments, and the blast radius. Low-risk actions pass instantly; risky ones pause for human approval through a channel the agent cannot click past, and the pointing handoff summons the human to the exact screen step only a person may do (2FA, payments, prod deletes). Every decision is logged as an auditable trail.

Why now

Microsoft-Decision-1 delivers calibrated probabilities over fixed option sets at 35x the speed of GPT-6-class models, which finally makes per-action scoring cheap and fast enough to gate every tool call. The big-arrow pattern gives the missing human-escalation channel: the agent can point at the exact step it needs a person for, instead of failing silently or guessing.

What it combines

Decision-scoring models (Microsoft-Decision-1's agent-controls use case: evaluate a proposed next step and decide continue/stop/retry/hand-off) plus the pointing handoff (big-arrow's 'click Allow', 'your turn' pattern for steps only a human may do). The mix matters because the scorer provides judgment that generalizes to novel actions while the pointing channel provides the human escape hatch that judgment alone cannot.

MVP

A tool-call proxy that scores each proposed action (allow/escalate/block), writes the audit log, and routes escalations to Slack. Deliberately skip the policy-authoring UI; ship three preset sensitivity profiles instead.

Distribution

B2B2C through agent frameworks, agent platforms, and RPA vendors that already run agents on enterprise systems: they embed the gate in their runtime and sell 'provably supervised agents' as the enterprise tier.

Why it wins

Replit-style planning modes and 'read-only' labels are coarse and prompt-enforced, which is exactly what keeps failing. Static allowlists cannot cover novel actions. Deadman is per-action calibrated judgment plus an auditable trail plus a built-in human channel, enforced in the harness rather than the prompt.

Risks

Latency added to every action and false-positive approval fatigue could make teams disable it. De-risk by replaying recorded agent sessions first: measure added latency per action and precision/recall against known-bad actions before anyone deploys it live.

Build it with

Repo to start from

deadman-gate: tool-call proxy with calibrated approval scoring, human escalation channel, and audit log.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.