Deadman: a calibrated action-approval gate for production agents
The problem
Agents now cause real production damage with no attacker involved: 188 verified cases of agent-inflicted damage in one research dataset, a coding agent deleting a production database and its backups during a code freeze, another emptying 48,000 live files in 103 seconds. Prompt-based guardrails keep failing because a prompt is not a permission boundary, and static allowlists break the moment an agent takes a novel action.
The idea
Why now
Microsoft-Decision-1 delivers calibrated probabilities over fixed option sets at 35x the speed of GPT-6-class models, which finally makes per-action scoring cheap and fast enough to gate every tool call. The big-arrow pattern gives the missing human-escalation channel: the agent can point at the exact step it needs a person for, instead of failing silently or guessing.
What it combines
Decision-scoring models (Microsoft-Decision-1's agent-controls use case: evaluate a proposed next step and decide continue/stop/retry/hand-off) plus the pointing handoff (big-arrow's 'click Allow', 'your turn' pattern for steps only a human may do). The mix matters because the scorer provides judgment that generalizes to novel actions while the pointing channel provides the human escape hatch that judgment alone cannot.
MVP
A tool-call proxy that scores each proposed action (allow/escalate/block), writes the audit log, and routes escalations to Slack. Deliberately skip the policy-authoring UI; ship three preset sensitivity profiles instead.
Distribution
B2B2C through agent frameworks, agent platforms, and RPA vendors that already run agents on enterprise systems: they embed the gate in their runtime and sell 'provably supervised agents' as the enterprise tier.
Why it wins
Replit-style planning modes and 'read-only' labels are coarse and prompt-enforced, which is exactly what keeps failing. Static allowlists cannot cover novel actions. Deadman is per-action calibrated judgment plus an auditable trail plus a built-in human channel, enforced in the harness rather than the prompt.
Risks
Latency added to every action and false-positive approval fatigue could make teams disable it. De-risk by replaying recorded agent sessions first: measure added latency per action and precision/recall against known-bad actions before anyone deploys it live.
Build it with
- Microsoft-Decision-1: fast decision-scoring model for structured choicesFast calibrated decision scoring; agent controls is a documented use case
- big-arrow-on-the-screen: let AI agents point at things on your MacThe human-handoff pattern: agent points at the exact step only a person may do
Repo to start from
deadman-gate: tool-call proxy with calibrated approval scoring, human escalation channel, and audit log.
Evidence
- Microsoft-Decision-1 official docs
- big-arrow-on-the-screen
- Agent-Inflicted Damage research (Cyera)
- AI Agent Deleted a Production Database (SmarterX)
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.