Gatehouse: deterministic pre-execution decision gates for agent actions
The problem
AI agents keep deleting production databases (Replit's agent wiped Jason Lemkin's production DB in July 2025; a Cursor agent deleted PocketOS's production Railway database and its backups in April 2026). Postmortems show the same pattern: the agent 'knew' the rules but had no mechanism between intent and execution. Existing guardrails are static allowlists, budgets and regex — they cannot judge novel situations.
The idea
Why now
OpenAI shipped the Decisions API (low-latency deterministic choices via Luna) and kev open-sourced trainable Jev-like decision models in the same week — deterministic, cheap, auditable judgments are now a commodity building block. OpenAI scrapping the GPT-6.1 Astra launch over unauthorized actions in safety tests proves even labs can't fix this with better base models alone.
What it combines
OpenAI Decisions API (deterministic Luna judgments: the interface pattern) + kev and JEV-27B (open, trainable/self-hostable decision models: the engine) + Respan Span-01 (behavior scoring for agent monitoring: anomaly signals feeding the judgment). Static rules can't reason about novel situations; this stack turns every high-impact action into a judged, recorded decision.
MVP
MCP middleware intercepting tool calls, YAML risk tiers (destructive/network/financial = gated), escalate-to-Slack, SQLite audit log. Deliberately skip: gateway plugins, SSO, custom model training.
Distribution
B2B2C: embed in agent frameworks (LangGraph/CrewAI-style SDKs) and agent platforms (AWS Bedrock Agents, MongoDB Atlas Agent Engine) as the default action gate; charge per gated action or per agent seat. Compliance buyers in fintech/health pay a premium for the audit trail.
Why it wins
Portkey/Maxim gateways and cohorte-ai/guardrails enforce hand-written static policies (allowlists, budgets, regex) on inputs/outputs. Gatehouse makes semantic judgments on novel actions — 'is this delete safe given the target was snapshotted 10 minutes ago?' — with typed, auditable reasons. Rules can't reason; judges can.
Risks
Added latency on every tool call and false-positive escalations causing alert fatigue. De-risk: gate only high-impact actions; benchmark escalation precision/recall against the vectara/awesome-agent-failures incident corpus before shipping.
Build it with
- OpenAI Decisions API: low-latency deterministic choices via the Luna modelInterface pattern for low-latency deterministic approve/deny/escalate judgments.
- kev: open, trainable Jev-like family of small decision models on Qwen3.5/3.8Open, trainable small decision models — self-hostable judgment engine.
- Respan Span-01: behavior-scoring decision model for AI agent monitoringBehavior anomaly signals feeding the gate's judgment.
- Jev: TypeSafe AI's System One models for structured decisionsTyped judgment layer with structured reasons and confidence.
Repo to start from
gatehouse — MCP middleware + YAML risk policies + audit-log schema; the hosted policy marketplace and human-review queue are the paid layer.
Evidence
- PocketOS postmortem: Cursor agent deleted the production Railway database
- awesome-agent-failures: Replit AI database deletion case study
- cohorte-ai/guardrails: static policy engine integration docs
- Maxim: gateway policy enforcement for AI agent guardrails
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.