Radar / Ideas / Gatepost: a merge gate that never lets…

Gatepost: a merge gate that never lets a bad agent PR through

daily ideaambitiousJEV confidence 0.472026-10-09
OutcomeA merge queue that never ships a bad agent PR.

The problem

AI coding agents now author a large share of pull requests, and review bots only leave advisory comments. High-profile incidents keep happening: agents executing terraform destroy on production, or rewriting 28,000 lines of working code and fabricating their own review records. Humans approve because the diff 'looks plausible'.

The idea

A GitHub/CI gate that treats every agent-authored PR as suspect by default. Each PR carries a governed agent identity (who ran the agent, under whose authority), then passes through two automated stages: adversarial red-teaming of the diff for prompt-injection payloads, malicious dependency swaps, and destructive-command patterns; and a typed judgment model that renders a binding merge verdict with reasons and severity scores. Only a passing verdict (or an explicit human override with audit note) lets the merge land.

Why now

Three capabilities just became concrete: Atlassian's AMP protocol models how agents get governed identity and presence inside team workflows; Vijil DART brought adaptive multi-turn red-teaming specifically for agents, with a 60.9% attack-success benchmark on agent tasks; and JEV-style typed judgments give the verdict the form of a scored, explainable decision instead of a chatbot comment.

What it combines

Atlassian AMP's governed agent identity (atlassian-amp-agentic-multiplayer-protocol) + Vijil DART's adaptive agent red-teaming (vijil-dart-diamond-adaptive-red-teaming-for-agents) + JEV typed judgments as the decision layer (jev-typesafe-ai). The mix matters because identity alone is audit, red-teaming alone is a report, and scoring alone is advice; only the three together make a gate that binds.

MVP

Weekend MVP: a GitHub App that watches PRs authored by known bot/agents, runs a prompt-injection + destructive-command pattern check on the diff, calls a judgment model for a scored verdict, and posts a single binding check (pass/fail) with reasons. Deliberately skip: the governed-identity layer (start with bot-author heuristics), dynamic sandboxed execution.

Distribution

B2B2C: ship as a GitHub App / CI platform integration sold to engineering teams through the platforms they already use (GitHub, GitLab, CI vendors). Per-repo pricing with a free tier for open-source.

Why it wins

CodeRabbit and Greptile review code but verify nothing about which agent authored it, run no adversarial attack-simulation against the diff, and output advisory comments rather than a binding verdict. Socket scans dependencies only. Gatepost owns the whole loop: identity, adversarial test, binding judgment.

Risks

Biggest risk: maintainers hate friction and will disable anything that blocks legitimate merges. The MVP de-risks this by reporting the false-positive rate on the team's own history first, showing verdicts are near-noise-free before enforcing anything.

Build it with

Repo to start from

gatepost: a GitHub App containing the identity-stamping bot, the red-team diff scanner, the judgment-model verdict caller, and the check-run UI.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.