Shipgate: a merge queue that never ships a bad agent PR
The problem
AI agents now write most new code and it ships vulnerabilities: Snyk's research across 4,800 customers found a 108% increase in newly introduced issues as AI-generated code scaled; a Plugin4Shell zero-click RCE hit Codex, Claude Code, Gemini CLI, and Copilot; Stanford research showed developers using AI assistants create more vulnerabilities. Scanners built for human PRs catch patterns but do not understand agent behavior, and the human review every agent PR gets does not scale.
The idea
Why now
This week's unprecedented joint cyber-safety release from OpenAI, Anthropic, and Google DeepMind created the first shared baseline of cyber safeguards to check against; Visa open-sourced VVAH (agentic vulnerability discovery) on Sept 30; and Snyk Evo's 81.5% month-over-month growth proves enterprises are already paying for agent-native security.
What it combines
Joint-lab cyber safeguards (security/compliance) plus Visa VVAH agentic discovery (security) plus Jev-Omni typed judgments (models/decision): static scanning finds patterns, an attacker agent finds exploitable paths, and typed judgments turn findings into enforceable merge decisions. Together they form a full gate, not another scanner.
MVP
GitHub App for one language (Python): run static rules plus an attacker-agent pass on the diff and post a judgment verdict as a merge check. Deliberately skip auto-fix in v1, reputation scoring, and multi-language support.
Distribution
B2B2C: ship as a GitHub/GitLab marketplace app where developers already merge, then upsell to the platforms (GitHub Advanced Security, GitLab Ultimate) as their agent-PR gate.
Why it wins
Semgrep, Snyk, and Checkmarx scan artifacts after the fact; Snyk Evo adds agent supply-chain scanning but not per-PR adversarial discovery combined with a typed approval gate. Nobody puts agentic red-teaming and a judgment verdict inside the merge queue itself.
Risks
False positives block velocity and teams disable the gate. De-risk with a warn-only mode first, measuring precision on real PRs before ever enforcing.
Build it with
- OpenAI, Anthropic and Google DeepMind jointly unveil cyber-focused safety models and safeguardsThe shared cyber-safeguard baseline the gate checks PRs against.
- Visa open-sources VVAH: agentic vulnerability discovery and remediation harnessAgentic vulnerability discovery method to probe each PR diff like an attacker.
- Jev-Omni — 12B multimodal decision classifierTyped judgment layer that renders the approve/queue/block verdict with reasons.
Repo to start from
shipgate - a GitHub App that gates agent-authored PRs: agentic vulnerability discovery plus cyber-safeguard checks plus typed judgment verdicts as merge checks.
Evidence
- TechGig on Snyk Evo agent-native security growth
- InfoWorld on the Plugin4Shell coding-agent RCE
- BitRss on the joint labs cyber safety coordination
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.