MergeGate: a reliability gate that holds AI-written pull requests to a bar before humans ever open them
The problem
Coding agents now open PRs faster than teams can review them. LinearB's 2026 benchmark (via CodeRabbit) found AI-assisted PRs are 2.6x larger, take 4.6x longer to receive a first review, and only 32.7% are accepted within 30 days versus 84.4% for manual PRs. An instrumented study of 37,623 provenance-labeled PRs found Claude Code PRs wait a median 12.6 hours for first human review, and Devin PRs get reverted 14.5% of the time against 11.5% for humans.
The idea
Why now
Two unlocks landed within days of each other. Microsoft released 301,026 production Copilot coding-agent sessions (9.3M model calls) on Oct 1, the first public corpus showing failure-driven compute amplification patterns in agent coding runs. And the decision-model category converged in the last two weeks: TypeSafe Jev, then AWS Strands Decider 2B (Oct 1, fully open), Cloudflare Clef, and Inception Mercury Decide all concluded that calibrated yes/no judgments at ~100ms make per-PR gating economically viable for the first time.
What it combines
Microsoft's 301k coding-agent traces (the raw material for learning what failing agent runs look like) + open-weights decision models like Strands Decider 2B (cheap, calibrated, sub-150ms judgments) + Cursor/Coding-agent plugin ecosystems (the distribution surface into developer workflows). Traces alone are research; decision models alone are a library; together they become an enforcement point: a gate that knows what failure smells like and can afford to sniff every PR.
MVP
Weekend scope: GitHub App webhook; deterministic rule pack (secrets, lockfile diffs, migration checks); Strands Decider 2B with a 12-question battery and configurable confidence thresholds; GitHub Checks API verdict (block/allow) with an evidence card per blocked PR. Deliberately skip: training on the traces (ship with hand-picked questions first), multi-platform support, analytics dashboards.
Distribution
B2B2C: GitHub/GitLab marketplace app plus native CI integration (GitHub Actions, GitLab CI). Sell per-seat or per-PR-credit to engineering orgs; the real wedge is platform partnerships with Git hosts and code-hosting vendors that want 'agent-safe' merge queues as a differentiator.
Why it wins
CodeRabbit Triage (launched Sept 15) scores and prioritizes PRs with deterministic classification plus AI ranking, but it does not hold the merge button. Greptile writes reviewer-style comments ($1 per deep TREX review) but advises rather than gates. A head-to-head benchmark on deliberately flawed payments code showed all four mainstream reviewers missing the production-breaking timezone bug. MergeGate is a bar, not a commenter: merge is blocked until the bar is cleared.
Risks
Every repo has its own conventions, so a generic question battery will misfire on unusual codebases. De-risk: ship with a small rules DSL plus pre-built packs for 3 popular stacks (TypeScript/Node, Python, Go), and make every blocked decision show its evidence so teams can tune thresholds from real blocked PRs in week one.
Build it with
- Microsoft releases 301,000 Copilot coding-agent tracesFailure signatures from real production agent runs to train the question battery
- Strands Decider 2BOpen, self-hostable decision model answering calibrated yes/no questions at ~115ms per gate check
- Inception Mercury Decide: structured decision model via the System One Decisions APIFallback decision layer via hosted API for teams that do not self-host
- Cursor Plugins (official marketplace)Distribution surface: reach developers inside the IDEs and harnesses where agent PRs originate
Repo to start from
mergegate: GitHub App + open-weights decision-model service that gates agent-written PRs behind a calibrated reliability bar, with deterministic rule packs and per-block evidence cards.
Evidence
- Your pull request review time is mostly waiting and rework (dev.to, cites LinearB 2026 benchmark and the 37,623-PR provenance study)
- AWS Strands Labs releases Strands Decider 2B (MarkTechPost)
- Microsoft opens Copilot coding-agent traces (RuntimeWire)
- Agentic Change Management by CodeRabbit (Futurum)
- 4-agent AI code review comparison on payments code (GitHub, rraj7/circular-payments-demo)
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.