Radar / Ideas / Evergreen: e2e tests that maintain…

Evergreen: e2e tests that maintain themselves

daily ideamoderateJEV confidence 0.452026-10-05
OutcomeA green e2e suite your team never hand-maintains — ship daily.

The problem

End-to-end suites are the highest-value and least-trusted tests in most codebases: they break en masse when the DOM shifts, teams stop believing the results, and the accepted workflow is hand-triaging flakes every morning. Engineers actively debate whether E2E tooling is flaky enough to flip the testing pyramid's payoff. Every 'AI test' product on the market heals at the selector level while the brittle scripts remain — nobody sells a suite that maintains itself.

The idea

Goal-driven e2e testing as a system that writes, judges, and repairs itself. Tests are defined as natural-language goals, an agent drives the app to reach them, a deterministic judge emits pass/fail from concrete state (final URL, DOM presence, API/DB state — no LLM vibes-grading), and a trace-analysis loop clusters recurring failures and rewrites flows when the UI changes. Ships as a GitHub App plus a one-line GitHub Actions integration with per-repo pricing, so the first run happens inside the team's existing CI. Open-source the runner to seed the Playwright community; a green zero-maintenance badge sells it upward.

Why now

Three capabilities converged: natural-language app driving with decision-model executors is now open source, trace-analysis agents can group similar failures across runs instead of dumping logs, and small calibrated decision classifiers give deterministic pass/fail verdicts at cents per thousand — replacing both brittle scripts and LLM-graded assertions. The 'self-healing' incumbents all stopped at smarter selectors, leaving the goal-driven plus deterministic-judge plus repair-loop combination unclaimed, and the open-source, CI-native, dev-priced slot is empty.

What it combines

Combines e2e-by-testerarmy (natural-language goal-driven app driving with decision-model executors) + litellm-lens (AI agents that analyze traces, group similar failures, and link back to runs) + upstage-solar-decide-decision-model-from-an-established-lab-built-on-solar-mini-4 (calibrated decision model as the cheap deterministic pass/fail judge). Driving alone gives you flaky agents, judging alone gives you assertions, clustering alone gives you dashboards — together they form the self-maintaining loop: drive, judge deterministically, cluster failures, repair flows.

MVP

Build: a Playwright-MCP agent driving a small demo app from 5 natural-language goals, a deterministic judge emitting pass/fail from concrete state (URL, DOM presence, API/DB state — rule-based at first, no LLM grading), a JSON trace log with screenshots and DOM diffs per run, and a nightly GitHub Action posting results as a check. Deliberately skip: actual auto-repair (surface repair suggestions first), any UI/dashboard, multi-browser matrices, auth-heavy flows, and the trace-clustering layer (hand-cluster the first 20 failures).

Distribution

Who pays: engineering teams of 5-50 shipping daily on Playwright or Cypress, where one flaky suite blocks merges and eats a dev's morning. How it reaches them: a GitHub App in the GitHub Marketplace plus a one-line GitHub Actions integration, so the first run happens inside their existing CI with zero new infrastructure; open-source the runner and seed the Playwright community with migration guides; land at per-repo pricing ($49-199/month, no per-seat enterprise quotes) and let the green badge sell upward to the CTO.

Why it wins

Momentic grades verdicts with LLMs instead of a deterministic layer; QA Wolf is a managed human service that moves maintenance cost to the vendor instead of eliminating it; testRigor, Mabl, Autify, and Virtuoso are script- or record-centric with 'self-healing' limited to smart locators — none combine goal-driven execution, a deterministic judge, and a trace-clustering repair loop as a self-maintaining system.

Risks

Goal-driven execution with no scripts makes pass/fail inherently fuzzy — if the judge mis-grades even occasionally, teams stop trusting the suite and the self-maintaining promise collapses. The MVP de-risks this by grounding verdicts in deterministic state assertions with the full trace attached to every verdict for human audit, deferring any LLM grading entirely.

Build it with

Repo to start from

selfheal-e2e — the goal-driven Playwright runner, the deterministic pass/fail judge, the JSON trace format, the demo app with 5 goal-defined flows, and the nightly GitHub Action workflow.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.