Radar / Ideas / BlastRadius: verified architecture…

BlastRadius: verified architecture maps plus decision-scored incident triage

daily ideamoderateJEV confidence –2026-10-10
OutcomeWhen an alert fires, the on-call engineer sees a verified blast-radius map of the failing service with the likely culprit ranked, instead of 20-40 minutes of dashboard-hopping.

The problem

The investigation layer is the MTTR gap: engineers spend 20-40 minutes correlating signals across services, dashboards, and deploys before any fixing begins. Architecture diagrams rot within weeks of being drawn, and runbooks assume failure modes that no longer exist. Existing AI SRE tools investigate from the model's own knowledge rather than from a trustworthy map of the actual system.

The idea

A GitHub Action regenerates a verifiable, source-grounded architecture map (typed JSON IR, deterministic compile) on every merge to main. A PagerDuty/Opsgenie webhook handler overlays the alerting service on the fresh map and runs a decision-scoring model over candidate components, ranking likely culprits with calibrated probabilities. The responder gets one screen: the map, the blast radius, the ranked suspects, and links to the exact files.

Why now

Archify showed agents can generate verifiable architecture diagrams from a codebase on demand, so maps can stay fresh automatically. Microsoft-Decision-1 delivers calibrated per-option scoring at 35x the speed of GPT-6-class models, making per-alert ranked classification cheap enough to run on every page. The two together make topology-grounded triage a deploy hook, not a research project.

What it combines

Verifiable architecture diagrams (agent-generated, source-grounded, deterministic) plus fast decision-scoring models (calibrated probabilities over fixed option sets). The mix matters because the map removes the hallucination surface that plagues AI SRE tools, and the decision model turns the map into a ranked answer fast enough for a 3am page.

MVP

GitHub Action that regenerates the map on merge to main, plus a PagerDuty webhook that overlays the alerting service and returns ranked components with probabilities. Deliberately skip auto-remediation; this release only points.

Distribution

B2B through incident-management marketplaces: a PagerDuty/Opsgenie integration sold to SRE teams at mid-size SaaS companies, priced per incident triaged or per seat.

Why it wins

incident.io AI SRE and Sherlocks.ai investigate with AI but from the model's training knowledge, not a verified map of your infra. Rootly and FireHydrant automate process, not topology-aware triage. BlastRadius grounds every suggestion in a map regenerated from the code that is actually deployed.

Risks

Map regeneration cost and latency on large monorepos could make the 'fresh on every deploy' promise false. De-risk by measuring generation time on three real repos first and caching maps per deploy, regenerating only what changed.

Build it with

Repo to start from

blastradius: deploy hook that regenerates verified architecture maps plus an alert webhook that ranks suspect components.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.