RedSquad: the always-on adversary for production AI agents
The problem
Production agents ship untested against adversarial users, and the cost of failure is no longer theoretical: prompt injection turned a dealership chatbot into a viral liability, and courts have held companies liable for their agents' mistakes. Existing red-teaming options are either self-serve tools that teams configure and forget, or consulting-grade assessments sold per engagement. Nobody sells a continuous, always-on red-team wired into a team's CI with an automated verdict loop.
The idea
Why now
The building blocks only just matured into a service shape: open attack frameworks now cover 50+ vulnerability classes with research-backed methods, trace-analysis agents can cluster recurring failures across runs instead of dumping raw logs, and small calibrated decision classifiers can judge pass/fail at cents per thousand verdicts instead of GPT-graded vibes. Meanwhile procurement pressure is doing the marketing: OWASP's dedicated agentic-AI threat catalog standardizes what to test against, and agent platforms are adding red-teaming reports natively, which proves the gap is real.
What it combines
Combines deepteam-open-source-red-teaming-framework-for-llm-apps (the open attack framework, 50+ vuln classes) + visa-open-sources-vvah-agentic-vulnerability-discovery-and-r (the agentic discovery-to-remediation harness pattern) + litellm-lens (trace analysis that groups similar failures instead of dumping logs) + jev-typesafe-ai-s-system-one-models-for-structured-decisions (calibrated yes/no/choice classifiers as the deterministic pass/fail judge). Each piece alone is a tool; together they become a managed adversary: attack generation, failure clustering, deterministic verdicts, and remediation tracking in one loop.
MVP
Build: a curated attack pack runner (30-50 attack classes) that hits any HTTP agent endpoint, a deterministic pass/fail judge per class using a calibrated small classifier, a minimal verdict dashboard (score by class, failing transcripts, severity), and a GitHub Action that runs the pack on PR/push and posts results as a PR comment. Deliberately skip: auto-opening remediation PRs (comment-only at first), multi-turn crescendo attacks beyond 2-3 turns, compliance report exports, and any runtime guardrail product.
Distribution
The buyer is the agent platform team or AI startup shipping agents into enterprise pilots — the group that gets asked 'is this safe?' in procurement and has no answer. Reach them B2B2C: plugins in the LangChain, CrewAI, and Vercel AI SDK ecosystems put the service one click from every agent builder, and a GitHub App that runs the attack pack on every PR turns security testing into CI furniture. Procurement pressure (EU AI Act, enterprise security reviews) does the outbound once reference logos exist.
Why it wins
Giskard sells per-engagement expert assessments, Promptfoo is a self-serve tool you run yourself, Lakera builds runtime guardrails that block attacks rather than testing for them, and ZioSec is the closest live analog but lacks the CI-native auto-remediation loop. RedSquad is the continuous, developer-adoptable service: always-on, open attack pack, verdicts in CI.
Risks
Judge unreliability is the killer risk: noisy or false-positive verdicts destroy trust in a security product faster than missing attack classes. The MVP de-risks it by freezing a small, human-audited attack pack and calibrating each judge against labeled transcripts before widening coverage.
Build it with
- DeepTeam: open-source red teaming framework for LLM appsOpen attack framework with 50+ vulnerability classes and 20+ attack methods as the base attack pack.
- Visa open-sources VVAH: agentic vulnerability discovery and remediation harnessAgentic discovery-to-remediation pipeline pattern for turning findings into verified fixes.
- LiteLLM Lens — AI agents that analyze agent traces inside the gatewayTrace-analysis pattern for clustering recurring failures across runs instead of dumping raw logs.
- Jev: TypeSafe AI's System One models for structured decisionsCalibrated yes/no/choice classifiers as the cheap, deterministic pass/fail judge layer on every attack verdict.
Repo to start from
redsquad — the open attack-pack runner: attack configs, endpoint harness, calibrated judge definitions, the GitHub Action, and a starter verdict-dashboard template.
Evidence
- DeepTeam: open-source LLM/agent red-teaming framework
- OWASP Agentic AI Threats and Mitigations
- Giskard: managed AI red-teaming assessments
- Promptfoo: automated red teaming for agents
- HN: Show HN — free adversarial security testing for agents
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.