Radar / Ideas / CallSentry: every customer call…

CallSentry: every customer call compliance-scored in real time

daily ideamoderateJEV confidence 0.52026-10-02
Outcomeevery customer call compliance-scored in real time, not 2% sampled after the fact

The problem

Contact-center QA teams manually review 1-3% of calls, so compliance violations hide in the unheard 98%. Manual review costs $5-10 per evaluation and reviewers disagree 20-30% of the time, while enterprise conversation-intelligence suites start at $3,000-10,000/month with 3-6 month implementations - out of reach for mid-market BPOs and regulated-industry contact centers.

The idea

A self-hosted QA layer that taps call audio and scores 100% of conversations against a typed compliance scorecard while the call is still live. Streaming transcription produces partial hypotheses within ~100ms; a small open decision model answers yes/no questions per utterance (was the disclosure read? was consent captured? does this need escalation?) with calibrated confidence. Low-confidence or failed checks route to a human reviewer with transcript evidence attached. Ships as a container that plugs into existing CCaaS audio streams, with a dashboard showing per-agent compliance trends and coaching clips.

Why now

Microsoft's MAI-Transcribe-2-Streaming delivers first transcription hypotheses in just over 100ms across 60 languages, making real-time scoring feasible. Open decision models like kev (0.8B-27B, one-forward-pass typed answers) and Strands Decider 2B (Apache-2.0, ~100ms on local hardware) collapse the per-decision cost that made 100% coverage uneconomical with LLM APIs. EU AI Act obligations applicable from August 2026 require documented, challengeable AI QA - calibrated typed decisions are auditable in a way black-box scores are not.

What it combines

Streaming voice transcription (microsoft-mai-transcribe-2-streaming-mai-voice-2-1) + typed decision models (kev-open-trainable-jev-like-family-of-small-decision-models-on-qwen3-5-3-8, strands-decider-2b). Transcription alone produces text that still needs expensive human or LLM review; decision models alone have no real-time text feed to judge. The combination creates a compliance safety net that operates inside the call at ~100ms per decision on cheap hardware - something neither capability delivers alone, and something the post-call incumbents structurally cannot match.

MVP

Weekend scope: SIP/RTP audio tap into streaming transcription, a 5-question typed scorecard (disclosure, consent, escalation, tone, resolution) answered by a small decision model running locally, and a simple dashboard listing flagged calls with transcript snippets. Deliberately skip: real-time agent coaching prompts, languages beyond English and Hindi, CRM integrations, and multi-tenant billing.

Distribution

B2B2C: embed as a compliance add-on inside CCaaS platforms that already carry the call audio (Exotel, Ozonetel, Ameyo in India; Aircall, JustCall, CloudTalk globally). The platform resells per-minute scoring to its tenants and the startup takes a revenue share. Who pays: mid-market BPOs and collections, insurance, and lending contact centers where a missed disclosure is a regulatory fine.

Why it wins

Observe.AI, CallMiner, and Cresta sell per-seat platforms ($80-250/agent/month) aimed at 200+ seat enterprises, with months-long implementations and post-call analytics heritage. CallSentry is real-time, self-hosted, priced per scored minute with no seat minimum, and every score ships with a calibrated confidence plus transcript evidence - built for the mid-market centers the incumbents price out.

Risks

Biggest risk is transcription accuracy on accented speech, background noise, and overtalk - bad transcripts produce bad scores. The MVP de-risks this by running a calibration pass over 200 real customer calls first and publishing word-error-rate and score-agreement numbers before any compliance claim is made.

Build it with

Repo to start from

callsentry - self-hosted real-time call QA: audio tap plus streaming transcription plus typed decision scorecards with calibrated confidence, in one Docker container.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.