Radar / Ideas / CallProof: the vishing firewall for…

CallProof: the vishing firewall for call centers

daily ideaambitiousJEV confidence 0.472026-10-05
OutcomeEvery call that touches money is verified human before the conversation continues.

The problem

AI voice cloning is now free and fully local, and attackers are using it at scale: a finance worker wired $25.6M after a deepfaked CFO video call, and a Ferrari executive was nearly social-engineered by a cloned CEO voice. Call centers and voice agents that handle money, KYC, and account changes have no in-line defense — existing detection is either post-call forensics or enterprise identity suites with six-figure contracts and long sales cycles.

The idea

A real-time audio-path gate that sits in the call's media stream and flags synthetic or cloned voices in under a second, before money moves or sensitive data is discussed. Per-chunk risk scores feed the call router, which can hold, route to a human, or escalate the call when the synthetic probability crosses a threshold. Sold as a metered API plus a one-line SIP/media-stream hook for Twilio and CCaaS platforms, priced per monitored minute — a BPO or mid-market contact center can switch it on per queue with no enrollment, no voiceprints, and no hardware.

Why now

The same stack that made voice agents cheap made the defense deployable in-line: open cloning models commoditized the attack, full-duplex voice models brought sub-second audio reasoning, and small multimodal decision classifiers can now score real-vs-synthetic audio at decision-model latency and cost. Market formation is visible: voice-detection checks are scaling toward billions annually, the FCC has ruled against AI voices in robocalls, and FinCEN has issued deepfake-fraud alerts to financial institutions.

What it combines

Combines voicestudio-open-source-fully-local-elevenlabs-alternative-with-voice-cloning-dubbing-and-dictation-in-646-languages (the commoditized attack: free local cloning in 646 languages) + qwen-audio-3-1-realtime-plus-full-duplex-voice-model-with-think-act-speak-loop (cheap full-duplex audio reasoning) + jev-omni (12B multimodal decision classifier — real-vs-synthetic classification at decision-model latency) + elevenlabs-eleven-v4-v4-turbo-new-emotive-tts-architecture-with-a-100ms-real-time-variant (the ~100ms streaming bar that in-line gating must beat). The mix matters because the attack and the defense are built from the same generation of models — the product is the shield assembled from the sword's own supply chain.

MVP

Build: a Twilio Media Streams WebSocket hook receiving 8kHz PCM audio, chunked into 1-2 second windows, scored by an open-source audio spoof-detection classifier as a calibrated decision model, posting per-chunk risk scores to a webhook the call router uses to hold/route/escalate. Ship a demo dashboard showing live scores per active call. Deliberately skip: caller enrollment, voiceprint identity matching, multi-modal video, on-prem deployment, and any custom model training. The weekend proves the audio-path integration and the sub-second scoring loop.

Distribution

The payer is any business whose agents touch money by phone: banks and insurers (account-takeover losses), BPOs running collections, onboarding, and KYC-adjacent flows, and the CCaaS platforms they rent (Five9, Genesys, NICE). The wedge is B2B2C embedding: a one-line SIP/media-stream hook and metered API so platforms turn it on per queue, priced per monitored minute. Banks adopt it as a compliance-grade loss-prevention line item after one fraud write-down; BPOs adopt it as client-mandated QA infra. One Genesys app-store listing is worth a hundred direct enterprise sales cycles.

Why it wins

Pindrop is an enterprise identity platform with fortune-500 pricing and long sales cycles; Reality Defender is scan/API-based detection, not an always-on in-call gate; Resemble Detect alerts analysts rather than gating the conversation; ValidSoft and Phonexia sell identity verification for known customers, not a zero-enrollment synthetic-voice gate priced per monitored minute for any call center.

Risks

Detection accuracy against production telephony audio (8kHz, codecs, accents, background noise) will not match benchmark claims, and false positives that interrupt legitimate calls kill adoption faster than missed attacks. The MVP de-risks this by testing on real call audio from day one and gating on a threshold that only flags high-confidence synthetic calls, keeping precision the headline metric.

Build it with

Repo to start from

callproof-gateway — the open Twilio media-stream hook plus the real-vs-synthetic scoring service and the risk-score webhook contract; a reproducible, self-hostable reference for in-line vishing gating.

Evidence

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.