CallProof: the vishing firewall for call centers
The problem
AI voice cloning is now free and fully local, and attackers are using it at scale: a finance worker wired $25.6M after a deepfaked CFO video call, and a Ferrari executive was nearly social-engineered by a cloned CEO voice. Call centers and voice agents that handle money, KYC, and account changes have no in-line defense — existing detection is either post-call forensics or enterprise identity suites with six-figure contracts and long sales cycles.
The idea
Why now
The same stack that made voice agents cheap made the defense deployable in-line: open cloning models commoditized the attack, full-duplex voice models brought sub-second audio reasoning, and small multimodal decision classifiers can now score real-vs-synthetic audio at decision-model latency and cost. Market formation is visible: voice-detection checks are scaling toward billions annually, the FCC has ruled against AI voices in robocalls, and FinCEN has issued deepfake-fraud alerts to financial institutions.
What it combines
Combines voicestudio-open-source-fully-local-elevenlabs-alternative-with-voice-cloning-dubbing-and-dictation-in-646-languages (the commoditized attack: free local cloning in 646 languages) + qwen-audio-3-1-realtime-plus-full-duplex-voice-model-with-think-act-speak-loop (cheap full-duplex audio reasoning) + jev-omni (12B multimodal decision classifier — real-vs-synthetic classification at decision-model latency) + elevenlabs-eleven-v4-v4-turbo-new-emotive-tts-architecture-with-a-100ms-real-time-variant (the ~100ms streaming bar that in-line gating must beat). The mix matters because the attack and the defense are built from the same generation of models — the product is the shield assembled from the sword's own supply chain.
MVP
Build: a Twilio Media Streams WebSocket hook receiving 8kHz PCM audio, chunked into 1-2 second windows, scored by an open-source audio spoof-detection classifier as a calibrated decision model, posting per-chunk risk scores to a webhook the call router uses to hold/route/escalate. Ship a demo dashboard showing live scores per active call. Deliberately skip: caller enrollment, voiceprint identity matching, multi-modal video, on-prem deployment, and any custom model training. The weekend proves the audio-path integration and the sub-second scoring loop.
Distribution
The payer is any business whose agents touch money by phone: banks and insurers (account-takeover losses), BPOs running collections, onboarding, and KYC-adjacent flows, and the CCaaS platforms they rent (Five9, Genesys, NICE). The wedge is B2B2C embedding: a one-line SIP/media-stream hook and metered API so platforms turn it on per queue, priced per monitored minute. Banks adopt it as a compliance-grade loss-prevention line item after one fraud write-down; BPOs adopt it as client-mandated QA infra. One Genesys app-store listing is worth a hundred direct enterprise sales cycles.
Why it wins
Pindrop is an enterprise identity platform with fortune-500 pricing and long sales cycles; Reality Defender is scan/API-based detection, not an always-on in-call gate; Resemble Detect alerts analysts rather than gating the conversation; ValidSoft and Phonexia sell identity verification for known customers, not a zero-enrollment synthetic-voice gate priced per monitored minute for any call center.
Risks
Detection accuracy against production telephony audio (8kHz, codecs, accents, background noise) will not match benchmark claims, and false positives that interrupt legitimate calls kill adoption faster than missed attacks. The MVP de-risks this by testing on real call audio from day one and gating on a threshold that only flags high-confidence synthetic calls, keeping precision the headline metric.
Build it with
- Jev-Omni — 12B multimodal decision classifier12B multimodal decision classifier as the low-latency real-vs-synthetic audio judgment layer.
- Qwen-Audio-3.1-Realtime-Plus: full-duplex voice model with think-act-speak loopFull-duplex audio reasoning with think-act-speak loop for the in-call analysis path.
- VoiceStudio: open-source, fully-local ElevenLabs alternative with voice cloning, dubbing and dictation in 646 languagesReference threat model: free, local voice cloning that the detector must catch.
- ElevenLabs Eleven v4 + v4 Turbo: new emotive TTS architecture with a 100ms real-time variantThe ~100ms streaming latency bar the in-line gate must operate under.
Repo to start from
callproof-gateway — the open Twilio media-stream hook plus the real-vs-synthetic scoring service and the risk-score webhook contract; a reproducible, self-hostable reference for in-line vishing gating.
Evidence
- FBI IC3 PSA: Senior US Officials Impersonated in Malicious Messaging Campaign
- CNN: Finance worker pays out $25 million after video call with deepfake 'chief financial officer'
- Biometric Update: Deepfake defense draws new capital as voice detection checks near 5.5B by 2028
- Pindrop: Deepfake Defense for Calls, Virtual Meetings, & Contact Centers
- ValidSoft: Voice Verity real-time synthetic voice detection
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.