Modulate Raises $25M to Combat Deepfake Voices with AI
Key Takeaways Modulate has secured $25 million in new funding to advance its AI-powered audio analysis platform. The investment targets the escalating threat of deepfake voice fraud, impersonation,...
Key Takeaways
- Modulate has secured $25 million in new funding to advance its AI-powered audio analysis platform.
- The investment targets the escalating threat of deepfake voice fraud, impersonation, and the need for enhanced oversight of AI-driven conversations.
- The company’s Velma platform, powered by its Ensemble Listening Model (ELM), analyzes spoken audio for various indicators, including synthetic generation, with reported high accuracy in deepfake detection benchmarks.
- The funding will expand research, developer tools, partnerships, and deployment options, aiming to integrate voice intelligence into security and communication workflows across multiple industries.
Modulate Secures $25M to Enhance Deepfake Voice Detection and Audio AI Capabilities
In a significant move to counter the growing threat of deepfake voice technology, Modulate has successfully raised $25 million in a funding round. This capital infusion is earmarked to scale the company’s advanced audio-native AI platform, which addresses critical issues such as deepfake voice fraud, impersonation schemes, and the need for safer, more accountable automated conversations within enterprises.
Table Of Content
The investment round was spearheaded by Future Ventures, with additional participation from Hyperplane and Lakestar. This latest funding brings Modulate’s total capital raised to $60 million. The company plans to leverage these funds to accelerate its AI research, bolster engineering efforts, develop new developer tools, forge strategic partnerships, and expand deployment options for organizations seeking to integrate sophisticated voice intelligence into their security and communications frameworks.
Combating Sophisticated Voice Threats
The urgency for such technologies has intensified as increasingly convincing synthetic speech makes traditional telephone-based verification methods unreliable. Modulate’s innovative approach moves beyond mere transcript analysis. Its technology meticulously examines the nuances of how something is spoken, scrutinizing elements such as emotion, tone, emphasis, speaker intent, conversational patterns, and tell-tale signs of synthetic generation.
This detailed analysis is crucial in combating sophisticated vishing attacks and executive-impersonation scams. In these scenarios, the spoken words might appear innocuous, but underlying factors like an artificial voice, a manipulated sense of urgency, or deceptive intonation can expose the fraudulent nature of the interaction.
According to an announcement published by Modulate, the core of their offering is Velma, a real-time conversation-understanding platform. Velma is engineered to identify a spectrum of events, including fraud attempts, harassment, indicators of customer dissatisfaction, policy violations, and operational failures by voice AI agents.
Velma is powered by Modulate’s proprietary Ensemble Listening Model (ELM), which orchestrates over 100 specialized audio models. This architecture avoids routing all tasks through a single foundation model, a design choice Modulate asserts can achieve up to 1,000 times greater inference efficiency. This significantly reduces the computational, memory, and energy demands typically associated with comprehensive audio analysis.
Performance and Applications
A key security capability of the platform is deepfake detection. Modulate reported an impressive average equal error rate of 1.104% across 14 Speech DF Arena datasets as of August 19, 2026. This translates to an approximate 98.9% detection accuracy, securing Modulate the top position on that specific benchmark snapshot. The 316-million-parameter model can initiate streaming verdicts within 2.5 seconds of speech commencement.
However, Modulate advises that production systems rarely operate precisely at the equal-error threshold. They recommend that customers fine-tune detection thresholds using their own labeled recordings, as overly strong detection can lead to an increase in false positives.
Modulate currently processes over 10 million hours of audio monthly, with a cumulative total exceeding 600 million hours. The company also achieved the leading position on Hugging Face’s Open ASR Leaderboard for transcription in July, outperforming 88 other evaluated models. Their batch transcription API starts at $0.03 per hour, while deepfake detection is priced at $0.25 per hour.
The new capital will facilitate the expansion of Modulate’s APIs and software development kits, the creation of industry-specific models, the integration of new partner solutions, and support for diverse deployment environments. The company is targeting a broad range of applications, including fraud prevention, healthcare security, contact center oversight, moderation for gaming and social platforms, child safety initiatives, and supervision of AI voice agents.
Real-time audio analysis offers significant advantages for defenders, enabling them to challenge suspicious callers, escalate high-risk sessions, or intervene in harmful conversations proactively, rather than relying on post-event recording reviews. For security teams, Modulate emphasizes treating this technology as a crucial detection layer rather than a definitive proof of identity. The company also notes that benchmark leadership does not guarantee identical performance on compressed telephone audio, unfamiliar languages, noisy environments, replay attacks, or novel voice generators. The current arena primarily measures English and Mandarin and does not fully account for streaming behavior or end-to-end narrowband telephony.
Carter Huffman, Modulate’s chief executive, highlighted that voice is rapidly becoming a primary AI interface, generating challenges that simple transcripts cannot resolve. With this latest funding, Modulate is asserting its belief that advanced audio understanding will become a fundamental infrastructure component for authenticating callers, monitoring automated agents, and detecting sophisticated manipulation. As synthetic voices continue to evolve, the integration of audio AI with multi-factor verification, stringent transaction controls, and human oversight will remain vital to prevent a single detection score from becoming a critical point of failure.
What You Should Do
- Evaluate your current authentication and fraud detection systems for vulnerabilities to sophisticated voice impersonation and deepfake attacks.
- Consider integrating advanced audio analysis solutions as a layer in your security architecture, focusing on real-time detection capabilities.
- Implement multi-factor authentication for critical transactions and sensitive data access, ensuring that voice is only one component of verification.
- Educate employees and customers about the risks of deepfake voice fraud and social engineering tactics, emphasizing the importance of verifying unusual requests through alternative, trusted channels.
- Regularly review and update your incident response plans to include protocols for addressing suspected deepfake or voice impersonation incidents.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



No Comment! Be the first one.