Skip to content

Modulate’s $25M round tests voice AI’s scalability beyond LLMs

Modulate has raised $25 million to scale Velma, its "audio-native voice intelligence model," which the company says achieves 95% precision across 76 genres. The round, led by Future Ventures, reflects growing investor curiosity about startups exploring alternatives to large language models (LLMs) for voice applications. Whether Velma can outperform LLMs at understanding real conversations remains to be seen as the company moves beyond demos and into production.

The funding arrives as voice AI startups seek ways to stand out in a competitive market. Many competitors rely on LLMs fine-tuned for speech, an approach some argue may not be optimized for audio-specific tasks. Velma’s developers suggest it is designed to handle functions like transcription, emotion analysis, deepfake detection, and compliance enforcement more efficiently. If the model delivers on these promises, it could be useful in latency-sensitive or cost-constrained scenarios, such as call centers or live moderation. While Modulate hasn’t shared details about customer adoption, its emphasis on precision across genres hints at potential interest from industries like music and entertainment, where generic speech models often struggle.

Still, the 95% precision claim invites skepticism. Modulate hasn’t released the dataset or methodology behind the figure, and genre-based benchmarks may not account for real-world variability. When covering Treble’s funding last month, we noted the difficulty of testing voice AI in uncontrolled environments—an area where Velma will need to prove its reliability. Earlier coverage of Modulate’s $25 million raise framed it as a compliance-focused tool, but the company’s latest messaging suggests a broader ambition to demonstrate technical advantages over broader AI models. Whether this approach will hold up in practice remains an open question.

The round also highlights a broader trend in AI funding, with some investors backing specialized models over large-scale solutions. Instinct’s $350 million raise last month showed demand for expansive AI platforms, while Modulate’s smaller funding round reflects a different strategy. However, the market tends to favor startups that can deliver tangible results. Apate.AI, which raised $8.15 million to combat scam calls, faces similar pressure to demonstrate its bots’ effectiveness in real-world scenarios. Modulate’s challenge will be showing that Velma isn’t just another voice model but one that meets its promises in live deployments.

What to watch next: Whether Modulate announces customers, particularly in industries like music or gaming where genre-specific precision is valuable. If the company secures deals with platforms handling diverse audio content, it could validate its approach. If not, the $25 million may appear as a bet on an unproven concept. Either way, the raise underscores that voice AI remains an evolving space—and investors are still willing to back startups aiming to solve its challenges.

Sources: digitalmusicnews.com

“Modulate’s funding signals investor interest in small, specialized audio models, but its claims of high precision across genres will need real-world validation to justify the bet.”
— StartupReader
ShareLinkedInXWhatsApp

Read the original reporting

The outlets below did the original reporting.

Related briefs

This brief was drafted automatically from the sources above and published under our editorial policy. Spotted an error? Tell us.