Editor's Note: Inside AI talked with Carter Huffman, CEO and Co-Founder of Modulate, for an exclusive, deep-dive Q&A exploring Modulate's hidden capibilities, architecture powering their platform, and how the software actually work. To read his raw, unfiltered insights on their technical roadmap, you can jump to the Interview section here.
September 28, 2026, (Inside AI) — The frontier audio artificial intelligence sector achieved a major financial milestone today as a prominent industry leader secured a massive capital injection. Today, the Boston based voice intelligence developer Modulate has announced a successful $25 million funding round to scale its infrastructure. Pioneering venture capital firm Future Ventures led the investment, expanding total institutional backing to $60 million. Key multi-stage international investors Hyperplane and Lakestar joined the round to accelerate global distribution.
Modulate plans to deploy the fresh capital across AI/ML research, product and engineering, developer relations, and strategic partnerships, while continuing to broaden the APIs, models, software development kits (SDKs), and deployment options available to software developers.
The expansion arrives at a critical moment for the global software sector. While massive investments flow into synthetic systems that mimic human speech, understanding these interactions remains a challenge.

Building the Understanding Layer for Voice AI
Traditional large language systems rely entirely on text transcripts, which strip away vital acoustic patterns. This technical gap exposes enterprises to advanced voice spoofing threats and unchecked conversational compliance errors. Modulate addresses this vulnerability through its flagship Velma platform, an advanced real-time audio analytics engine. The technology scans direct signals including emotion, intent, tone, emphasis, synthetic speech generation, and conversational behavioral habits.
Those signals can be used independently or composed to recognize higher-level events, from fraud attempts and AI agent failures to harassment, customer dissatisfaction, and policy violations.
Modulate’s flagship Velma platform is the leading model for understanding conversations, with 2x greater accuracy than traditional LLMs at detecting true positive results and 7x fewer false positive results.

The core technology behind the platform utilizes an unconventional structural design. Instead of operating a single massive foundation framework, Modulate utilizes a proprietary Ensemble Listening Model architecture. This setup orchestrates more than 100 highly specialized audio processing frameworks simultaneously.
The central supervisor combines specific neural layers depending on the input acoustic characteristics. This system achieves massive efficiency gains over traditional single model approaches. It drops computing, memory, and electricity overhead substantially during enterprise scaling. The technological framework demonstrates up to 1,000 times greater efficiency than a standard multimodal system, making real-time analysis cost-effective for large digital communications platforms.
“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can't be solved from a transcript,” said Carter Huffman, CEO and co-founder of Modulate. “We're already using audio-native AI to protect organizations from deepfake attacks, help voice agents understand emotion and respond with more empathy, identify dangerous behavior in online conversations, and monitor whether voice agents are actually performing the way they're supposed to. Underneath all of that are more than a hundred specialized models working together to understand what's really happening across audio, with dramatically less cost and compute than traditional large models.”
| Acheivements | Details |
| New Capital Raised | $25 Million |
| Lead & Participating Investors | Future Ventures (Lead), Hyperplane, Lakestar |
| Monthly Audio Volume Processed | >10 Million Hours / Month |
| Cumulative Audio Analyzed | >600 Million Hours Total |
| Hugging Face Open ASR Benchmark | Ranked #1 (Transcription API priced at $0.03/hr batch) |
| Hugging Face Deepfake Benchmark | Ranked #1 (98.9% accuracy on public benchmark data) |
| Velma Detection Accuracy vs. LLMs | 2x higher true positive accuracy, 7x fewer false positives |
| Inference Efficiency vs. Monolithic LLMs | Up to 1,000x greater efficiency |
This data scaling has translated directly into public technical achievements. The startup secured #1 spot on the Hugging Face Open ASR Leaderboard for automated speech recognition. It also ranks first on the platform for deepfake speech benchmark verification metrics.
The financial performance numbers reveal an aggressive infrastructure scaling strategy. The enterprise batch processing API operates at $0.03 per hour, challenging legacy voice vendor costs. The deepfake speech validation framework maintains a 98.9% accuracy score on standard public evaluation testing datasets. This accuracy makes the software highly viable for deployment within critical security networks.
In-Depth Executive Insights: Q&A with Modulate CEO Carter Huffman
To better understand how audio-native AI functions under the hood and addresses real-world vulnerabilities, we asked CEO Carter Huffman to address critical technical, operational, and competitive questions.

Question:1
Inside AI: Standard transcription tools focus heavily on what text is being said. When detecting advanced audio deepfakes, what are the most critical vocal signals (like emotion, tone, or breath) that your models catch which traditional text-based software completely misses?
Carter Huffman: "Modulate combines vocal emotion, emphasis, stress, tone, environmental noise, non-speech audio, audio quality, and other factors with transcription in order to get a complete understanding of a voice conversation, just the way a human listens. A transcript tells you what someone said, but not necessarily what is actually happening in the conversation.
The important part is that we do not treat any one signal as definitive. A synthetic voice, for example, is not automatically malicious; it could be assistive technology. But if that synthetic voice is also trying to impersonate someone, create urgency, manipulate the person on the other end or bypass a security procedure, the combination of those signals tells you something much more meaningful about the interaction.
This shows up in our models via stressed emotion signals, suspicious background audio and other non-text factors. That is the fundamental difference between transcription and audio-native understanding."
Question:2
Inside AI: While ranking #1 on Hugging Face is a massive win, public benchmarks use clean datasets. In the messy real world like calls with bad reception or background noise, how resilient is the Velma platform against audio compression and low-quality microphones?
Carter Huffman: "Real-world audio is actually where a lot of our technology was forged. And our models consistently win benchmarks on both clean and noisy audio. Modulate got its start in gaming, which is an incredibly messy environment: cheap microphones, background noise, different accents and languages, people talking over one another, and highly variable audio quality. That forced us very early to build models that could perform outside of laboratory conditions.
Our ensemble architecture, Velma, also helps here because we can route audio toward models that are better suited to a particular language, accent or audio condition rather than expecting one giant model to perform equally well on all audio. We have multiple specialized models available even within individual capabilities such as transcription, and the system can compare outputs, discard weaker results and focus on the models that are performing best for that particular conversation."
Question:3
Inside AI: Big Tech is moving toward massive, all-in-one multimodal foundation models. What prevents a giant like OpenAI or Google from training their next model to replicate Velma's context layer, and how does your 100+ model ELM architecture protect you?
Carter Huffman: "Humans pay attention to dozens of different audio factors when listening to a conversation, which powers our ability to rapidly understand voice - Modulate’s ELM architecture follows the same pattern to extract all of the signal in every fragment of audio, and our models are trained to produce high accuracy results efficiently and deterministically, which generative foundation models struggle with (e.g. hallucinations).
Our advantage is that we have made a fundamentally different architectural tradeoff. Instead of asking one enormous model to do everything, we build highly specialized models that are exceptionally good at narrow audio tasks, then orchestrate them together to create a much richer understanding of the conversation. That lets us get the flexibility of a larger system without paying the compute cost of running a giant foundation model for every task.
It also lets us move very quickly. If a customer needs a new capability, we do not necessarily need to retrain the entire system. We can build a targeted model, add it to the ensemble and update the orchestration layer so the platform knows when to use it. That approach also has a major infrastructure advantage. These models are small enough to run on older GPUs and, in some cases, consumer hardware, rather than requiring frontier chips, enormous memory footprints or hyperscale data centers."
Question:4
Inside AI: Your Ensemble Listening Model orchestrates over 100 specialized audio models simultaneously. From an engineering standpoint, how do you keep the compute overhead and latency low enough for real-time processing across millions of hours of audio?
Carter Huffman: "The key distinction is that Modulate leverages more than 100 specialized models available to the ensemble, but we do not blindly run all of them on every conversation, and our architecture tolerates delay between models which allows us to asynchronously incorporate new information and build understanding, even if results from one model are delayed for one inference step. The orchestration layer determines which models are appropriate for the task and routes work accordingly.
Where multiple signals are required, we run those models in parallel rather than creating a long serial pipeline. The combination of parallel processing, intelligent routing, and stateful orchestration is what allows us to deliver real-time analysis without the compute overhead you would expect from running 100+ large models in sequence.
The broader engineering philosophy is simple: use the smallest model that can solve the problem well, and only invoke the compute you actually need. That is how we can achieve extremely high efficiency while still analyzing conversations in real time and at very large scale."
Voice AI Enters the Real World
Social communications and online gaming platforms integrate the software to track child grooming conversations. The architecture uses specialized vocal markers to flag predatory conversational structures before harm escalates. It also provides advanced safety features for high-risk human customer service agents. The system applies dynamic voice masking to prevent callers from identifying agents across outside digital networks. This deployment diversity reflects a massive growing market demand from external developers.

Modulate is now expanding the team and infrastructure needed to meet growing demand from developers and partners building voice applications across security, customer experience, communications, AI agent supervision, and trust and safety.
Steve Jurvetson, Co-founder of Future Ventures and Board member of SpaceX AI, said, “Modulate has gained a significant technical lead in audio-native AI, and the market opportunity is expanding quickly. The team has proven these models in some of the most demanding voice environments in the world, and we're now seeing the need for that technology to expand well beyond where it started into AI agents, security, customer experience, and more. This investment will help Modulate move faster, grow the team, put its models into the hands of more developers and partners, and establish audio intelligence as a foundational layer of the AI stack.”
The company is actively developing a suite of custom software development kits (SDKs) and API integrations. These tools will expose sophisticated conversational metrics to development teams without forcing them to train underlying neural models. The expanded developer relations group will provide direct optimization support across multiple cloud infrastructure setups. The company will also introduce fresh deployment environments to support specialized edge computing scenarios. This architectural flexibility allows enterprise clients to run sensitive acoustic classification tasks locally without transmitting raw voice data over outside networks.
“Developers shouldn't have to rebuild the audio intelligence layer every time they create a new voice experience,” Huffman added. “Our mission is to build the models and infrastructure that let them focus on the application they want to create. The opportunity facing audio-native AI is expanding incredibly quickly. We’ve built the technology and proven it at scale, and this investment lets us grow the team and move faster to meet that demand.”
The broader macroeconomic environment indicates an accelerating shift toward multimodal computing across global business sectors. As digital voice agents handle increasingly sensitive financial transactions, verification integrity becomes paramount.
The proliferation of automated open source cloning software has lowered the technical barrier for cybercriminals. Standard legacy firewalls cannot isolate synthetic voice waves injected into a telephone system stream. By providing an explicit algorithmic layer focused entirely on behavioral characteristics, the startup addresses a deep systemic weakness in corporate network authentication practices.
About Modulate
Modulate is a frontier audio AI company building audio-native models that power the next generation of voice AI, enabling machines to understand the nuances in human conversation beyond the words being spoken. Its Velma platform is powered by Modulate’s Ensemble Listening Model (ELM) architecture to understand signals including emotion, tone, intent, synthetic speech and conversational behavior and create a powerful layer of audio intelligence for voice applications.
Founded by MIT alumni, Modulate’s technology is used across AI agents, fraud and deepfake detection, customer experience, trust and safety, and other emerging voice applications.