800.553.8359 info@const-ins.com

AI voice models are having a moment, and Alex Smola is seeing it up close.

As the founder of Santa Clara, Calif.-based Boson AI, a startup that is set to release its first speech-to-speech model named Higgs RealTime, Smola says voice models have finally reached the threshold needed to drive the next leap in human-machine interaction.

A former distinguished scientist at Amazon and a leading machine learning researcher, he launched Boson after recognizing that the industry was moving beyond text-only interfaces. This is driven by the belief that multimodal AI, including audio and vision, delivers a far richer user experience. Beyond capability, he is also aiming for price competitiveness. Boson’s mission is to offer live audio products significantly cheaper to develop and deploy than rivals like OpenAI.

“What we have is about an order of magnitude more affordable [than competitors], and I would say it’s nonetheless very competent to use,” Smola told Fortune. Boson says their models are one-tenth of the cost of others. 

AI voice technology has become a primary front for OpenAI, Meta, and other industry leaders racing to build the next evolution of chatbots. Teams across the sector are focused heavily on full-duplex systems, or technology that enables natural, fluid back-and-forth conversations where users can interrupt the AI mid-sentence. OpenAI Chief Executive Sam Altman recently posted that he now talks with ChatGPT more than he texts with it, noting that the company’s “new voice model really crossed a threshold.”

Smola’s key differentiator for Boson is its “full-stack” model that he says makes their tech cheaper. The company says it allows enterprise customers to run systems within their own data centers and train custom voice and video models from scratch. That approach gives enterprise clients direct control over data residency and model execution in sensitive business environments, he said. In general, Smola has been an advocate for open-weight, customizable models, aligning with Meta and other open-source advocates. But since he’s also targeting business customers, he’s ultimately building proprietary models.  

The three-year-old startup has raised $70 million and is backed by Chinese entrepreneur Su Hua and the venture arm of Singapore-based investment firm Temasek. Smola is initially targeting clients across finance, telecommunications, healthcare, and insurance.

“A lot of machine interaction for customer support and sales… will become automated with voice agents,” he said. “They can be strictly superior [to humans].”

Aside from the high costs full-duplex audio can incur because it increases use of graphics processing units, a central hurdle in building voice interfaces is latency. While a one-second response lag is acceptable in text chat, that same pause feels awkward and disruptive in spoken conversation. To bridge the gap, Smola says Boson is training its AI for the nuances of the “human condition,” including the ability to parse emotional tone like cheerfulness or passive-aggressiveness, as well as rapid speech patterns.

Still, Boson faces formidable incumbents. Microsoft recently announced the latest iteration of its voice model, while OpenAI debuted GPT-Live, a family of full-duplex audio models. Meta’s new Muse Spark model is similarly optimized for voice, with the company claiming it “lets you talk naturally with the assistant—interrupt, switch topics, or swap languages.” While Microsoft is positioning its audio models around workplace productivity and OpenAI is bidding to become the default for app developers, Meta is targeting hardware integration across its lineup of wearable devices.

In the future, Smola envisions robots equipped with animated faces and souped-up voice models becoming household staples, à la the sci-fi blockbuster I, Robot. For now, the enterprise value lies in practical automation, he said, such as voice models maintaining total context during sales calls to generate real-time logs and personalized records faster than human teams can.

Ultimately, Smola sees embodied AI as the end goal. Robots combining reasoning, vision, and voice will “really be the ultimate technical frontier.”

See you tomorrow,

Sebastian Herrera
Email: sebastian.herrera@fortune.com
@SebasAHerrera
Submit a deal for the Term Sheet newsletter here.

Joey Abrams curated the deals section of today’s newsletter. Subscribe here.

This story was originally featured on Fortune.com