How Machines Learned to Talk – OpenCV Live! 227

Video by OpenCV via YouTube
How Machines Learned to Talk - OpenCV Live! 227

Akshat Mandloi, co-founder of Smallest.ai, joins OpenCV Live to answer a question most of us have wondered on hold with customer service: why can you still tell it’s a bot?

His answer isn’t "the model’s too small." After three generations of voice AI, less than 1% of the voice market is automated, and Akshat argues the problem is structural. Today’s agents listen, then think, then speak. People do all three at once, and interrupt each other while they’re at it.

He’ll walk through how the field got here, from the old ASR-to-LLM-to-TTS pipeline to full-duplex models that can hear while they talk, and why measuring "does this sound human" is still harder than it looks. Then the good part: how Smallest.ai built a speech model that scores 96% on Big Bench Audio and an agent that holds its own against frontier models at about a twentieth of the size, and where they’re heading next with a model that predicts a conversation instead of reacting to it.

OpenCV is a 501(c)(3) registered non-profit in the United States. See how you can support open source CV & AI: http://opencv.org/support/

Watch along for your chance to win during our live trivia segment, and participate in the live Q&A session with questions from you in the audience.

Become a paid member of the channel to help us make more episodes https://www.youtube.com/channel/UCkrcW82Y2kbgU-U9RaYfgxw/join

Got a cool project of your own? Send it to us and you may be featured https://www.jotform.com/form/233105358823151

Thumbnail by Natalia de la Rosa natdlrs.com

Source