Deep Learning with PolyAI

Can AI really hear a call the way a person does?

Team PolyAI

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 21:33

PolyAI just launched Dialog-RSN-1, its first audio-native model, and your host Nikola Mrkšić sat down with its builder, Matt Henderson, to unpack why it’s a game-changer for building voice agents.

Most voice AI either flattens a call into a transcript and loses the audio, or goes fully speech-to-speech and gives up control of the voice. Dialog-RSN-1 does neither. It hears the raw audio directly, decides when to speak, and keeps text-to-speech separate so the voice stays under your control, all in under 300 milliseconds.

Nikola and Matt get into what audio-native really means, how auto-reasoning keeps it fast, and why it beats every other real-time model on quality and quickness. Hear the full episode, and see how PolyAI builds dialog agents that hear the whole call at https://poly.ai?utm_source=youtube&utm_medium=podcast&utm_campaign=podcast&utm_content=podcast

  • Follow PolyAI on LinkedIn
  • Watch this and other episodes of the Deep Learning pod on YouTube

People on this episode