Talking Turns: Benchmarking Audio Foundation Models
The paper addresses a key challenge in conversational AI: making interactions with voice assistants feel natural and interactive. Current systems often use simple methods, like waiting for a period of silence, to decide when to speak. However, human conversation is much more complex, involving a fluent succession of turns, subtle cues, interruptions, and minimal long silences or overlapping speech.... The authors propose a novel evaluation protocol to assess an AI's turn-taking capabilities. Their goal is to measure if an AI understands when to listen, speak, interrupt, or provide feedback (like "uh-huh") in a way that mimics natural human-human conversation. They use this protocol to test existing spoken dialogue systems and other audio FMs, revealing significant room for improvement