EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception
Inside AsembleAI: DeepTech, AI & Science by Mac & Sam
Episode notes
Recorded live at the Ai4 Podcast Pavilion, Sam wraps Day One with Aditya Grover, Co-Founder & CTO of Inception, on why the next generation of LLMs won't look anything like the ones we use today.
What's Covered:
"Every Millisecond Matters" — Why latency, not intelligence, is the real bottleneck holding back voice agents and multi-step AI agents alike.
How Mercury Actually Generates Text — Instead of predicting one token at a time like every autoregressive model, Mercury generates a rough draft of the full response and refines it into coherence — diffusion, applied to language instead of images.
Solving Voice AI's Impossible Tradeoff — Fast-but-lower-quality, or high-quality-but-too-slow: Aditya explains how Mercury 2 finall ...