Why Speech Recognition Was Solved Before Conversation Was — Shawn Wen

From the creator

Tsung-Hsien (Shawn) Wen, CTO of PolyAI, tells Tim Scarfe why voice agents are harder than text agents. Voice adds time, and a good conversation depends on adapting to the person on the line, not just on reasoning to the best answer. Shawn describes an audio-native model (Dialog-RSN-1) that first predicts a turn-taking signal, then replies in text with citations, and writes the transcript last so enterprises can audit it. Along the way: training on real, noisy calls with synthetic noise added, and why over-cleaned audio made the new model worse. Latency, and what a voice agent should do while it thinks. Why a voice with a hint of regional accent beats a generic one. Why public benchmarks fall short for voice, why enterprises want to own their agent harness, and whether behaviour belongs in the harness or in the weights. The last stretch is about working with agents: cognitive debt, the shift from producing content to checking it, Wispr Flow, building tools that agents can use, and whether slop is in the eye of the reader. This episode was produced in partnership with PolyAI. https://poly.ai CHAPTERS 0:00 Why voice agents are harder than text 4:27 What enterprises want, and why PolyAI built its own model 9:04 How an audio-native voice model works 15:21 Training data, spectrograms and synthetic noise 20:36 The cocktail party problem and the future of turn-taking 25:21 Latency, adaptive reasoning and keeping callers' trust 31:33 Voices, personality and the uncanny valley 36:32 How do you benchmark a voice agent? 42:00 Harness engineering and owning the intelligence 45:09 Well-specified problems and auditable agents 49:46 Weight adaptation and cognitive debt 56:31 Agents at work: Wispr Flow, voice and tool building 1:02:50 The next decade of voice, and what counts as slop REFERENCES The Bitter Lesson: http://www.incompleteideas.net/IncIdeas/BitterLesson.html [9:05] Retrieval-augmented generation: https://arxiv.org/abs/2005.11401 [13:23] Mel scale: https://en.wikipedia.org/wiki/Mel_scale [17:16] Victor Zue: https://en.wikipedia.org/wiki/Victor_Zue [17:45] Cocktail party effect: https://en.wikipedia.org/wiki/Cocktail_party_effect [20:39] Speaker diarisation: https://en.wikipedia.org/wiki/Speaker_diarisation [21:16]

Choose to Build with AI
Matched to AI Agents

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.