Speech Recognition Is Not a Solved Problem — Pavan Muddireddy

From the creator

Pavankumar Reddy Muddireddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why, after everything that has landed in the last two years, the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system. IN PARTNERSHIP WITH MISTRAL AI: --- This episode was produced in partnership with Mistral AI. Mistral AI: https://mistral.ai/ --- The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, fed into the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model changes shape into a dual-stream decoder that consumes audio and emits text at the same time, at a configurable target delay down to 160ms, with slower streams running in parallel for anything that can afford to wait for more context. On the generation side, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi's split of semantic and acoustic codebooks, and places finite scalar quantisation and flow matching in it. Tim presses on the engineering priors underneath: why a mel spectrogram instead of the raw waveform, what noise augmentation actually buys, and the point at which acoustic overfitting becomes somebody's fine-tuning problem. Then the failure modes. Speaker diarisation is emitted autoregressively as part of the transcript rather than by a separate head, which makes streaming diarisation fragile in a particular way — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, a single out-of-distribution mistake can compound into looping or skipped segments, which is what DPO is there to correct: the negative supervision that pre-training and SFT cannot provide. The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions do not describe a solved problem, they describe scaffolding, with a sharp quality drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable — fine-tune the ASR for your acoustics, log what compliance demands, keep it on your own hardware. The interface has its own limit: voice alone is cognitive debt, because absorbing information and deciding in one serial audio stream is much harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one. --- TIMESTAMPS: 00:00:00 Cold open: the state changed 00:00:46 Why Mistral moved into audio 00:09:27 Inside Voxtral: the trunk, the encoder and dual streams 00:20:22 Speech that works in real time 00:30:52 How a voice becomes tokens 00:39:59 Flow matching, FSQ and the new codec 00:52:51 When speech models lose the speaker 01:03:23 Correcting hallucinations with preferences 01:12:12 Controlling synthetic speech 01:20:06 Why cascades still win 01:29:25 Speech in the wild 01:33:46 Audio models as interfaces 01:37:54 Why voice still needs a screen --- REFERENCES: paper: [00:01:42] Mistral 7B https://arxiv.org/abs/2310.06825 [00:09:38] Voxtral https://arxiv.org/abs/2507.13264 [00:14:41] Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) https://arxiv.org/abs/2212.04356 [00:19:11] Voxtral Realtime https://arxiv.org/abs/2602.11298 [00:21:52] Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling https://arxiv.org/abs/2509.08753 [00:30:52] Voxtral TTS https://arxiv.org/abs/2603.25551 [00:32:38] SoundStream: An End-to-End Neural Audio Codec https://arxiv.org/abs/2107.03312 [00:34:59] Flow Matching for Generative Modeling https://arxiv.org/abs/2210.02747 [00:37:03] High Fidelity Neural Audio Compression (EnCodec) https://arxiv.org/abs/2210.13438 [00:37:42] Moshi: a speech-text foundation model for real-time dialogue (Mimi) https://arxiv.org/abs/2410.00037 [00:39:05] Finite Scalar Quantization: VQ-VAE Made Simple https://arxiv.org/abs/2309.15505 [01:03:33] Direct Preference Optimization: Your Language Model is Secretly a Reward Model https://arxiv.org/abs/2305.18290 benchmark: [01:15:40] ElevenLabs v3 and Flash comparison in Voxtral TTS https://arxiv.org/abs/2603.25551 dataset: [00:46:14] Mozilla Common Voice datasets https://commonvoice.mozilla.org/en/datasets organization: [00:01:24] Mistral AI https://mistral.ai/ [00:50:47] Hugging Face https://huggingface.co/

Choose to Build with AI
Matched to AI Video Generation

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.