This video covers the latest wave of open audio tooling, from Mistral's Voxtral 4B text-to-speech model to Cohere Transcribe for speech recognition and the Hugging Face infrastructure used to run large-scale transcription workflows. It walks through live demos, browser-based transcription with Transformers.js, and a practical UV-script pipeline built on storage buckets, HF Mount, and HF Jobs. If you're building speech apps or batch transcription systems, this is a fast overview of the current open stack.
---
Demo Links
👉 Voxtral TTS: https://huggingface.co/spaces/mistralai/voxtral-tts-demo
👉 Cohere Transcribe: https://huggingface.co/spaces/CohereLabs/Cohere-Transcribe-WebGPU
👉 UV scripts for transcription: https://huggingface.co/datasets/uv-scripts/transcription
---
🤓 Topics Covered
- Voxtral 4B text-to-speech
- Cohere Transcribe speech-to-text
- Hugging Face audio pipelines
---
⏱️ Timestamps
0:00 Open audio models and demos
2:44 What Hugging Face storage buckets are
3:39 How HF Mount works
4:02 HF Jobs and wrap-up
Choose to Build with AI
Matched to Voice and Audio
AI Maker Residence 3
The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code.
Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.
◆ Fri 09 Oct 2026◆ KOKO Cafe, London◆ With Nick Sarafa