Intro to Mixture of Experts | Aritra Roy Gosthipaty | HF Podcast #2

From the creator

In this episode, Alejandro sits down with Aritra Roy Gosthipaty from the Hugging Face Transformers team to talk about mixture-of-experts models, why dense models still matter, how synthetic data changes training, and what coding agents are changing for working engineers. They discuss Mixtral, DeepSeek-V2, Switch Transformers, vLLM, Inference Providers, TinyAya, data curation, local inference limits, and how practitioners should think about coding with agents without losing core engineering skill. More from Aritra: - Hugging Face profile — https://huggingface.co/ariG23498 Chapters 00:00 Meet Aritra Roy Gosthipaty 00:45 How Aritra joined Hugging Face 03:05 What a Developer Advocate on Transformers works on 04:00 What mixture-of-experts models are 08:26 Why MOEs matter now 11:36 Where dense models still win 15:00 Where to start learning and training MOEs 18:44 Synthetic data, data engines, and data quality 22:18 Why MOEs are still hard to run locally 23:11 How coding tools changed engineering work 25:49 Do agents weaken creativity and skill? 28:29 Should beginners rely on coding agents? 33:21 What coding will look like in a year 35:25 The biggest recent AI wow moments If you enjoyed the episode, subscribe for more conversations about open models, infrastructure, and the future of AI. --- Sources / References • Aritra Roy Gosthipaty — https://huggingface.co/ariG23498 • Mixture of Experts Explained — https://huggingface.co/blog/moe • Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — https://arxiv.org/abs/1701.06538 • vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — https://arxiv.org/abs/2309.06180 • DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — https://arxiv.org/abs/2405.04434 • Mixtral of Experts — https://arxiv.org/abs/2401.04088 • Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — https://arxiv.org/abs/2101.03961 • Hugging Face Inference Providers — https://huggingface.co/docs/inference-providers/index • Unsloth Docs — https://unsloth.ai/docs • MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases — https://arxiv.org/abs/2402.14905 • Tiny Aya: Bridging Scale and Multilingual Depth — https://arxiv.org/abs/2603.11510 • Aya 23: Open Weight Releases to Further Multilingual Progress — https://arxiv.org/abs/2405.15032 • Self-Instruct: Aligning Language Models with Self-Generated Instructions — https://arxiv.org/abs/2212.10560 • The Curse of Recursion: Training on Generated Data Makes Models Forget — https://arxiv.org/abs/2305.17493 • DeepSeek-V3 Technical Report — https://arxiv.org/abs/2412.19437

Choose to Build with AI
Matched to AI Podcast

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.