MoE Token Routing Explained: How Mixture of Experts Works (with Code)

From the creator

This video dives deep into Token Routing, the core algorithm of Mixture of Experts (MoE) models. Slides: https://huggingface.co/ariG23498/moe-routing-algorithm/blob/main/MoE_routing_algorithm.pdf Colab Notebook: https://huggingface.co/ariG23498/moe-routing-algorithm/blob/main/routing_mechanism.ipynb Chapter Timestamp: Introduction: 00:00 Laying the Foundation for Mixture of Experts (MoE): 00:09 Focus on Token Routing: 00:50 What is a Mixture of Experts Layer?: 02:36 Problem Statement and Configurations: 04:48 Compute Router Logits: 08:31 Sparsity and Selecting Top K Experts: 10:54 Normalizing Logits to Router Probabilities: 12:43 Slot Selection: 14:39 Dropping Oversubscribed Tokens: 16:51 Updated Normalized Token Weights: 20:36 Updated Slot Selection and Token Slots: 21:34 Final Weight Matrix Construction: 24:35 Conclusion: 32:41 Correction: As @denisflavius5365 rightly suggested the router matrix in the slides should be of the shape 3x4 and not 4x4.

Choose to Build with AI
Matched to Neural Networks

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.