NEURAL NETWORKS ARE WEIRD! - Neel Nanda (DeepMind)

From the creator

SPONSOR MESSAGES: *** CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. https://centml.ai/pricing/ Neel Nanda leads the mechanistic interpretability team at Google DeepMind. At 26, he's become one of the most prominent researchers working on the question of what's actually going on inside neural networks -- systems that can win IMO medals and write complex software, but which nobody actually designed or understands. This nearly four-hour conversation is a deep technical dive into the field. Nanda explains why machine learning is fundamentally weird: we produce artifacts that do impressive things, but unlike conventional software, no one wrote the code or planned the architecture. His team's goal is reverse-engineering these systems by finding the internal structures and algorithms that emerge during training. The discussion covers the mechanics of sparse autoencoders at length -- how they decompose model activations into interpretable feature vectors, the mathematical foundations (ReLU vs TopK activation functions), scaling laws for feature learning, and the engineering challenges of running them at the scale of frontier models. Nanda walks through the Golden Gate Claude experiment (amplifying a single feature to make Claude obsessed with the Golden Gate Bridge), induction heads (the circuits responsible for in-context learning), and activation patching as a causal intervention technique. On AI safety, Nanda is pragmatic. He argues that mechanistic interpretability gives us genuine empirical evidence about questions that are otherwise stuck in philosophical debate -- do models have goals? Do they deceive? He also discusses the limitations: sparse autoencoders haven't yet demonstrated capabilities beyond what fine-tuning already achieves, and at sufficient model complexity, models could potentially facade interpretability measurements. The conversation covers his path from pure maths at Cambridge through Anthropic to DeepMind, and why he thinks hands-on coding matters more than reading papers for new researchers entering the field. --- REFERENCES: person: [00:00:00] Neel Nanda - Personal Website https://www.neelnanda.io/ tool: [00:35:00] TransformerLens https://github.com/TransformerLensOrg/TransformerLens paper: [01:00:31] A Mathematical Framework for Transformer Circuits https://transformer-circuits.pub/2021/framework/index.html [01:01:40] In-context Learning and Induction Heads https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html [01:21:06] Scaling Monosemanticity https://transformer-circuits.pub/2024/scaling-monosemanticity/ [01:33:27] Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 --- LINKS: Full Transcript: https://app.rescript.info/share/acb415fa59ae2d2909d60d761c8f4ff4 Download PDF transcript: https://app.rescript.info/api/public/sessions/c3a4bf1e32a46ce7/pdf NEEL NANDA: https://www.neelnanda.io/ https://scholar.google.com/citations?user=GLnX3MkAAAAJ&hl=en https://x.com/NeelNanda5

Choose to Build with AI
Matched to Neural Networks

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.