AI Interpretability, Safety, and Meaning - Nora Belrose
SPONSOR MESSAGES: CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. https://centml.ai/pricing/ Nora Belrose, Head of Interpretability Research at EleutherAI, delivers a wide-ranging conversation that moves from the mathematical foundations of concept erasure in neural networks to fundamental questions about consciousness, AI safety, and Buddhist philosophy. The technical core centers on LEACE (LEAst-squares Concept Erasure), a method Belrose developed for surgically removing targeted information from neural network representations. She explains how LEACE emerged from connecting two prior approaches (RLACE and spectral attribute removal) through a mathematical equivalence proof, and demonstrates its applications in both fairness-oriented debiasing and interpretability research. A key finding: language models remain functional even after erasing part-of-speech information from every layer, suggesting robust reliance on redundant cues. Belrose then presents her ICML paper on simplicity biases in deep learning, showing that neural networks learn to exploit statistical moments in order -- first means, then covariances, then higher-order statistics. This has implications for understanding when and why concept erasure techniques may backfire against sufficiently deep models. The second half pivots to AI safety, where Belrose delivers a detailed critique of "counting arguments" used to predict AI misalignment. She argues these arguments rely on the principle of indifference applied to poorly-defined outcome spaces, drawing an analogy to an identical argument structure that would absurdly predict all neural networks must overfit. She connects this to broader questions about goal attribution, agency, and whether instrumental convergence arguments hold up under scrutiny. The conversation concludes with an exploration of 4E cognition, Evan Thompson's philosophy of mind, Belrose's departure from effective altruism, and her growing interest in Buddhist philosophy as a framework for thinking about meaning in a post-automation world. --- REFERENCES: Paper: [00:00:00] Episode Shownotes https://www.dropbox.com/scl/fi/38fhsv2zh8gnubtjaoq4a/NORA_FINAL.pdf?rlkey=0e5r8rd261821g1em4dgv0k70&st=t5c9ckfb&dl=0 [00:05:00] LEACE Paper https://arxiv.org/abs/2306.03819 [00:06:40] RLACE Paper https://arxiv.org/abs/2201.12091 [00:08:20] Spectral Attribute Removal https://arxiv.org/abs/2012.14424 [00:15:00] Pythia Models https://arxiv.org/abs/2304.01373 [00:20:30] LoRA https://arxiv.org/abs/2106.09685 [02:00:00] Holden Karnofsky https://forum.effectivealtruism.org/posts/T975ydo3mx4YnRv4J/ea-is-about-maximization-and-maximization-is-perilous Company: [00:01:37] CentML https://centml.ai/pricing/ [00:01:37] Tufa AI Labs https://tufalabs.ai/ [00:02:20] EleutherAI https://www.eleuther.ai/ Person: [00:02:20] Nora Belrose https://norabelrose.com/ [01:03:00] Evan Thompson https://evanthompson.me/ --- LINKS: Full Transcript: https://app.rescript.info/share/79d69cf24406cc36d8f7e8eee389e3ae Download PDF transcript: https://app.rescript.info/api/public/sessions/61e64e1737593802/pdf Nora Belrose: https://norabelrose.com/ https://scholar.google.com/citations?user=p_oBc64AAAAJ&hl=en https://x.com/norabelrose