This is what happens when you let AIs debate
Akbir Khan, AI researcher and ICML 2024 Best Paper winner, joins Tim Scarfe to discuss his groundbreaking work on using debate between language models to improve AI truthfulness and oversight. Khan explains how pitting two LLMs against each other in structured arguments helps non-expert judges arrive at more accurate answers than simply querying a single model — a result with profound implications for supervising AI systems that may eventually surpass human capabilities. The conversation moves through the mechanics of scalable oversight and the "sandwiching protocol" that operationalises the problem of checking entities smarter than their supervisors. Khan describes how debate naturally surfaces the cruxes of disagreements, making complex expert judgments more accessible to laypeople. The discussion broadens into the relationship between intelligence and agency, the risks of deceptive alignment and reward tampering, and whether the current trajectory of AI development constitutes the early stages of a Cambrian explosion in artificial minds. Khan and Scarfe also explore open-ended AI systems, Kenneth Stanley's arguments against objective-driven optimization, and the philosophical terrain mapped by thinkers like Aaron Sloman and Francois Chollet on the space of possible minds and measuring intelligence. The episode closes with a frank exchange on cultural evolution, memetics, and whether the intelligence that matters most for AI safety is the kind we can measure. --- REFERENCES: person: [00:00:00] Akbir Khan https://akbir.dev/ other: [00:00:00] MLST Episode Shownotes PDF https://www.dropbox.com/scl/fi/sjekivbg3ok6qugsv2p1u/AkbirKhan.pdf?rlkey=ewiyvq0aq7mjvql4u7os0jos2&st=vblhp7af&dl=0 [00:06:05] OpenAI Superalignment Team https://openai.com/index/introducing-superalignment/ [00:19:10] DeepMind Responsible AI https://deepmind.google/about/responsibility-safety/ paper: [00:00:40] Akbir Khan et al. - Debating with More Persuasive LLMs https://arxiv.org/html/2402.06782v3 [00:08:10] Sam Bowman - Scalable Oversight in AI Systems https://arxiv.org/abs/2211.03540 [00:10:35] Sam Bowman - Artificial Sandwiching Protocol https://www.alignmentforum.org/posts/nekLYqbCEBDEfbLzF/artificial-sandwiching-when-can-we-test-scalable-alignment [00:14:35] Janus - Simulators https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators [00:21:30] Eliezer Yudkowsky - AI FOOM Debate https://intelligence.org/files/AIFoomDebate.pdf [00:21:45] Sammy Martin - Discontinuous AI Progress https://www.alignmentforum.org/posts/5WECpYABCT62TJrhY/will-ai-undergo-discontinuous-progress [00:24:35] Nora Belrose - Counting Arguments vs AI Doom https://www.lesswrong.com/posts/YsFZF3K9tuzbfrLxo/counting-arguments-provide-no-evidence-for-ai-doom [00:25:35] Evan Hubinger - Deceptive Alignment https://www.lesswrong.com/posts/zthDPAjh9w6Ytbeks/deceptive-alignment [00:26:50] Anthropic - Reward Tampering https://www.anthropic.com/research/reward-tampering [00:34:58] Ryan Greenblatt et al. - AI Control https://arxiv.org/pdf/2312.06942 [00:37:20] Aaron Sloman - The Space of Possible Minds https://www.cs.bham.ac.uk/research/projects/cogaff/sloman-space-of-minds-84.pdf [00:38:45] Francois Chollet - On the Measure of Intelligence https://arxiv.org/abs/1911.01547 [00:42:45] Jonathan Cook et al. - Artificial Generational Intelligence https://arxiv.org/abs/2406.00392 video: [00:03:28] Yann LeCun on Machine Learning Debates https://www.youtube.com/watch?v=OKkEdTchsiE book: [00:16:35] Thomas Suddendorf - The Gap https://www.amazon.in/GAP-Science-Separates-Other-Animals/dp/0465030149 [00:32:35] Kenneth Stanley - Why Greatness Cannot Be Planned https://www.amazon.co.uk/Why-Greatness-Cannot-Planned-Objective/dp/3319155237 [00:42:30] Richard Dawkins - The Selfish Gene https://www.amazon.co.uk/Selfish-Gene-Richard-Dawkins/dp/0192860925 --- LINKS: Full Transcript: https://app.rescript.info/share/c75b45c237ad52154d276851af8f812b Download PDF transcript: https://app.rescript.info/api/public/sessions/571abed49dab3e7f/pdf Akbir Khan: https://x.com/akbirkhan https://akbir.dev/