MIGHT THE ROBOTS TAKE OVER? [Prof. Yoshua Bengio]
SPONSOR MESSAGES: *** CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. https://centml.ai/pricing/ Turing Award winner Professor Yoshua Bengio sits down with Tim Scarfe for a wide-ranging discussion on AI safety, the architecture of intelligence, and what happens when machines become as capable as our best researchers. Bengio argues that the current race toward agentic AI systems carries fundamental risks that most people underestimate. He lays out how instrumental convergence and reward tampering could lead systems to develop goals misaligned with human interests -- not through malice, but through the basic logic of optimization. His proposed alternative: powerful AI tools that function as scientific oracles rather than autonomous agents, systems that can revolutionize medicine and science without needing goals of their own. The conversation covers international AI governance and the game theory driving the US-China competition, the military implications of frontier AI capabilities, and why hardware-enabled governance may be the most promising path toward verifiable AI treaties. Bengio draws direct parallels to nuclear nonproliferation and warns that data centers will become strategic military assets. On the technical side, Bengio discusses the return of recurrent architectures (the 'Were RNNs All We Needed?' paper), his work on GFlowNets for probabilistic inference with discrete structures, a complexity-based theory of compositionality that tries to formalize what was previously just intuition, and the question of whether System 2 reasoning needs to be designed from first principles rather than bolted onto existing architectures. Throughout, Bengio repeatedly emphasizes what he does not know -- a deliberate stance he contrasts with the overconfidence he sees driving dangerous policy decisions. --- REFERENCES: paper: [00:00:15] AI Risk Statement https://www.safe.ai/work/statement-on-ai-risk [00:23:10] Reward Tampering and AI Safety https://hdsr.mitpress.mit.edu/pub/w974bwb0 [00:44:30] Can a Bayesian Oracle Prevent Harm? https://arxiv.org/abs/2408.05284 [00:52:00] Hardware-Enabled AI Governance Memo https://yoshuabengio.org/wp-content/uploads/2024/08/FlexHEG-Memo_August-2024.pdf [01:33:20] GFlowNet Foundations https://arxiv.org/abs/2111.09266 [01:35:05] Complexity-Based Compositionality Theory https://arxiv.org/abs/2410.14817 [01:37:50] Discrete Attractor States in Neural Systems https://arxiv.org/abs/2302.06403 person: [00:03:50] Professor Yoshua Bengio https://yoshuabengio.org/ tool: [00:40:45] Munk Debate on AI Existential Risk https://munkdebates.com/debates/artificial-intelligence [00:56:07] 2018 Turing Award https://awards.acm.org/about/2018-turing [01:12:35] EU AI Act Code of Practice https://digital-strategy.ec.europa.eu/en/news/meet-chairs-leading-development-first-general-purpose-ai-code-practice --- LINKS: Full Transcript: https://app.rescript.info/share/1c4f010852d0db9a34dbeb043c0da26b Download PDF transcript: https://app.rescript.info/api/public/sessions/baff3f8846202250/pdf Yoshua Bengio: https://x.com/Yoshua_Bengio https://scholar.google.com/citations?user=kukA0LcAAAAJ&hl=en https://yoshuabengio.org/ https://en.wikipedia.org/wiki/Yoshua_Bengio