It's Not About Scale, It's About Abstraction
MLST is sponsored by Tufa Labs: Are you interested in working on ARC and cutting-edge AI research with the MindsAI team (current ARC winners)? Focus: ARC, LLMs, test-time-compute, active inference, system2 reasoning, and more. Future plans: Expanding to complex environments like Warcraft 2 and Starcraft 2. Interested? Apply for an ML research position: benjamin@tufa.ai Francois Chollet, creator of Keras and the ARC-AGI benchmark, delivers his AGI-24 keynote on why scaling LLMs will not get us to AGI. He walks through concrete failure modes -- LLMs that break on trivial rephrasing of memorized problems, that pattern-match the Monty Hall problem without parsing the actual numbers, that solve Caesar ciphers only for key sizes found in online examples. The failures all point the same way: LLM performance tracks task familiarity, not task complexity. Chollet introduces his Kaleidoscope Hypothesis: the world looks infinitely complex on the surface, but it is built from a small set of repeating atoms of meaning. Intelligence, in his framing, is the process of mining experience to extract those atoms and recombining them to handle genuinely novel situations. This is what the ARC benchmark is designed to test -- abstraction and reasoning that cannot be memorized. The talk closes with a proposal: combine deep learning (good at perception and pattern recognition) with discrete program synthesis (good at precise, compositional reasoning). Neither approach alone gets there, but the hybrid might. Chollet points to early results on ARC from Ryan Greenblatt and others as evidence that the research community outside big labs may be where the next breakthrough comes from. --- TIMESTAMPS: 00:00:00 LLM Limitations and Composition 00:12:05 Intelligence as Process vs. Skill 00:17:15 Generalization as Key to AI Progress 00:19:59 Introduction to ARC-AGI Benchmark 00:26:10 The Kaleidoscope Hypothesis and Abstraction Spectrum 00:34:05 Limitations of Transformers and Program Synthesis 00:39:59 Applying Combined Approaches to ARC Tasks 00:44:20 State-of-the-Art Solutions and Future Directions --- REFERENCES: paper: [00:01:15] On the Measure of Intelligence https://arxiv.org/abs/1911.01547 [00:03:30] Embers of Autoregression https://arxiv.org/abs/2309.13638 [00:05:30] Monty Hall problem https://www.tandfonline.com/doi/abs/10.1080/00031305.1975.10479121 [00:06:20] LLM Training Dynamics Analysis https://arxiv.org/abs/2205.10770 [00:07:33] GPT-4 Technical Report https://cdn.openai.com/papers/gpt-4.pdf [00:10:20] Faith and Fate: Limits of Transformers on Compositionality https://arxiv.org/abs/2305.18654 [00:10:25] The Reversal Curse in LLMs https://arxiv.org/abs/2309.12288 [00:10:52] LM-Polygraph: Uncertainty Estimation for LLMs https://arxiv.org/abs/2311.07383 [00:20:34] Baldur: Whole-Proof Generation https://arxiv.org/abs/2303.04910 [00:34:00] Core Knowledge in Infants https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf [00:44:20] Hypothesis Search with LLMs for ARC https://arxiv.org/abs/2309.05660 tool: [00:20:10] ARC-AGI GitHub Repository https://github.com/fchollet/ARC-AGI [00:22:15] ARC Prize https://arcprize.org/ book: [00:33:30] Thinking, Fast and Slow https://www.amazon.com/Thinking-Fast-Slow-Daniel-Kahneman/dp/0374533555 --- LINKS: Full Transcript: https://app.rescript.info/share/c8b5bacdf1ffefab4f65060edc295d4a Download PDF transcript: https://app.rescript.info/api/public/sessions/b537d0b92ae48338/pdf [0:20:10] ARC-AGI: GitHub repository (François Chollet) https://github.com/fchollet/ARC-AGI