The Gap Between Humans and Machines Is ___ [Dr. Max Bartolo]

From the creator

SPONSOR MESSAGES: *** Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. Dr. Max Bartolo from Cohere discusses the gap between model capabilities and genuine robustness: why next-token prediction can produce impressive results yet still fail on slightly reformulated… --- TIMESTAMPS: 00:00:00 Model Reasoning and Consistency Verification 00:03:25 Influence Functions and Distributed Knowledge Analysis 00:10:28 AI Application Development and Model Deployment 00:14:24 AI Alignment and Human Feedback Limitations 00:20:15 Human Evaluation Challenges and Factuality Assessment 00:27:15 Cultural and Demographic Influences on Model Behavior 00:32:43 Adversarial Examples and Model Robustness 00:41:54 DynaBench and Dynamic Benchmarking Approaches 00:50:02 Benchmarking Challenges and Data-Centric Evaluation 00:55:15 Cohere Command A Development Process 01:00:26 Model Quantization and Performance Evaluation 01:05:18 Reasoning Capabilities and Training Progression 01:13:48 Context Windows and Enterprise Applications --- REFERENCES: person: [00:00:00] Max Bartolo Website https://www.maxbartolo.com/ company: [00:00:00] Cohere https://cohere.com/command [00:12:10] Command A Model https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 paper: [00:03:25] Procedural Knowledge in Pretraining Drives Reasoning in LLMs https://cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20 [00:04:15] Influence Functions in Machine Learning https://arxiv.org/abs/1703.04730 [00:08:05] Studying Large Language Model Generalization with Influence Functions https://arxiv.org/abs/2308.03296 [00:16:15] Human Feedback is not Gold Standard https://arxiv.org/abs/2309.16349 [00:27:15] The PRISM Alignment Dataset https://arxiv.org/abs/2404.16019 [00:32:50] Adversarial Examples Are Not Bugs, They Are Features https://arxiv.org/abs/1905.02175 [00:43:00] DynaBench: Rethinking Benchmarking in NLP https://aclanthology.org/2021.naacl-main.324.pdf [00:50:15] Sara Hooker on Compute Limitations https://arxiv.org/html/2407.05694v1 [00:53:25] DataPerf: Benchmarks for Data-Centric AI https://arxiv.org/abs/2207.10062 [01:04:35] DROP: A Reading Comprehension Benchmark https://arxiv.org/abs/1903.00161 [01:07:05] GSM8k https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k [01:09:30] ARC-AGI Challenge https://github.com/fchollet/ARC-AGI --- LINKS: Full Transcript: https://app.rescript.info/share/163bf5e7338685f635fc0b8d6920005c Download PDF transcript: https://app.rescript.info/api/public/sessions/7167a679366f97f8/pdf REFS: [00:03:10] Research at Cohere with Laura Ruis et al., Max Bartolo, Laura Ruis et al. https://cohere.com/research/papers/procedural-knowledge-in-pretraining-drives-reasoning-in-large-language-models-2024-11-20 [00:04:15] Influence functions in machine learning, Koh & Liang https://arxiv.org/abs/1703.04730 [00:08:05] Studying Large Language Model Generalization with Influence Functions, Roger Grosse et al. https://storage.prod.researchhub.com/uploads/papers/2023/08/08/2308.03296.pdf [00:11:10] The LLM ARChitect: Solving ARC-AGI Is A Matter of Perspective, Daniel Franzen, Jan Disselhoff, and David Hartmann https://github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf [00:12:10] Hugging Face model repo for C4AI Command A, Cohere and Cohere For AI https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 [00:13:30] OpenInterpreter https://github.com/KillianLucas/open-interpreter [00:16:15] Human Feedback is not Gold Standard, Tom Hosking, Max Bartolo, Phil Blunsom https://arxiv.org/abs/2309.16349 [00:27:15] The PRISM Alignment Dataset, Hannah Kirk et al. https://arxiv.org/abs/2404.16019 [00:32:50] How adversarial examples arise, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry https://arxiv.org/abs/1905.02175 [00:43:00] DynaBench platform paper, Douwe Kiela et al. https://aclanthology.org/2021.naacl-main.324.pdf [00:50:15] Sara Hooker's work on compute limitations, Sara Hooker https://arxiv.org/html/2407.05694v1 [00:53:25] DataPerf: Community-led benchmark suite, Mazumder et al. https://arxiv.org/abs/2207.10062 [01:04:35] DROP, Dheeru Dua et al. https://arxiv.org/abs/1903.00161 [01:07:05] GSM8k, Cobbe et al. https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k [01:09:30] ARC, François Chollet https://github.com/fchollet/ARC-AGI [01:15:50] Command A, Cohere https://cohere.com/blog/command-a [01:22:55] Enterprise search using LLMs, Cohere https://cohere.com/blog/commonly-asked-questions-about-search-from-coheres-enterprise-customers

Choose to Build with AI
Matched to Neural Networks

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.