New Research: LMArena is "rigged"
We expose the serious problems with Chatbot Arena, the AI industry's most influential leaderboard. We showcase the recent paper "The Leaderboard Illusion" which shows that it's actually being gamed by the biggest players in tech. We comment on the recent revelation that Mark Zuckerberg openly admitted that Meta manipulated the arena by fine-tuning models specifically to top the charts. The paper uncovers how select companies receive preferential treatment, including extensive private testing and the ability to retract underperforming models. These companies also get a massive data advantage, as their models are sampled far more often, allowing them to train on arena data to boost their scores. [HEDGING] Meta's actions were actually useful to expose a serious flaw in the arena and Zuck 100% owned it, and was super transparent about it afterwards. SPONSOR MESSAGES: ========= Prolific: Quality data. From real people. For faster breakthroughs. https://prolific.com/mlst?utm_campaign=98404559-MLST&utm_source=youtube&utm_medium=podcast&utm_content=script-4 Tufa AI Labs are hiring for ML Engineers and a Chief Scientist in Zurich/SF https://tufalabs.ai/ ========= Interactive Transcript: https://app.rescript.info/public/share/I8k9saMwSQbkpL7rbZRiWCKnFHzgSLBi2atvnVACVm8 TOC: 00:00:00 - The Serious Problem with Chatbot Arena 00:01:45 - Mark Zuckerberg Admits to "Hacking" the Leaderboard 00:03:26 - Goodhart's Law: Why the Benchmark is Broken 00:04:00 - Andrej Karpathy on Judging Subtle Model Improvements 00:06:45 - The Origin Story of Chatbot Arena 00:10:00 - How ELO Ratings Work (and Their Flaws) 00:15:15 - The Explosive "Leaderboard Illusion" Paper by Cohere 00:16:45 - Finding #1: Preferential Private Testing 00:18:00 - Finding #2: Unfair Data Advantage for Big Tech 00:19:15 - Finding #3: Massive Gains from Fine-Tuning on Arena Data 00:20:45 - Why Some Models Are Sampled Overwhelmingly More Than Others 00:22:15 - Cohere's Proposed Solutions to Fix the Arena 00:23:45 - Analysis: How Repetitive Are User Prompts? 00:25:56 - Final Thoughts & The Arena's Controversial Response REFS: The Leaderboard Illusion [2025] Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D'Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, Sara Hooker https://arxiv.org/abs/2504.20879 LM Arena response: https://blog.lmarena.ai/blog/2025/our-response/#:~:text=Published&text=Recently%2C%20a%20writeup%20titled%20%E2%80%9CThe,ongoing%20discussions%20with%20the%20authors. Elo Uncovered: Robustness and Best Practices in Language Model Evaluation [2023] https://arxiv.org/pdf/2311.17295 Meriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker, Marzieh Fadaee Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference [2024] https://arxiv.org/abs/2403.04132 Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, Ion Stoica Too much efficiency makes everything worse: overfitting and the strong version of Goodhart's law https://sohl-dickstein.github.io/2022/11/06/strong-Goodhart.html Jascha Sohl-Dickstein Primary/Senior Authors-- https://x.com/singhshiviii https://x.com/mziizm https://x.com/sarahookr