Benchmarking LLMs at the Game Of Science (Eleusis)
A card game ♠️♥️ to benchmark AI models at scientific discovery 🔬🧪 📝 Blog post : https://huggingface.co/spaces/huggingface/eleusis-benchmark 00:00 LLMs and science 01:55 The Eleusis Game 04:09 Our benchmark 08:35 Performance comparison 11:49 Overcaution vs Recklessness 14:05 The Key Chart 17:03 Calibration of confidence 20:40 LLMs and Occam's razor 21:58 Conclusion and learnings