Why Kimi Suddenly Started Sounding So Much Like Claude? — Ilia Shumailov & Alexander Panfilov
Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs. The core bug sounds deceptively simple: providers return encrypted reasoning state so conversations can be resumed or forked. But those blobs can be replayed across users and sibling models. A smaller model can ask the provider to decrypt the trace, then repeat the hidden reasoning in plain text. The discussion covers leaked private data, a broadly reusable jailbreak, poisoned agent traces, chain-of-thought monitoring, responsible disclosure, and possible defenses. Ilia Shumailov is an AI and security researcher, formerly at Google DeepMind, who completed his Cambridge PhD under Ross Anderson. Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, working on AI safety, adversarial machine learning, and LLM red-teaming. They close by separating the demonstrated jailbreaking threat from ordinary benign distillation, and by arguing for controlled experiments over sweeping claims. --- TIMESTAMPS: 00:00:00 Intro montage 00:01:33 Portable encrypted thought and decoded reasoning 00:24:55 How the attack works and what it means 00:39:04 Doom, defense, and scientific restraint --- REFERENCES: paper: [00:00:00] Stealing Reasoning Traces from Proprietary LLM APIs https://arxiv.org/abs/2608.09867 [00:09:22] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety https://arxiv.org/abs/2507.11473 [00:11:30] Reasoning Models Don’t Always Say What They Think https://www.anthropic.com/research/reasoning-models-dont-say-think [00:37:22] PostTrainBench: Can LLM Agents Automate LLM Post-Training? https://arxiv.org/abs/2603.08640 [00:41:02] Large-scale online deanonymization with LLMs https://arxiv.org/abs/2602.16800 other: [00:09:28] OpenAI and Hugging Face partner to address security incident during model evaluation https://openai.com/index/hugging-face-model-evaluation-security-incident/ [00:10:22] Claude, GPT, and Gemini All Struggle to Evade Monitors https://metr.org/notes/2025-08-22-claude-gpt-gemini-struggle-evade-monitors/ tool: [00:42:08] Isabelle proof assistant https://isabelle.in.tum.de/ --- RESCRIPT: https://app.rescript.info/share/07fc38276e0823dc9b8986c32e202c7f