Anthropic Researcher Quits With A Terrifying AI Warning

From the creator

Link to our newsletter: https://bitbiased.ai/ Anthropic researcher Jacob Coxon just resigned after saying his former employer is “far and away” the most responsible company in the AI race. He also says Anthropic isn’t cutting safety corners. So why walk away before his equity even started vesting? The answer is more complicated than the headlines calling Coxon Anthropic’s “safety chief” or a “whistleblower.” He was a pretraining researcher who worked inside frontier AI development at both OpenAI and Anthropic, including research tied to GPT-4o. His warning isn’t that Anthropic has secretly abandoned safety. It’s that competitive pressure could eventually make even the most cautious AI lab move faster than it believes is safe. And the timing matters. Five days earlier, OpenAI released GPT-6 Astra, posting major benchmark results including 99.9% on ARC-AGI-3 and 97.6% on FrontierMath Tier 4. OpenAI president Greg Brockman said it was “not unreasonable” to feel we are entering the AGI era — but OpenAI’s official launch materials did not formally classify Astra as AGI. The more consequential detail may be cybersecurity. Under OpenAI’s Preparedness Framework, Astra was classified “Critical” for cyber capability after controlled tests where it autonomously discovered vulnerabilities and constructed exploit chains against hardened targets. But that comparison needs context. An independent evaluator recorded zero successful attacks against fully hardened targets, and Astra failed all seven of the hardest “Elite” challenges it faced. “Critical” describes a specific capability threshold — not a model capable of hacking anything on demand. That tension leads directly to Coxon’s deeper concern: recursive AI development. Frontier models are already getting better at coding, debugging machine-learning systems, optimizing training infrastructure, and automating pieces of AI research. If AI increasingly helps build the next generation of AI, development cycles could accelerate — potentially creating systems that become harder to monitor or control. But the evidence also shows an important limitation. OpenAI’s own safety evaluation says Astra does NOT reach its “High AI Self-Improvement” threshold and cannot independently design and execute frontier-scale AI training. On RE-Bench, AI agents can outperform humans on short research tasks, while humans regain a major advantage when given much longer time horizons. So the real argument isn’t whether runaway self-improvement has already arrived. It’s what happens if competitive pressure pushes AI labs toward that point faster than their safety processes can handle. Anthropic itself has argued that voluntary corporate promises may ultimately be insufficient and that governments could need authority to stop deployments presenting serious catastrophic risks. Meanwhile, an autonomous-agent cybersecurity incident involving an early Claude Opus build illustrates the paradox: increasingly capable systems can create unexpected failures, while responsible monitoring can still catch and investigate them. Coxon’s resignation doesn’t provide a smoking gun or a neat ending. Instead, it raises a harder question: if even researchers inside the companies they consider most responsible are worried about where the race leads, who ultimately decides how quickly frontier AI should move? CHAPTERS 00:00 Why an Anthropic Researcher Walked Away 01:01 He's Not Who the Headlines Say He Is 02:30 What Happened Five Days Earlier 03:46 The Number That Should Actually Worry You 05:04 The Fear Underneath the Fear 07:12 So What's Actually Being Argued About? 08:39 The Ending Coxon Doesn't Have #anthropic #openai #gpt6astra #aisafety #artificialintelligence

Choose to Build with AI
Matched to GPT-6 Astra

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.