GPT-6 Astra Just Made the AGI Debate REAL
Link to our newsletter: https://bitbiased.ai/ OpenAI’s GPT-6 Astra scored an almost unbelievable 99.95% on ARC-AGI-3 — a benchmark designed to test whether AI can learn genuinely unfamiliar tasks. But change the testing harness, and that same model drops to 62.71%. That gap might tell us more about the state of AGI than the headline score itself. GPT-6 Astra is OpenAI’s new flagship model, and its capabilities are difficult to dismiss. It can operate real software, build and debug programs, work inside Blender, Unreal Engine and KiCad, browse the web, and tackle research-level mathematics. Under a controlled FrontierMath Erdős evaluation, Astra even solved two open mathematical problems that competing frontier models failed to solve. But the deeper question is whether extraordinary capability is the same thing as general intelligence. On ARC-AGI-3, Astra’s 99.95% result came through OpenAI’s Provider Adapter, which preserves internal reasoning state and manages context automatically. Through ARC Prize’s provider-neutral Standard harness, it scored 62.71%. ARC creator François Chollet still described Astra’s interactive reasoning as a “step-function change” in capability — but ARC Prize explicitly warns that saturating a bounded benchmark does not prove AGI. The contradictions continue elsewhere. Astra scores around 96% on GPQA Diamond, dramatically outperforming the PhD-expert baseline cited by the benchmark creators. Yet on Humanity’s Last Exam with tools, it reaches only 57.2%, while Claude Fable 5.1 scores 65% in OpenAI’s comparison. A model capable of producing formally verified new mathematics can still fail more than 40% of an extremely difficult expert-level exam. Then there’s Astra as a digital worker. OpenAI demonstrations show the model building complex projects across Codex, Blender and Unreal Engine 5, inspecting its own work and correcting errors. It has generated a routed PCB inside KiCad and can execute real terminal-based engineering and scientific workflows. But these demonstrations often involve human direction, corrections and scaffolding — very different from handing an autonomous employee a goal and coming back days later. That distinction becomes critical because OpenAI’s own Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” The autonomy evidence isn't there yet. A UK AISI evaluation put Astra’s 50%-reliability mathematics time horizon at roughly 30.9 minutes. Artificial Analysis found an approximately 51% hallucination rate at maximum effort on its adversarial AA-Omniscience benchmark. Astra also still fails substantial portions of difficult software and professional-work evaluations. And cybersecurity introduces another complication. Astra is the first broadly deployed OpenAI model to reach the company’s Critical cybersecurity capability threshold. On fresh vulnerabilities from June through August 2026, it scored 39% versus 5.5% for the previous model and discovered two previously unknown vulnerabilities during testing. So is GPT-6 Astra actually AGI? There’s now a serious case that models have crossed into something qualitatively different from the chatbots we were using only a few generations ago. Astra can reason across unfamiliar environments, operate professional tools, perform research-level mathematics and produce useful work across radically different domains. But there’s an equally serious case that we’re confusing peak intelligence with dependable intelligence. Astra’s performance can change dramatically depending on its scaffolding. It still hallucinates, still fails difficult bounded tasks, and there is no published representative evidence that it can reliably own even a full unsupervised working day across unfamiliar professional goals. Under a broad definition of general intelligence, the AGI era may have started. Under OpenAI’s own stricter definition? Not yet. CHAPTERS 00:00 GPT-6 Astra and the AGI Question 01:52 What OpenAI Actually Confirmed 03:16 The Benchmark That Breaks If You Change the Harness 05:01 Solving Math Problems That Were Still Open to Humanity 06:57 The Exam That Exposes the Gap 08:14 Astra as a Digital Worker, Not Just a Chatbot 10:42 Where It Still Falls Apart 13:32 So What Does AGI Even Mean Here 14:48 The Case For, and the Case Against 16:23 Is This Actually AGI #openai #gpt6 #gpt6astra #agi #artificialintelligence