GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?
Visit AISeeKing (my second channel about more ai stuff) : https://youtube.com/@aiseeking In this video, I’ll be breaking down the launch of GPT-6 Astra, OpenAI’s claims about entering the AGI era, and what the available benchmarks actually reveal. While Astra delivers major improvements in computer use, cybersecurity, terminal workflows, and long-horizon tasks, its coding performance, general intelligence scores, pricing, and benchmark comparisons make the overall story far more complicated. -- Key Takeaways: 🚀 GPT-6 Astra delivers major improvements in computer use, autonomous workflows, cybersecurity, and interactive reasoning. 🧩 Its 99.9% ARC-AGI-3 score was achieved with OpenAI’s Provider Adapter, while the provider-neutral result was 62.7%. 🧮 Astra performs exceptionally well on FrontierMath Tier 4, but solving unsolved mathematical problems remains rare and extremely expensive. 💻 Coding results are mixed, with Astra winning Terminal-Bench 4.0 but remaining close to existing models on DeepSWE and FrontierCode. 📊 Artificial Analysis gives Astra the same Intelligence Index score as GPT-5.6 Sol, despite Astra’s significantly higher API pricing. ⚡ Astra uses fewer tokens during long coding tasks, making it a potentially efficient autonomous coding agent despite its higher per-token cost. 🖥️ Computer use may be the real GPT-6 breakthrough, with substantial gains on OSWorld, ScreenSpot Pro, and AutomationBench. 🔐 Astra is OpenAI’s first model to reach its Critical cyber capability threshold, although its most advanced capabilities will be restricted. 🧠 OpenAI describes Astra as its most aligned model, but its system card also suggests that its internal reasoning may be harder to monitor. ⚖️ Overall, GPT-6 Astra appears to be a specialized agentic leap rather than the universal intelligence leap suggested by the AGI marketing.