Gemini 4 Argon Beats Fable 5.1 & GPT 6 Astra
Google says Gemini 4 Argon beats GPT-6 Astra and Fable 5.1 at coding. One row down its own table, it comes last. Here is what the benchmarks, the price and the access actually say. On Google's own DeepSWE v1.1 numbers, Gemini 4 Argon scores 77.9% against GPT-6 Astra's 74.1% and Fable 5.1's 67.4%. On the very next row of the same table, FrontierSWE v2, it comes last of the three (55.0% against 65.5% and 56.3%), and it also trails both on Terminal-bench 4.0. Google's three DeepSWE figures come from three different places: its own run with a mini-swe agent harness, a public leaderboard and a system card. Independent evaluators back Argon as a front-runner, not a sweeper. Artificial Analysis scores it 53 on its Intelligence Index, level with GPT-6 Astra (and Fable 5.1). Vals ranks it first on the Vals Index at 68.90%, ahead of Fable at 65.83% and Astra at 63.13%, yet Fable wins Vals' Tax Agent Bench, 77.64% to 76.23%. Google's long-job showcase: its agents profiled, edited and measured again until the libgav1 video decoder ran 2.7 times faster than an earlier Rust port with identical video output, still short of hand-tuned C++. Separately, Argon's output limit rises from 64K to 1M tokens; Artificial Analysis tested it by pausing and resuming long answers across calls, and Vals ran it with a 262K limit. Cost: at the introductory $2 / $10 per million input / output tokens, Artificial Analysis measured $1.99 per index task for Argon against $3.26 for Astra; the saving comes from the lower price, not fewer tokens, since Argon writes about 62K output tokens per task against Astra's 27K. At the standard $4 / $20, it projects $3.98, more than Astra. No end date for the introductory price is confirmed. Access: Argon is rolling out to trusted cyber defenders through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers next and no date. Verdict: keep Fable or Astra for day-to-day work; if you have partner access and run long multi-step coding jobs, test Argon on those jobs; everyone else should shortlist it for a matched trial on their own code and bills. Media credits: Rust logo by the Rust Foundation, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) Stock footage (Pexels): Julia Repnikova, Fabio Reis de Abreu Chapters: 0:00 Intro 0:53 Coding 2:29 Independent tests 4:01 Long jobs 6:48 Cost 8:18 Access 10:06 Conclusion Tools & resources mentioned: - Google: Gemini 4 Argon announcement (Sep 30, 2026): https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ - Google: Gemini 4 Argon model evaluation (PDF): https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf - Artificial Analysis: Gemini 4 Argon evaluation: https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs - Vals AI: Gemini 4 Argon results: https://www.vals.ai/models/google_gemini-4-argon - libgav1 (Google's open-source video decoder): https://chromium.googlesource.com/codecs/libgav1 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #gemini4argon #aibenchmarks #gpt6astra