xAI's Grok 5: Elon Musk's 1 MILLION GPU Ambition!

From the creator

https://bitbiased.ai/ai-automation-services Elon Musk said Grok 5 would launch in Q1 2026 with 6 trillion parameters and a 10% chance of achieving AGI. But it's now October, Grok 5 still hasn't launched, and xAI has never officially confirmed those numbers. Meanwhile, Grok 4.7 just delivered an enormous benchmark jump: Terminal-Bench performance surged from 20.3% to 37.6%, an improvement of roughly 85% over Grok 4.6. And that might be our clearest indication yet of what xAI is building toward with Grok 5. But there's a catch hiding inside those benchmark results. xAI has released three major model updates in roughly ten weeks: Grok 4.5 in July, Grok 4.6 in August, and Grok 4.7 on September 21. The latest release introduces a new, larger base model, extended reinforcement learning, and training focused on complex tasks that can take hours to complete. That's a significant shift from simply improving an existing model. And it raises an important question: Is Grok 4.7 an early preview of the technology behind Grok 5? We examine what xAI has actually confirmed about Grok 5, what Elon Musk has claimed, and why the distinction matters. The rumored 6-trillion-parameter architecture, Q1 release window, and AGI predictions remain unverified. The only official confirmation is that Grok 5 was in training. Then we break down Grok 4.7's performance across coding, engineering, medical reasoning, and autonomous agent benchmarks. According to xAI's published results: - Terminal-Bench: 20.3% → 37.6% - CursorBench: 40.4% → 46.3% - DeepSWE: 65.2% → 71.0% - EEBench: 53.0% → 64.0% - HealthBench Pro: 48.5% → 56.7% Those gains look impressive, especially for long-running coding agents. But the tests were conducted by xAI using high reasoning effort, with limited public details about the evaluation methodology and no independent reproduction of the reported results. And despite the improvements, Grok 4.7 still trails competitors in several important areas. Claude Fable 5.1 reportedly scores 57.9% on Terminal-Bench versus Grok's 37.6%. GPT-5.6 Sol remains ahead on medical reasoning, while Grok takes a notable lead in engineering reasoning with its 64.0% EEBench score. The competitive picture is much more complicated than a single benchmark chart suggests. Pricing adds another dimension. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, compared with $4 and $20 for GPT-5.6 Sol on xAI's comparison chart. With a 500,000-token context window, image input, function calling, web search, X search, code execution, and integration into Cursor and Grok Build, xAI is positioning Grok as a serious option for developers building AI agents. But can those agents actually complete complicated work reliably? We investigate xAI's claims about Grok Bot, including a reported internal workflow involving hundreds of agentic steps and a five-person team merging more than 100 pull requests daily. We also explain why impressive internal demonstrations aren't the same as independently verified real-world reliability. Then there's the infrastructure behind everything. xAI has officially reported access to more than one million H100 GPU equivalents. Its massive Colossus computing infrastructure provides important context for the company's rapid release cycle, although the specific hardware training Grok 5 remains unconfirmed. Ultimately, Grok 5 faces a much harder challenge than delivering another impressive announcement. It needs to close the gap with Claude and GPT-5.6 on demanding agentic benchmarks, improve reliability on long-running tasks, and potentially move beyond the text-only output limitations of Grok 4.7. So is xAI approaching a genuine breakthrough, or is the gap between Musk's promises and the actual product getting harder to ignore? Watch the full breakdown to understand what Grok 4.7 reveals, what the benchmarks leave unanswered, and what Grok 5 would actually need to deliver. CHAPTERS 00:00 Grok 5 Is Late — But Grok 4.7 Changes the Picture 01:44 What xAI Has Actually Said About Grok 5 03:19 Three Releases in Two Months 04:39 What the New Training Actually Changed 05:51 The Benchmark Jumps, and the Catch Hiding in Them 08:02 Where Grok 4.7 Still Loses 09:54 What It Costs and What You Can Build 10:59 Does It Work Outside the Charts? 12:10 The Machine Behind It 13:42 What Grok 5 Has to Beat #Grok5 #Grok47 #xAI #ElonMusk #ArtificialIntelligence

Stop watching · start building
Matched to Grok 5

Makers residency 4th edition

Your idea can be anything. At our last workshops, people built plant-based nutrition businesses, a way to measure planetary resilience, crochet headwear and a way to connect lonely people living in the same building. An asset manager went on to raise an additional £25M on their fund within nine months. The fourth workshop is a full day at KOKO with Nick Sarafa, learning to build with Claude Code. Bring the thing you keep meaning to start. We’ll bring the pizza.

◆ Fri 13 Nov 2026 ◆ KOKO, London
Makers residency 4th edition
Live event
Makers residency 4th edition
Fri 13 Nov 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.