Grok 4.7 (Fully Tested): This is ACTUALLY SAD! (WORST THAN GLM-5.3 FLASH!)
https://bambooed.ai is in Alpha. Use Coupon Code: 1000KINGS & get 50% OFF for the first 12 months (applicable for first 1000 users only). In this video, I’ll be reviewing Grok 4.7, SpaceXAI’s latest model for coding, agents, and knowledge work. I’ll compare it with Grok 4.6, Grok 4.5, Astra, and other leading models while examining its benchmark scores, token usage, task costs, frontend quality, 3D capabilities, and instruction-following performance. -- Key Takeaways: 🚀 Grok 4.7 delivers meaningful improvements on some difficult engineering and agentic tasks. 📊 Official benchmarks show gains on CursorBench 4.0 and Terminal-Bench 4.0, but several individual benchmark scores decline. 💸 Despite having the same token pricing as Grok 4.6, Grok 4.7 can cost significantly more per completed coding task. ⏳ Average agent runtime increases considerably, resulting in longer waits for finished work. 🎨 Frontend quality and visual judgment remain disappointing and inconsistent for a flagship model. 🧊 Grok 4.7 performs well on some 3D tasks, but weaker results show that it still requires close supervision. 🧠 Instruction-following issues and repetitive loops make the model less reliable as a default coding agent. 🛠️ Grok can still be useful as a specialized subagent for clearly defined engineering assignments. 🏆 Grok 4.7 scores 62 out of 80 on KingBench 3, placing it behind SWE 2 and Astra. 👍 Overall, Grok 4.7 is more capable in certain areas, but its higher task costs, slower completion times, and uneven output make it difficult to recommend.