Minimax M3.1 Flash (Fully Tested): Okay, this MODEL is PRETTY GOOD!
In this video, I’ll be testing MiniMax M3.1 Flash on KingBench 3 to see how it handles interactive simulations, 3D objects, SVG illustrations, games, math, local model training, and more. I’ll review all eight results, test their main interactions, inspect key issues, and compare the model with GPT-6 Sol, Grok 4.7, SWE 2, Opus 5.5, and the rest of the leaderboard. -- Key Takeaways: 🚀 MiniMax M3.1 Flash scores 53 out of 80, or 66.25 percent, on KingBench 3. 📈 This is a 35-percentage-point improvement over the earlier MiniMax M3 result. 🪑 The 3D folding table is the strongest visual result, earning nine out of ten. 👁️ The contact lens case looks polished and includes working interactive caps, but its compartments need improvement. 🏹 The archery game has attractive scenery, but misplaced targets, animation errors, and a faulty timer hurt the experience. 🛗 The elevator simulation’s main interaction fails because a code error prevents passengers from spawning. 🐼 The local Gemma 2 2B training project includes real data, LoRA adapters, saved weights, and a working interface, but its dataset has reliability issues. ⌚ The 3D watch tracks two time zones and has smooth hands, but its hour markers are incorrectly positioned. 🏆 M3.1 Flash shows major progress, although it still trails stronger models such as GPT-6 Sol, SWE 2, and Opus 5.5 overall. ⚠️ The results show why generated applications should be tested carefully before they are used in larger projects.