Mistral Large 4 (Le Chonk - Fully Tested): Is IT ACTUALLY The BEST OPEN MODEL!?
In this video, I'll put Mistral Large 4, also known as Le Chonk, through my eight KingBench 3 prompts in OpenCode through OpenRouter, giving it a fresh folder for every task and judging only what it actually delivered. I'll go through the elevator simulation, the 3D contact lens case, the Three.js folding table, the panda SVG and the bow-and-arrow game, opening every result in the browser. Then I'll cover the counting problem, the local Gemma 2B fine-tuning job with its web UI, and the 3D wristwatch that only worked after a repair. Finally, I'll go over the permission-corrected score, why the retries don't count, and where I'd actually use this model. -- Key Takeaways: 🧪 I ran the same eight KingBench 3 prompts in OpenCode through OpenRouter, one fresh folder per task, and scored the delivered result. 🛗 The elevator simulation works properly, with all five spawned riders delivered and no browser errors, for 9.5 out of 10. ⚠️ The default contact lens and counting sessions ended without a file or an answer, so both score zero. 🔁 A separate no-reasoning run did build a working contact lens case, but I kept that retry outside the total. 🪑 The Three.js folding table folds and unfolds smoothly with the slider and earns 8.5 out of 10. 🐼 The panda SVG is clean and cute but the bite is ambiguous, and the archery game works but its targets are tiny, so 8 and 6. 🧠 With permissions fixed, it built a 510-example dataset, trained a LoRA adapter for 400 steps and served a new panda fact on every refresh for 10 out of 10. ⌚ The wristwatch page loaded blank because of a broken import map; one repair run fixed it, but that 6 out of 10 isn't counted. 📊 The permission-corrected score is 42 out of 80, or 52.5%, which is 5.25 out of 10 overall. ✅ My verdict: try Le Chonk for contained visual prototypes with a browser and test loop nearby, but I want it to finish more reliably before it becomes my default coding model.