GPT-5.2 Low Found What Opus 4.5 Missed
GPT-5.2 Low just beat Opus 4.5 on a real code review. I had Opus write a 52,000-token audit prompt for my multiplayer game YeeBall, then ran the exact same prompt through every major model: GPT-5.2, Opus 4.5, Gemini 3 Flash, and Codex Max. The model with the lowest reasoning effort caught sprite rendering issues that Opus missed. Meanwhile, Gemini Flash went completely rogue and started editing files when I only asked for a review. For procedural tasks like audits and checklists, low reasoning actually beats overthinking. LEARN TO SHIP Idea to Deployed Cohort (waitlist): https://startmy.ai Ship a real app in 3 weeks with auth, payments, and database. Next batch kicks off in February; 20 spots. Key moments 00:00 The cheap AI model won 01:35 Testing Gemini 3 Flash for code review 02:36 Gemini's internal dialogue is exposed 04:51 Why low-reasoning models can work better 06:52 Comparing 5 different AI models live 08:25 The surprising winner: GPT-5.2 low 10:36 Nuance: Where smart models still shine 11:52 The final verdict on all tested models Join our Discord: https://rfer.me/discord My LLM Rules and Prompts Repo: https://github.com/rayfernando1337/llm-cursor-rules GET THE TOOLS Factory AI (Droid): https://factory.ai Cursor IDE: https://cursor.com CONNECT WITH RAY X (Twitter): https://x.com/RayFernando1337 #GPT5 #Opus45 #AICodeReview #Cursor #FactoryAI #Droid #GeminiFlash #CodexMax #AIComparison #Programming