Deepseek V4.1 Flash (Fully Tested): 200 TPS & Beats Astra!? (+New Architecture Overview)
checkout my second channel AISeeKing for more cool ai stuff: https://www.youtube.com/@aiseeking In this video, I’ll be testing the newly launched DeepSeek V4.1 Flash with thinking enabled at maximum effort. It achieved 81.25 percent on my KingBench 3 coding tests while generating an estimated 221 tokens per second, including reasoning. I’ll cover its new architecture, benchmark results, API pricing, coding performance, and major improvements over my original thinking-disabled test. -- Key Takeaways: 🚀 DeepSeek V4.1 Flash introduces a new architecture, native visual understanding, and lower API pricing. 🧠 Enabling thinking at maximum effort increased my KingBench 3 score from 53.75 to 81.25 percent. ⚡ OpenCode records suggest an estimated generation speed of around 221 tokens per second, including reasoning. 🏗️ The model uses a 552-billion-parameter mixture-of-experts backbone alongside 196 billion Engram memory parameters. 💻 Major improvements appeared in the elevator simulation, archery game, local fine-tuning workflow, and wristwatch tasks. 🧮 The model correctly solved the restricted-moves math problem and verified its answer of 20,460. 🐼 It completed a 700-iteration LoRA fine-tuning run and created a working local Gemma 2B panda-fact application. ⌚ The wristwatch generation earned ten out of ten with accurate time-zone, daylight-saving, and half-hour-offset handling. 💸 DeepSeek V4.1 Flash starts at fifteen cents per million uncached input tokens and sixty cents per million output tokens during off-peak hours. 👍 Overall, DeepSeek V4.1 Flash is a fast, affordable, and highly competitive model for coding when thinking is configured correctly.