Hy4-Preview Just Shrunk From 1.5TB to 200GB
Want to make money and save time with AI? Join here: https://www.skool.com/ai-profit-lab-7462/about Video notes + links to the tools π https://www.skool.com/ai-profit-lab-7462/about Get a FREE AI Course + Community + 1,000 AI Agents π https://www.skool.com/ai-seo-with-julian-goldie-1553/about Get a FREE AI SEO Strategy Session: https://go.juliangoldie.com/strategy-session?utm=julian Tencent shrunk its 770B-parameter HY4 Preview from 1.5TB to 213GB β but the number everyone's leaving out of the thumbnail changes the whole story. I break down exactly how the mixed-precision quantization works, what it actually scores on benchmarks, and the one line in the docs that stops most people from running it. If you care about open models and where local AI is really headed, this is the honest version. 00:00 Intro β 1.5TB to 200GB, and the missing number 00:40 The Real Math β Why HY4 Preview is 1.56TB 01:06 213GB β The compressed build 01:14 Quantization Explained β In plain English 01:39 The Smart Trick β Mixed precision, not brute force 02:00 The Photo Analogy β Why some layers get protected 02:09 Where the Cuts Go β Experts vs down projection 02:27 Benchmark Results β 85% smaller, ~1 point lost 03:38 The Catch β You still need 214GB of VRAM 03:57 The Better Option β Why Q4_K_M is the default 04:16 Storage vs VRAM β The mistake everyone makes 04:24 Real Speed β 20 tokens/sec on 8 data center GPUs 04:39 Who It's Actually For β Not your desktop 05:19 Inside the Model β MoE, 49B active per token 05:49 Sparse Attention & 1M Context β The efficiency stack 06:06 It Won't Just Run β Stock llama.cpp fails 06:41 The Bigger Pattern β 4 techniques, one direction 07:20 The Real Takeaway β Spend precision where it matters