Can A GTX 1080 Ti Run A 35B AI Model?
Can a $150 used GTX 1080 Ti run a 35B AI model? Learn how quantization and mixture-of-experts architecture make local AI feasible on old GPUs. A used GTX 1080 Ti holds 11GB of VRAM but a standard 35B parameter model like Qwen 3.5 needs 21.4GB at Q4_K_M quantization, nearly double. This video tests whether you can actually run that model on the cheap card, and at what speed. The key is that models like Qwen 3.5-35B-A3B and Qwen 3.6-35B-A3B are mixture-of-experts: only 3 billion of the 35 billion parameters activate per token, so you can keep the active experts on the GPU while the idle specialists sit in system RAM. Real hardware testing shows a GTX 1080 Ti achieves 17, 24 tokens/second on a 35B MoE model using llama.cpp with careful layer offloading, three times faster than you read. The video walks through what a graphics card needs to run any LLM, how expert offloading works, measured speed on real hardware, what happens when you shrink the model to 2-bit quantization to fit it entirely on card (trading quality for ~65 tok/s), whether the setup handles long coding prompts, and the honest cost-per-performance verdict against a modern RTX 5060 Ti (16GB, ~$750). You'll also see PXA, a new project built specifically for Pascal-era cards. This is for builders deciding whether a bargain used GPU is worth the upgrade, or whether your existing CPU and RAM already does the job. Chapters: 0:00 Intro 1:11 Basics 3:29 Experts 5:55 Speed 8:20 Shrinking 11:08 Coding 13:13 Worth it? 15:59 Conclusion Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - Unsloth - Ollama - ik llama.cpp - PXA: https://github.com - willitrunai.com: https://willitrunai.com - InventiveHQ Lab - Phoronix About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local ai #quantization #llm #gpu #mixture-of-experts