"You need a 24 GB GPU for serious local LLMs in 2026."
Everyone repeats this. It's not true anymore.
In this video I go over new methods to run a 35B-parameter model on an RTX 4060 Ti 8 GB: • 41 tok/s at 16k context • 24 tok/s at 200k context. This approach uses a relatively common ryzen system with 64gb of vram and llama cpp.
Let me know what you think in the comments below!
X post - https://x.com/edgaraveloso/status/2050145746272321717
Rent an nVidia GPU on Vast AI (15% OFF) https://cloud.vast.ai/?ref_id=74601
For sponsorships or collaboration inquiries please contact aifluxcontact@proton.me
Games footage: https://x.com/ZoldenGames OR https://t.co/u24MwZRsd7
Chapters 🎬:
0:00 - intro
0:32 - qwen 3.6 in smaller spaces
1:24 - post breakdown
2:00 - initial benchmarks
2:20 - how does it work?
2:30 - MoE offload
3:15 - suggested GPUs
3:35 - best context for agents
3:55 - flash attention
4:40 - 3070 also good?
5:44 - reddit users try this
6:28 - gpu specs
7:00 - my conclusion
Choose to Build with AI
Matched to Qwen
AI Maker Residence 3
The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code.
Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.
◆ Fri 09 Oct 2026◆ KOKO Cafe, London◆ With Nick Sarafa