Can A 78B Local AI Run On A Regular Gaming PC?
Run a 78B local AI on your gaming PC: Kolibri-1, mixture of experts, llama.cpp, and why it does not replace Claude Aleph Alpha released Kolibri-1 on October 3, 2026: a free 78-billion-parameter mixture-of-experts model with about 3.46 billion parameters active per token. Its official minimum is datacenter hardware (two A100 or H100 cards, or one H200, B200 or B300) for a roughly 78 GB footprint. Yet Hob-forge's unofficial 4-bit file (47.5 GB) ran on an AMD Ryzen 7 7800X3D with 128 GB of RAM and no graphics card at about 14 tokens per second, using 46.6 GB of RAM, and its smoke checks (a German explanation, a math word problem, a tool call) came back right. The 78 billion total parameters set the memory bill; the 3.46 billion active ones make processor speed possible. On paper, 64 GB of RAM should hold the 4-bit file, but only 128 GB has actually been tested. Split across an RTX 3060 (12 GB) and an Intel Arc Pro B60 (24 GB), Eliasfpv28's smaller 3-bit file reached about 50 tokens per second on a short test. You have to patch and compile llama.cpp yourself: Hob-forge's page says Ollama, LM Studio and other llama.cpp-based apps will not load the file until they support the architecture, and llama.cpp's own feature request (#29922) is still open. In Aleph Alpha's own table (high reasoning effort), Kolibri-1 leads the 3B-active class (75.5 English, 70.8 German overall), but the dense Qwen3.8 27B scores 80.2 and 79.9, and its 4-bit file, about 16.5 GB, fits one 24 GB graphics card. The table never compares Kolibri with Claude or any paid model, and we found no independent repeat of the scores yet. Verdict: yes, a gaming PC can run it if it has the memory (64 GB of RAM on paper, 128 GB tested) and you are willing to compile; in its maker's table it is the top scorer among models a processor can keep up with. No, it does not replace Claude. With a big graphics card, run a model that fits in it and wait for llama.cpp, then the one-click apps, to add support. The Stack did not run the model; every result shown is its owner's report. Sources and media credits: Aleph Alpha (Kolibri-1 model card and evaluation table); Hob-forge, Eliasfpv28, webmp3 and unsloth (GGUF files on Hugging Face); llama.cpp feature request #29922; Amine Raji (aminrj.com) RAM and graphics card install footage: Archetype Origins, CC BY 3.0, via Wikimedia Commons Datacenter footage: Lawrence Systems, CC BY 3.0, via Wikimedia Commons llama.cpp marks: github.com/ggml-org/llama.brand Stock footage: Pepino8A (Pixabay), Mario Aranda (Pexels) Chapters: 0:00 Intro 2:07 The Test 4:26 Size vs Speed 6:09 Your PC 10:19 Software 12:24 The Scores 16:10 Verdict Tools & resources mentioned: - Kolibri-1 (Aleph Alpha): https://huggingface.co/Aleph-Alpha/Kolibri-1 - Hob-forge Kolibri-1-GGUF (4-bit, the CPU test): https://huggingface.co/Hob-forge/Kolibri-1-GGUF - Eliasfpv28 Kolibri-1-Q3_K_S-GGUF (3-bit, two cards): https://huggingface.co/Eliasfpv28/Kolibri-1-Q3_K_S-GGUF - webmp3 Sakura-MicroQuality-Kolibri-1-GGUF (2-bit): https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-GGUF - llama.cpp feature request #29922: https://github.com/ggml-org/llama.cpp/issues/29922 - llama.cpp: https://github.com/ggml-org/llama.cpp - Qwen3.8 27B GGUF (unsloth): https://huggingface.co/unsloth/Qwen3.8-27B-GGUF - Amine Raji's RTX 3090 test (Qwen3.6): https://aminrj.com/posts/llamacpp-qwen36-35b/ - Ollama - LM Studio About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #kolibri1 #mixture of experts #local AI #llama.cpp #gaming PC