4B Parameter Local AI Model: Is It Worth It?
4B parameter local AI: NeoHorse-1 vs Qwen 3.5 benchmarks show agent harness routing beats raw scale, but only for single-move tasks. NeoHorse-1-4B is a post-trained variant of Alibaba's Qwen3.5-4B released in September 2026 using a routing harness curriculum to boost agentic performance. The model is free, Apache 2.0 licensed, ships as a 2.4 GB four-bit file via Ollama, and runs on a laptop with 262K native context. Across ten benchmarks, BFCL v4, VitaBench, tau2-Bench, PinchBench, WorkBuddy Bench, QwenClawBench, HumanEval, LiveCodeBench, IFBench, and IFEval, the macro average improved from 58.94 to 64.87 at 4B and 65.60 to 69.04 at 9B. On agentic harness tests it leads its weight class; HumanEval jumped from 87.20 to 96.95. But WorkBuddy Bench, which tests real repository tasks, only climbed from 24.62 to 34.41, revealing the model's core weakness: it stops after a single attempt and doesn't enter repair loops. A controlled data ablation showed routing-harness recordings alone (versus public synthetic data Toucan) lifted the five-benchmark average by 6.26 points, proving live trajectories are the active ingredient. However, those benchmark scores came from full-precision 8.41 GB checkpoints on SGLang v0.5.17 with reasoning enabled, not the compressed copy you download locally. The KV cache for full context costs 8 GB of memory, exceeding the model weights. The routing harness itself is proprietary; open-source copies lack that feedback loop. This video digs into what those numbers actually mean on real work. Best for builders evaluating whether small local models can handle agentic workflows, and anyone curious whether post-training beats scale for coding tasks. Chapters: 0:00 Free model, Apache licensed, runs offline 0:46 One command to launch, no containers 1:39 Context window's hidden 8GB cache bill 3:11 Routing harness predicts task complexity 5:17 Real data beat synthetic training data 6:37 The wins hide a deeper split 8:51 Single attempt, then it quits 10:40 Following instructions before writing code 11:58 The benchmarks assume full precision 13:11 Scaling crawls at diminishing returns 14:25 Where 4B actually works in production Tools & resources mentioned: - NeoHorse-1-4B on Hugging Face: https://huggingface.co/TokenRhythm/NeoHorse-1-4B - NeoHorse-1 on Ollama: https://ollama.com/TokenRhythm/neohorse-1 - arXiv paper 2609.08183: https://arxiv.org/abs/2609.08183 - NeoHorse GitHub: https://github.com/TokenRhythm/NeoHorse - Qwen3.5-4B base model: https://huggingface.co/Qwen/Qwen3.5-4B About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #localai #aiagents #qwen