Can A 29MB Model Replace Your Cloud LLM?
Needle 3's 29MB model beats DeepSeek V4 Flash, but only after training on DroidCall, on one narrow phone-command task. Here's what the benchmark really shows. Cactus Compute released Needle 3, a 29MB on-device AI model built for tool calling and function calling on phones, watches, smart homes, and embedded systems, not general chat. The headline claim is real but scoped: after fine-tuning on the DroidCall dataset, Needle 3's full model scored 70% on DroidCall's 200-question test in Cactus's own scoring, against 60.5% for DeepSeek V4 Flash, and even the 8MB four-layer version reached 62.5%: four more right answers than DeepSeek. Straight out of the box, however, DeepSeek won all six benchmarks Cactus ran. Inside, about 70 million of its 121 million parameters are lookup tables, so Cactus says it does the arithmetic of a 50-million-parameter model; most of its numbers are stored in about 2 bits instead of 16; and it is built as a "ladder": every depth from 2 to 20 layers is a working model on its own, and the first four layers alone are 8MB. A byte-level grammar built from the app's own action definitions constrains every token, so the app never receives a half-finished command, and Cactus says its engine is under 1MB. But the trade-off is real: Cactus's CTO says it can struggle with implied references and multistep reasoning, and it has no world knowledge to fall back on; one developer who fine-tuned it for a game-database tool got the right tool-call shape 32% of the time, against 91% for their fine-tuned Google FunctionGemma. The model and code ship free under Apache 2.0; hosted fine-tuning costs $19 for three training runs. This video breaks down the actual benchmark numbers, how a neural network fits into 29MB, and whether Needle makes sense for your project, ideal for builders shipping on-device AI, and anyone curious how benchmark claims actually hold up. Chapters: 0:00 Intro 1:43 The job 3:42 The claim 6:40 Size 8:50 Worth it? 11:19 Conclusion Tools & resources mentioned: - Needle 3: https://huggingface.co/Cactus-Compute/needle3 - Cactus Compute GitHub: https://github.com/cactus-compute/needle - Cactus's Needle 3 release page: https://cactuscompute.com/needle - DroidCall paper (arXiv): https://arxiv.org/abs/2412.00402 - DeepSeek V4 Flash model card: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash - Show HN launch thread (Cactus CTO answers): https://news.ycombinator.com/item?id=49748553 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #agentic AI #on-device AI #tool calling