This 35MB Model Replaces Your Tool Calling LLM
Needle 3: A 35MB parameter model that runs function calling on a CPU in 66 milliseconds, with no GPU and no JSON parser. In this video I walk through Cactus Needle 3, run a live home automation demo, break down the simple attention network architecture, and show you a Colab notebook you can run yourself. I also cover four gotchas I hit while testing it, because knowing where a model breaks matters more than knowing where it shines. Notebook link below. Let me know in the comments what you'd use this for. Blog: https://cactuscompute.com/needle Notebook: https://tinyurl.com/yewmt3za My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 00:00 - Needle 3 & Automation Models 00:57 - Live Home Automation Demo (66ms Response) 02:37 - How Needle 3 Compares to Traditional LLMs 03:42 - Architecture: Simple Attention Network & Slicing 06:40 - Running Function Calls on Real Queries 07:54 - Gotchas: State Reset & History Leakage 09:08 - Importance of Default Function Values 10:01 - Handling Messy Input vs. Structured Text 10:49 - Classification Limits vs. System 1 Models 11:18 - Using Needle 3 for Embeddings & Semantic Search 12:03 - Model Slicing & On-Device Automation Summary