Same Mac, Twice The Local AI Speed With One Trick

From the creator

Llama.cpp on Mac: does the Neural Engine really double local AI speed? One report claims 24.3 tokens/sec on M3, here's what that actually measures. The Apple Neural Engine is a fixed-function accelerator on every M3 Mac, but reaching it for local LLM inference isn't automatic. A researcher reports that optimizing data flow through the M3 Neural Engine lifted Llama 3.2 1B decode speed from 10.0 to 24.3 tokens per second, no new hardware required. That's compelling, but it's a single unreplicated benchmark measured only at the decoding stage, not the full response pipeline. The catch: mainstream local-AI stacks like llama.cpp (which uses Metal GPU backends) and MLX don't target the Neural Engine by default; they dispatch to CPU or GPU instead. Core ML offers a supported route through Apple's framework, and specialized projects like ANEMLL exist for ANE-focused workflows, but the direct hardware path that produced the reported speedup is undocumented, unsupported, and version-fragile, intended for research, not shipping software. Moving model weights across memory can become the bottleneck even when the accelerator has arithmetic headroom left, which explains why data routing matters. This video walks through what the M3 result actually measured, why software choice determines hardware access, how to run a fair test on your own Mac, and why unreplicated benchmarks deserve skeptical watching. Built for developers and builders running local models on Apple silicon who want to separate hype from testable claims. Chapters: 0:00 Intro 1:02 Basics 2:26 Speed 3:36 Bottleneck 4:50 Software 6:17 Your Mac 7:45 Conclusion Tools & resources mentioned: - llama.cpp: https://github.com/ggerganov/llama.cpp - MLX: https://ml-explore.github.io/mlx/ - Core ML: https://developer.apple.com/coreml/ - ANEMLL: https://github.com/rmalejczuk/anemll - arXiv 2606.22283: Apple Neural Engine: https://arxiv.org/html/2606.22283v1 About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #llamacpp #local ai #neural engine

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.