A Computer With No GPU Runs A Million-Token AI
Spark-X2.5-4B claims a million-token window on device. It does run with no graphics card, but reading a million tokens on a processor takes most of a day. Spark-X2.5-4B is a free, Apache-licensed 4B model from iFLYTEK's SparkLLM team; its makers say it and its 1.7B sibling are the only on-device models with a native window of a million tokens. Two independent testers ran it on plain processors with no graphics card, on short prompts: an 8-core Intel Xeon box with 16 GB of shared memory wrote about 6 tokens a second and read 30 to 36, and Javier Cañete's 18-core arm machine wrote about 47 and read up to about 254 (writing collapsed under 7 at 16 threads). It needs a runtime build from September 2026 onward (llama.cpp b10828, Ollama v0.34.1, LM Studio 2.34.0) and the compact 2.6 GB four-bit file; Ollama's latest tag is the full 8.2 GB file. Set the context size yourself: the default reaches for the full million. Memory: only 9 of the 36 layers keep notes (the KV cache) on the whole text, so by our arithmetic the cache costs about 36 KB per token: 1.2 GB at 32K tokens, 4.8 GB at 128K and just under 39 GB at a million, about 41 GB with the model, which a 64 GB desktop holds. A 32 GB machine fits it only with a compressed cache whose accuracy cost is unmeasured; a 16 GB machine tops out around a couple of hundred thousand tokens. Waiting: nobody has timed a long prompt on a processor. By our arithmetic from the measured speeds, reading cost grows faster than the length, so the fast arm machine needs about 3 minutes for 32K tokens, about 20 for 128K and about 13 hours for a million; the Xeon box would need about 4 days. Answers are only evidenced to about 100K tokens (the maker's own AA-LCR score, 56.3 against 57.0 for Qwen3.5-4B), and thinking mode can spend minutes of CPU time before an answer. Verdict: yes to the model, no to the million in everyday practice. The exception is a small graphics card, which read about 131K tokens in about 4 minutes where the fast processor needs about 20. On a CPU, use the 2.6 GB four-bit file on a current runtime, keep the window to tens of thousands of tokens and turn thinking off for quick jobs. The Stack did not run the model; measured figures belong to the named testers and makers, and the long-prompt times are our arithmetic. Sources and media credits: Claims and figures: SparkLLM launch post (DEV); XHToken/Spark-X2.5-4B model card, config and GGUF files (Hugging Face); Hugging Face discussions #3 (ljmodel), #21 (YOcean) and #26 (nakeep); Javier Cañete, jacar.es; peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF; litert-community/Spark-X2.5-4B; Artificial Analysis methodology; llama.cpp issue 28300 and pull request 27868 RAM install footage: Archetype Origins, Wikimedia Commons, CC BY 3.0 Graphics card install footage: Humaam, Wikimedia Commons, CC BY 3.0 Stock footage (Pexels): Kamil Zubrzycki, Tom Shearwood llama.cpp marks: github.com/ggml-org/llama.brand Chapters: 0:00 Intro 1:53 Running It 4:26 Memory 8:01 Reading Time 10:35 Answers 13:31 Verdict Tools & resources mentioned: - Spark-X2.5-4B model card (Hugging Face): https://huggingface.co/XHToken/Spark-X2.5-4B - Spark-X2.5-4B GGUF files: https://huggingface.co/XHToken/Spark-X2.5-4B-GGUF - llama.cpp: https://github.com/ggml-org/llama.cpp - Ollama (SparkLLM/Spark-X2.5-4B): https://ollama.com/SparkLLM/Spark-X2.5-4B - LM Studio: https://lmstudio.ai - CPU-only test by Javier Cañete (jacar.es): https://jacar.es/en/cpu-only-llms-arm64-maple-minicpm5-spark/ - CPU-only test (Hugging Face discussion #3): https://huggingface.co/XHToken/Spark-X2.5-4B/discussions/3 - Sharp-Spark-X2.5-4B-GGUF (cache and long-prompt measurements): https://huggingface.co/peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #local AI #LLM #CPU inference