No GPU, Still Runs A 27B AI Model

From the creator

Alibaba's RISC-V chip claims 30 tokens/second on Qwen 27B with no GPU, here's what that speed actually means, whether the hardware is real, and when you should ditch your graphics card. Alibaba's XuanTie team announced in August 2026 that their C950 RISC-V processor runs Qwen3.8-27B at over 30 tokens per second with no discrete GPU, achieving 1.9-second time-to-first-token. The headline is arresting, but the hardware is licensable IP, not a product you can buy, and no independent third party has reproduced the benchmark, every coverage source echoes Alibaba's own vendor test. The real story splits into two parts: the experimental RISC-V path (a 2027 hardware preview) and Qwen3.8-27B itself (deployable right now). The model is a dense 27-billion-parameter Apache 2.0 multimodal transformer already running on 24GB consumer GPUs via vLLM at 24.6 GiB quantized, or hosted serverless via Cloudflare Workers AI. On vision tasks, Roboflow's benchmarks rank it 32nd of 34 models at 61.2% average, weakest on OCR, which matters because text extraction is the most common production use case. For decode speed alone, 30 tokens/second is comfortably interactive; the real limit is memory bandwidth, not raw math throughput. If you want to run this 27B model locally this month, a standard consumer GPU or hosted endpoint both have hard, published numbers backing them. The no-GPU claim survives only as a data point about future silicon until independent reproduction on shipping hardware arrives. Built for builders and anyone evaluating local AI deployment tradeoffs between GPU, CPU, and hosted paths. Chapters: 0:00 The 30-token headline everyone's sharing 0:47 Why interactive speed matters for UX 2:32 The 24.6 GiB memory floor explained 4:13 Why the C950 isn't a product yet 6:11 When press releases echo without proof 7:58 The model you can actually use today 9:49 Where vision benchmarks crack 11:30 Five metrics GPU still wins on 13:59 Yes, but can you actually buy it? 15:40 What third-party proof would look like Tools & resources mentioned: - Qwen3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B - vLLM: https://docs.vllm.ai - Cloudflare Workers AI: https://developers.cloudflare.com/workers-ai/models/qwen3.8-27b - Roboflow Vision Evals: https://roboflow.com - Artificial Analysis Intelligence Index About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen #risc-v #localai

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.