Qwen 27B Can Run Easily... Across Multiple Devices

From the creator

SwarmLLM splits Qwen's 27B model across browser tabs using WebGPU and WebRTC, 10KB per token, no server, read the privacy cost. SwarmLLM is an open-source MIT-licensed project that runs Qwen 3.8 27B, Alibaba's 27.78-billion-parameter dense multimodal model, split across multiple web browsers and devices using pipeline model parallelism. Instead of requiring one machine with tens of gigabytes of VRAM, SwarmLLM divides the model's 64 transformer layers into contiguous ranges, assigning each slice to a different peer via WebGPU for local compute and WebRTC for peer-to-peer activation exchange. The headline demo splits the model across a MacBook and iPhone, with only a 10-kilobyte hidden-state vector crossing the network per generated word, the 15GB of weights never move. You learn how the architecture avoids downloading the full model (browsers fetch only their assigned layer ranges), why reading weights dominates latency at 82ms per token (near memory bandwidth ceiling), how Qwen 3.8's built-in draft block enables speculative decoding to reach 16 tokens/second on a single device, and why the WebRTC congestion-control overhead added 3x latency until the project optimized packet slicing. The trade-off is clear: the project's own security docs state that room participants can reconstruct 88.4% of prompts via prompt-inversion attacks on intermediate activations, splitting a model across strangers' devices means handing them a readable version of your question. Multi-device rooms add capacity but reduce speed (16 devices ran slower than 3 due to per-hop latency), and the entire runtime is still experimental, Chrome on macOS is the only tested host. This video is for builders exploring distributed LLM inference, developers curious about WebGPU and peer-to-peer ML, and anyone asking whether pooling consumer hardware can replace a single server. Chapters: 0:00 When 62 layers meet 2 on Wi-Fi 1:17 Inside the 64-block relay 2:09 The 10KB trick explained 2:55 Browsers fetch only what they run 3:52 Where 82ms goes every token 4:54 Guessing ahead to break the wall 6:02 The digital doorman's job 7:12 Web wins the native benchmark 8:17 The packaging that cost 3x speed 9:31 Corruption hiding in rounding 10:42 Your prompt, 88% recoverable 11:54 Why 16 devices got slower 13:18 The speed myth that's backward 14:32 The only tested setup right now 15:23 What the project refuses to promise 16:26 What SwarmLLM says it will not do Tools & resources mentioned: - SwarmLLM: https://github.com/Nehanth/swarmllm - Qwen 3.8 27B: https://huggingface.co/Qwen/Qwen3.8-27B - Qwen on NVIDIA NGC: https://catalog.ngc.nvidia.com/orgs/nim/qwen/models/qwen3.8-27b/ About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #swarmllm #qwen27b #webgpu

Choose to Build with AI
Matched to AI Coding

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.