Which Qwen 27B Actually Runs On Your Graphics Card
Qwen 3.8 27B thinks too much. UkisAI's Swift fine-tune cuts reasoning tokens 58% with 1% accuracy loss, see if it's actually faster, cheaper and better for local AI. Qwen 3.8 27B, Alibaba's open-weight 27B reasoning model, defaults to generating 200M tokens of private working notes, nearly 2.5× the median for similar models, forcing you to wait through every one before your answer arrives. UkisAI, a small European AI lab, fine-tuned Swift-Qwen3.8-27B by identifying and penalizing tokens linked to overthinking, reporting 58% fewer thinking tokens on GPQA-Diamond, a 41% median cut across nine benchmarks, and speed-ups ranging from ~20% on math word problems to ~33% on coding tasks in third-party testing. The trade-offs: Swift holds accuracy within ~1% on everyday tasks and graduate-level science, gains ~5 points on code under token limits (where the original hits the ceiling first), but loses nearly 5 points on competition math, UkisAI admits a math-token training bug. Simon Willison's independent test found the original on default spent 21 minutes generating image code; Swift's single outside timing (GSM8K and HumanEval via Q4 quantized GGUF) showed 20, 35% shorter waits with mixed accuracy. Both twins require identical 18, 24 GB graphics cards for Q4_K_M quantized deployment; the original stays Apache 2.0, while Swift's free license caps commercial use at US$1,000,000 annual revenue. You can also turn the original's built-in reasoning_effort dial to medium (cutting tokens by ~50% but losing ~4 points) for free, Swift sits between xhigh and medium, aiming for top-setting answers with partial speed gains. The full benchmarks, external checks, quantized drift measurements, and license terms are covered. This breakdown is for builders choosing between the original Qwen 3.8 27B, Swift, and the free reasoning_effort settings on local hardware. Chapters: 0:00 Intro 1:18 Basics 3:44 Speed 6:57 Quality 9:58 Settings 13:06 Cost 15:17 Verdict Tools & resources mentioned: - Swift-Qwen3.8-27B: https://huggingface.co/ukisai/Swift-Qwen3.8-27B - Qwen 3.8 27B: https://huggingface.co/Qwen/Qwen3.8-27B - llama.cpp - Ollama - LM Studio - Artificial Analysis: https://artificialanalysis.ai - Hugging Face: https://huggingface.co About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #qwen #local-llm #ai-infrastructure