Free Local AI is WILDLY Good Now (Replace Claude Code & Codex)

From the creator

Claude AI leads SWE-bench Verified by ~14 points, but free local models like Qwen3-Coder 30B now close the gap fast Claude AI's closed models still top SWE-bench Verified, but the best open-weight coding models have closed the gap to roughly 13-14 points, and this video breaks down exactly what that distance costs you in practice. The leaderboard fight covers Ornith-1.0-397B, DeepSeek V4 Pro, MiniMax M3, and Kimi K2.6 all clustered near 80%, plus Claude vs ChatGPT-style comparisons showing how fast open models are catching closed frontier releases from just a year ago. We walk through why two different leaderboards (BenchLM and Steel) disagree on rankings, why SWE-bench Pro was built after a contamination audit caught models reading test answers in advance, and how GLM 5.2 leads that harder open-weight field at 62.1%. You'll see why Anthropic's own engineering writeup admits SWE-bench scores measure a model plus its agent scaffold, not the raw model, which is the real reason Claude Code and Codex feel harder to replace than a benchmark number suggests. We cover the desk-class hardware ladder that actually matters: Qwen3-Coder 30B running at 220 tokens/sec on a single 24GB GPU, Qwen 2.5 Coder 32B and its 7B little sibling, and how to wire local models into a real coding loop using Ollama, llama.cpp, and editor extensions like Continue. This is for builders deciding whether to keep paying for Claude Code or Codex, or route bounded coding work through a free, offline model instead, and it ends with the actual cost/privacy tradeoff instead of a single leaderboard score. Chapters: 0:00 How Close Is Free To Frontier 0:35 The Free Models Beating Last Year's Best 2:12 When 49% Was The Record 3:23 Two Leaderboards Can't Agree 5:22 Claude Code's Hidden Second Half 7:03 Building An Eighty-Percent Rig At Home 9:11 What The Meter And The Plug Cost 10:29 Where The Real Gap Still Lives Tools & resources mentioned: - Claude Code: https://claude.com/claude-code - Codex - Ollama: https://ollama.com - llama.cpp: https://github.com/ggerganov/llama.cpp - Continue: https://continue.dev - Qwen2.5-Coder-32B-Instruct: https://github.com/QwenLM/Qwen2.5-Coder - DeepSeek-Coder-V2: https://github.com/deepseek-ai/DeepSeek-Coder-V2 - SWE-bench Verified: https://www.swebench.com - SWE-bench Pro: https://www.morphllm.com/swe-bench-pro - BenchLM Leaderboard: https://benchlm.ai/benchmarks/sweVerified About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #claudeai #claudecode #localllm #aicoding #opensourceai

Choose to Build with AI
Matched to AI Coding Benchmarking

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.