One Computer Can Serve AI To Every Device You Own
Run Claude Code on local hardware: one computer serves AI to every device via LM Studio's dual-API server, llama.cpp, and MCP tool-calling, no cloud needed. LM Studio and llama.cpp both expose Anthropic-compatible `/v1/messages` endpoints on the same server that answers OpenAI-compatible routes, letting terminal coding agents like Claude Code and Codex point to your own machine instead of the cloud. Set two environment variables, `ANTHROPIC_BASE_URL=http://localhost:1234` and `ANTHROPIC_AUTH_TOKEN=lmstudio`, and Claude Code runs against local hardware; the same port answers Codex, Hermes Agent, and OpenClaw with no reconfiguration. Context window matters: Claude Code needs ~25k tokens, OpenClaw ~50k, Hermes ~64k, because the agent's system prompt consumes most of the budget before your first instruction. LM Studio's LM Link feature routes requests through a Tailscale-backed encrypted network, letting a thin laptop run coding tools while the model stays on a powerful desktop in another room, your laptop becomes a mailbox, the address stays localhost:1234, and the model loads on whichever device holds it. Bypass accounts entirely by binding to your local network (0.0.0.0), but LM Studio warns: unauthenticated servers expose your model to anyone scanning your IP. Run the headless daemon `llmster` on a Linux box without a screen; models load just-in-time and evict after 60 minutes of idle time, so multiple editors (Zed, Cline, Continue.dev) can share one rig without RAM choking. Model Context Protocol (MCP), an open standard now governed by the Linux Foundation with Anthropic, OpenAI, Google, Microsoft, and Block as co-sponsors, lets your local model call tools and APIs, but OAuth credentials and file-system access live on your machine. llama.cpp's server implements both APIs via pull request #17570, gated with `--api-key` for security. For builders running local agents, code editors, and private inference without third-party vendor lock-in. Chapters: 0:00 Why agents only need an address 1:47 One port, two API languages 3:02 Four agents, zero reconfiguration 4:13 Context hunger and system prompts 5:04 The laptop as a glorified mailbox 6:12 Encrypted networks you don't control 7:36 WiFi door with no lock 8:48 Servers that don't need screens 9:44 Models summoned, not stored 10:54 Your local model calls out 12:03 The one authentication you must enable 13:04 AI in your pocket, with limits 14:05 llama.cpp's parallel path 14:58 Local changes everything Tools & resources mentioned: - LM Studio: https://lmstudio.ai - llama.cpp: https://github.com/ggml-org/llama.cpp - Ollama: https://ollama.ai - Open WebUI: https://openwebui.com - Tailscale: https://tailscale.com - Model Context Protocol (MCP): https://modelcontextprotocol.io - Claude Code: https://www.anthropic.com - Codex: https://openai.com/research/codex - Hermes Agent: https://nous.gitbook.io - OpenClaw: https://www.openclaw.ai - Locally (LM Studio Mobile): https://lmstudio.ai/locally - vLLM: https://github.com/vllm-project/vllm - Zed Editor: https://zed.dev - Cline: https://github.com/cline/cline - Continue.dev: https://continue.dev About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #lmstudio #claudecode #localai