The Free 975B AI Model You Can Own (Inkling AI)
Inkling AI is Mira Murati's free 975B open-weights model on Hugging Face, but running it yourself has a catch Inkling is the new 975-billion-parameter open-weights model from Mira Murati's Thinking Machines Lab, and it's the first credible American entry in an open-weights tier that's been dominated by Chinese labs. This breaks down what's actually inside the release: a mixture-of-experts architecture with 41 billion active parameters per token, a 1M-token context window, 45 trillion tokens of multimodal pretraining, and a controllable thinking-effort dial that trades compute for accuracy. Inkling debuted at 41 on the Artificial Analysis Intelligence Index, edging out Nemotron 3 Ultra, though Thinking Machines itself admits it's not the strongest model available, open or closed, and hands-on testers report it trailing GLM 5.2 and Kimi K2.6 on some benchmarks. The real story is what open weights actually buy you. Apache 2.0 licensing, day-zero support in vLLM, SGLang, llama.cpp, and Unsloth, and a companion fine-tuning platform called Tinker mean you can inspect, audit, and specialize the model instead of renting a closed API. But the hardware reality is brutal, hundreds of gigabytes of RAM even quantized, terabytes of VRAM for the full checkpoint, so most people end up renting it anyway through Together AI, OpenRouter, or Tinker, at output pricing roughly four times comparable open models. This is for builders and AI-curious viewers trying to figure out whether an open, ownable base model actually beats renting a frontier chatbot, and who Inkling is realistically built for. Chapters: 0:00 Her Free Trillion-Parameter Bombshell 0:42 What's Actually Inside The Beast 2:24 The Benchmark They Won't Oversell 4:15 Why An Empty Seat Made News 5:30 Can You Actually Run This Thing 6:53 The One Thing Renting Can't Give You 8:15 Free Until You Hit Run 9:25 So Who Really Owns This Model Tools & resources mentioned: - Thinking Machines Lab: https://thinkingmachines.ai - Inkling (Hugging Face weights): https://huggingface.co - Tinker (fine-tuning platform) - vLLM: https://github.com/vllm-project/vllm - SGLang: https://github.com/sgl-project/sglang - llama.cpp: https://github.com/ggerganov/llama.cpp - Unsloth: https://unsloth.ai - Artificial Analysis: https://artificialanalysis.ai - Together AI: https://together.ai - OpenRouter: https://openrouter.ai About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #inklingai #aiagents #mirmurati #opensourceai #aitools