Laguna S 2.1: The Best Local Agentic Coder?
In this video, we break down Poolside’s Laguna S2.1, an open-weights 118B MoE coding model (8B active per token) with a 1M-token context window, and why it performs above its size on agentic coding benchmarks like Terminal Bench 2.1. I cover how it was trained with reinforcement learning in FP8 on ~4,000 NVIDIA H200s in under nine weeks, plus how they tackled reward hacking on SWE-bench using an external LLM judge, prompt amendments, and network-blocked sandboxes. Thanks to @NVIDIADeveloper for DGX Spark. Laguna: https://poolside.ai/blog/introducing-laguna-s-2-1 DGX Spark: https://nvda.ws/3XIkwsh Try it out: https://chat.poolside.ai/ Pool Agent Harness: https://poolside.ai/get-started vLLM Serving: https://github.com/MiaAI-Lab/Laguna-S-2.1-DGX-Spark-RTX-6000-PRO DSpark video: https://youtu.be/eFgknPFK-g0 MoE Quantization paper: https://arxiv.org/pdf/2606.00206 My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: engineerprompt@gmail.com Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 00:00 Laguna S2.1 Overview 01:17 Benchmarks and Harness 02:13 Reinforcement Learning 03:34 Reward Hacking Fixes 05:16 Running on DGX Spark 06:55 NVFP4 Quantization 08:03 Speculative Decoding Speed 10:09 Pool Harness Demo 12:07 Verbose Reasoning Loops