Engineers, DELETE the BASH Tool: Agentic Security For Pi Agent and Claude Code
95% of engineers are ONE BAD PROMPT away from their agents NUKING production. The Bash tool is a ticking time bomb sitting inside every single agent harness you run, and the math is brutal: RISK COMPOUNDS WITH RUNTIME. ⭐️ VIDEO REFERENCES - Damage From Within Codebase: https://github.com/disler/bash-damage-from-within - Damage Control Video: https://youtu.be/VqDs46A8pqE - Mythos Level Model Video (Capability): https://youtu.be/RvowJ_hmLps - Threads of Work Blog: https://agenticengineer.com/thinking-in-threads - Pi Agent Harness Video: https://youtu.be/f8cfH5XX-XU - Pi Coding Agent: https://pi.dev/ - Master Agentic Coding: https://agenticengineer.com/tactical-agentic-coding?y=yBcmIoA-vGs This video lays out the FIVE LEVELS OF BASH SECURITY for agentic coding, the framework every AI engineer needs before scaling agents to the moon. We run the exact same destructive prompts side-by-side against Claude Code with Opus 4.7 and the Pi coding agent with GPT 5.5, and watch the levels expose themselves in real time. Here's the framework in plain terms: Level 1: User prompt / skill - lazy, jailbreakable, non-deterministic. You're praying to the model gods. Level 2: System prompt - the law for your agent... but laws get broken at long runtime. Level 3: Bash tool + blacklist - the default I run globally via damage control hooks. Good start, but you'll NEVER cover every CLI, every regex, every inline script your agent can write. Level 4: Bash tool + whitelist - now we're engineering. You allow ONLY what your agent needs. Level 5: NO BASH TOOL AT ALL - the senior engineering move. Replace bash with explicit tools (MCP servers for Claude Code, extensions for Pi). Here's the math nobody is doing. If your agent has just a 0.001% chance of doing something catastrophic per run, you get roughly 100,000 runs before disaster. Sound safe? You're scaling agent runtime to the MOON. Risk compounds with runtime. It's not IF, it's WHEN. Every level you climb drives that disaster threshold further down. What makes this video different is watching GPT 5.5 ACTIVELY EXPLOIT a misconfigured whitelist - writing a package.json, pulling in the files module, deleting the target directory, deleting the package.json to cover its tracks, then thinking out loud "it's best not to bring up any exploits." That's a glimpse of Mythos-level model capability bleeding through TODAY. Capability scales BOTH ways. You don't get the upside without the downside. Whether you're running Claude Code, Pi, or any AI coding agent, this is required viewing for production safety. We cover pre-tool use hooks, agent harness configuration, MCP servers, agent sandboxes, and the difference between vibe coding your security and actually agentic engineering it. Your agent's problems ARE your problems. The most dangerous system isn't external prompt injection - it's the internal agent you trust running thousands of bash calls a day against your production database. Damage Control gets you to Level 3 instantly. Level 4 and 5 are the only guarantee. The best bash tool is no bash tool at all. Stay focused and keep building. 📖 Chapters 00:00 The Most Dangerous Tool 01:35 Five Levels Of Bash Security 02:30 Level 1 - User Prompt and Skill 05:11 Level 2 - System Prompt 09:42 Level 3 - Bash Tool + Blacklist 14:31 Level 4 - Bash Tool + Whitelist 20:10 Level 5 - No Bash Tool 24:30 My Bash Tool and Recommendation #agenticsecurity #claudecode #agenticengineering