Pi Coding Agent Observability: HTML Specs with Gemini 3.5 Flash and GPT Image 2

From the creator

If you don't measure your agents, you're not engineering. You're gambling with tokens. Everyone is racing to run MORE agents. Almost no one can answer the only question that matters: what did they actually cost, and was it worth it? 🚀 Get The Code Pi Agent Observability: https://github.com/disler/pi-agent-observability 🔥 Watch More Top 1 Opportunity For Sr Engineers: https://youtu.be/2KcITKKJikA Pi to Pi Agent Communication: https://youtu.be/PIdETjcXNIk Claude Code Observability Video (Claude Code Multi-Agent Orchestration): https://youtu.be/RpUTF_U4kiw 🔗 Resources Pi Coding Agent: https://pi.dev/ Pi Coding Agent (pi-mono): https://github.com/earendil-works/pi-mono Gemini 3.5 Flash: https://deepmind.google/models/gemini/flash/ GPT Image 2: https://openai.com/index/introducing-chatgpt-images-2-0/ Claude HTML Plans: https://claude.com/blog/using-claude-code-the-unreasonable-effectiveness-of-html ✅ Master Agentic Coding Tactical Agentic Coding: https://agenticengineer.com/tactical-agentic-coding?y=o4KZH_KSqYQ In this video we run three Gemini 3.5 Flash Pi coding agents against the exact same prompt with three different spec types, Markdown, HTML, and enhanced visual HTML, then we watch the whole thing live through a Pi agent observability dashboard. This is how you measure the trade-off triangle: performance, speed, and cost. More useful tokens beat fewer useful tokens. The keyword is useful. 📐 Markdown vs HTML vs Visual Specs Anthropic dropped the viral post on the unreasonable effectiveness of HTML. OpenAI dropped a benchmark-gapping image model in GPT Image 2. So which spec should you actually use? Instead of guessing, we test. Three Pi coding agents, same code base, same task, three different spec formats. The surprising part: the Markdown agent burned MORE tokens than the HTML agent on one run. Variance? Better focus? You will never know unless you observe it. That is the entire point. 👁️ Agent Observability Is The Control Surface The architecture is dead simple. Stream every event to a centralized server, persist it to a DB, read it back in a UI. But what it unlocks is everything. Swim lane view, single agent view, and a race mode that lines up agent turns side by side. Every tool call, every token count, every system prompt, every trace. Have you ever actually seen the full system prompt your agent boots with, every skill bloating your context? LLM observability is not a nice to have once you are running product-focused agents thousands of times a day. Agent traces ARE your leverage. 🎨 Visual Specs With GPT Image 2 and Gemini 3.5 Flash All my plans are visual specs now. GPT Image 2 lets me embed real interface mockups directly into the plan, the agent reads them, and multimodal monsters like Gemini 3.5 Flash execute against them with ease. Gemini wins on multimodal, full stop, and at that cost-per-intelligence the performance, speed, and cost trade-off is fantastic for product agents. A picture is worth a thousand prompts, and visual specs crush the planning constraint of agentic engineering. 🚀 The Agentic Value Chain This is tokenomics. Step one, use the tokens. Step two, generate value from the tokens. Step three, capture the revenue from that value. Observability is what moves you up that chain. Running a fleet of AI coding agents is a great place to start and a terrible place to end. You need valuable agents, then you need to capture their value, and you cannot do either blind. The Pi coding agent is the harness that makes this composable, one extension at a time. Want to go from vibe coding to engineering production agentic systems? Own the fundamentals of specs, harnesses, and observability and you adapt to any model, any agent SDK, any update. That is the whole game. Principled AI Coding: https://agenticengineer.com/principled-ai-coding Stop thinking in single prompts. Start thinking in systems. Specs, harnesses, traces, multi-agent orchestration. These are the building blocks of the next decade. Stay focused and keep building. Dan 📖 Chapters 00:00 Measure Your Agents to Improve 01:41 Pi Coding Agent Observability 03:18 Product Agent Observability with Gemini 3.5 Flash 09:54 Know Your Agent's System Prompt 11:07 Markdown vs HTML vs Visual HTML Specs 20:49 Pi Coding Agent Tokenomics #agentobservability #picodingagent #agenticengineering

Choose to Build with AI
Matched to AI Agent Systems

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.