Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights

From the creator

The Artificial Analysis Index is lying to you. 😲 Not on purpose, but an index is a proxy of a proxy, and by the time ten benchmarks get mashed into one number, the only thing that actually matters (which model to run for YOUR work) is gone. βœ… PUSH YOUR AGENTIC ENGINEERING BEYOND Tactical Agentic Coding: https://agenticengineer.com/tactical-agentic-coding?y=9weiIHy9T_0 😲 Phase 3 Early Sign Up is Officially Available: https://agenticengineer.com/ia3-pre-launch?y=9weiIHy9T_0 πŸŽ₯ VIDEO REFERENCES β€’ Claude Fable & Mythos 5.1: https://www.anthropic.com/claude-fable-and-mythos-5-1 β€’ GPT-6 Astra: https://openai.com/index/gpt-6-astra/ πŸ“Š MY TOP 5 BENCHMARKS β€’ Terminal-Bench v4.0: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 β€’ APEX Agents Leaderboard: https://www.mercor.com/apex/apex-agents-leaderboard/ β€’ AutomationBench: https://artificialanalysis.ai/evaluations/automationbench-aa β€’ AA-Omniscience: https://artificialanalysis.ai/evaluations/omniscience β€’ DeepSWE v1.1: https://deepswe.datacurve.ai/blog/deepswe-v1-1 Here's the question every agentic engineer should be able to answer: if you could only pick FIVE benchmarks, which five would you pick? I'm answering it, and I'm ranking GPT-6 Astra, Claude Fable 5.1, and the open-weights field against every one of them. πŸš€ Most engineers pick a model off a leaderboard headline and never look again. That's a mistake that compounds. Choosing a model is a THREE dimensional problem: performance, cost, and speed, together, as one unit. On Terminal-Bench v4.0, Astra and Fable 5.1 look close on score. Then you pull the cost axis and Astra is roughly 4x cheaper per task. Same benchmark. Completely different decision. πŸ› οΈ The five ai agent benchmarks and what each one actually tells you: Terminal-Bench v4.0: The cleanest pure agentic coding benchmark. Real container, real harness loop, verifier checks final state. This is your raw ai coding agents signal. APEX Agents: Investment banking, management consulting, corporate law. Expert-authored tasks, knowledge-worker style prompts. Your proxy for every domain that ISN'T software engineering. AutomationBench: 600+ tasks across finance, HR, marketing, ops, sales, support. The killer detail: you must complete the objective WITHOUT tripping guardrails. That's alignment measured at the floor level, and the ranking flips hard when you turn violations back on. AA-Omniscience: The hallucination benchmark. Correct, incorrect, partial, or NOT ATTEMPTED, with zero penalty for saying "I don't know." One hallucination upstream poisons every agent downstream in a long-running pipeline. DeepSWE v1.1: Long horizon software engineering from short, realistic prompts. If an agent ships on a lazy prompt, imagine what it does with a real plan. πŸ’£ The controversial part: I throw out benchmarks that look like a flat line. AA Long Context Retrieval? Dead to me. Saturation means zero information gain. I hunt for VARIANCE, because variance is where the alpha in ai model selection lives. That's where you find Gemini 3.8 Flash hitting the cost sweet spot, GLM 5.3 outperforming Kimi K3 where the index says otherwise, and DeepSeek quietly failing domains you might be deploying into. πŸ’‘ Why THESE five: every one of them is a proxy for out-loop agentic engineering. Long horizon work, no human in the loop, honest agents shipping on your behalf. Alignment, truthfulness, cross-domain competence, raw engineering skill, and a cost curve you can actually afford to run. AGI you can't pay for is irrelevant. 🌟 The move is a model stack, not a model. Combine compute, don't select compute. Keep one eye on the index and one eye on your OWN index, the three or four benchmarks that map to the work you actually do. That's the best ai model for coding for you, and nobody else can compute it. Comment below with your favorite benchmark. Not the index. A real one. Stay focused and keep building. Dan πŸ“– Chapters 00:00 AI Benchmarks Are Getting Blurry 01:50 Top 5 Benchmarks For Agentic Engineering 02:19 Benchmark #1: Terminal Bench 08:13 Benchmark #2: APEX Agents 11:55 Benchmark #3: AutomationBench 20:17 Benchmark #4: AA-Omniscience 25:32 Benchmark #5: DeepSWE 29:50 Why These 5 Benchmarks Matter For AI Model Selection #agenticengineering #gpt6 #agenticcoding

Choose to Build with AI
Matched to AI Agent Systems

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional Β£25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

β—† Fri 09 Oct 2026 β—† KOKO Cafe, London β—† With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.