Free GitHub Repo That Catches AI Lying About Code

From the creator

AI agents hallucinate code claims constantly, reverify, a free GitHub repo, proves when deterministic tools catch the lies better than another AI ever could. Fifty-five percent of enterprise decision-makers cite agent reliability and hallucination as a top challenge. Reverify, a nine-day-old open-source project from GitHub user 2akouwu, splits the verification job: an LLM proposes, then a deterministic tool checks that claim against ground truth before it's trusted. Unlike the failure mode of one AI reviewing another, reverify's pattern separates the model's guess from the verification itself, comparing Claude's binary analysis predictions against actual file bytes, control-flow heuristics from angr, and real function behavior through equivalence testing. Across 71 Windows binaries, the textbook prologue hallucination rate hit 97%; reverify caught 69 false claims while passing every genuine match. The benchmarks run on GitHub's CI runners, not the author's machine, and independent contributor IMGillision discovered and fixed an ARM64 routing bug by testing on Nvidia Jetson hardware, proving the verification pattern holds under real-world conditions. The tool ships as both an MCP server for Claude Code and Codex CLI, and a CLI gate for continuous integration pipelines. But reverify is candid about its boundaries: a known session consumed 909,000 tokens before stopping, the ARM64 routing bug existed, and high-level structural claims like call graphs are marked 'derived' not 'verified.' Most critically, nobody has yet measured whether agents actually improve with this judge attached, the headline numbers measure the checker and the hallucinating model separately, never together. For builders adopting AI agents into production workflows, this repo offers a working answer to the question: which of your agent's claims can anything other than another model actually check? Chapters: 0:00 Claude's textbook guess falls flat 2:22 69 hallucinations caught in 71 binaries 3:15 Building a test that refuses to cheat 4:33 Why GitHub's machines matter more 6:13 Six bugs only remote runners found 7:32 The ARM64 bug an outsider caught 9:20 MZ headers prove nothing useful 11:05 Side-by-side code execution never lies 12:45 One pip install puts it in reach 14:04 Ground truth survives the context reset 15:42 Thirty-two others tried, one broke through 17:26 909,000 tokens with no kill switch 18:55 Five percent confidence, not certainty 20:10 Call graphs live in the weak tier 21:38 The improvement nobody's measured yet 22:52 Can AI ever truly check AI? 24:11 Which claims can actually be verified? Tools & resources mentioned: - reverify: https://github.com/2akouwu/reverify - reverify (PyPI): https://pypi.org/project/reverify/ - Claude Code - Codex CLI - angr - GitHub deterministic-verification topic About The Stack The Stack helps you build with AI. Each video takes one tool, model, or workflow and shows how it works in a few focused minutes, with the real benchmarks and real costs. We go deep on Claude Code and Cursor for AI coding, AI agents and MCP servers, the open-source AI tools and GitHub repos most people miss, RAG and vector search, fine-tuning, and running local LLMs on your own machine with Ollama and LM Studio. We compare models like ChatGPT and Claude, test AI automation with Zapier, Make, and n8n, and flag the tools that actually ship. Subscribe for new breakdowns: https://www.youtube.com/@the-stack-ai?sub_confirmation=1 #ai agents #llm verification #ai quality control #ai hallucination #deterministic verification

Choose to Build with AI
Matched to Agent Harness

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.