My Self Training Claude Code Setup: I Made Claude Code TRAIN ITSELF!

From the creator

Join the membership (members-only videos on every level from Bronze up): https://www.youtube.com/@AICodeKing/join In this video, I'll show you how to use AutoResearch with Claude Code to test whether Claude can improve the instructions it uses to fix bugs. We start with Karpathy's original AutoResearch idea and Udit Goenka's community plugin, then set up a bug-fix skill, a frozen evaluator, and three small Python bug tasks, and test the evaluator itself before paying for any runs. After that, we launch isolated trials in bare mode, write the /autoresearch loop prompt, and look at how to read candidate changes without getting fooled by overfitting to the practice tasks. Finally, we go over the real cost of the trials and how to adopt a winning version safely. -- Key Takeaways: 🧠 AutoResearch borrows Karpathy's loop: change one thing, check a fixed metric, then keep or discard the change. 🔀 Here the Claude model stays the same, and only the bug-fix skill's instructions are allowed to change. 🎯 "Make Claude better at coding" has no finish line, but "solve more of these frozen bug-fixing tasks under the same limits" can be checked. ✅ The evaluator is tested first: broken code must fail, a reviewed patch must pass, and deleted or edited tests must be rejected. 💻 Trials run in bare mode with --append-system-prompt-file, so they measure the instructions rather than memory, plugins or old sessions. ⚠️ Bare mode doesn't use your Claude subscription login, so trial calls need an API key or provider credentials. 🔍 A development score can rise just because the instructions overfit three examples, so the best version is checked on held-out tasks. 💸 Three iterations still means 48 coding trials, before counting the optimizer's own thinking, setup or rehearsals. ⚡ Use --max-budget-usd, --max-turns and an external timeout, and treat reported dollar amounts as estimates. 👍 Verdict: it's worth trying on a narrow workflow you repeat often and can grade reliably. Start small, keep the grader fixed, and trust fresh-task results over the optimizer's own description. -- Timestamps: 00:00 Intro 00:05 New members-only videos 00:24 Can Claude improve its own workflow? 01:10 What is AutoResearch? 01:52 Udit Goenka's AutoResearch plugin 02:21 The bug-fix skill experiment 02:43 Setup: installing the plugin 03:26 Creating the bug-fix skill 03:56 Building the evaluator with Claude Code 04:43 The development tasks 05:23 Testing the evaluator 05:55 Launching trials in bare mode 07:07 Baseline and scoring rules 08:16 The /autoresearch loop prompt 09:09 Candidate changes 10:27 How you can get fooled 11:11 Reading the results 12:15 What it costs 13:25 Adopting the winner 14:14 Verdict

Stop watching · start building
Matched to Claude Code

Makers residency 4th edition

Your idea can be anything. At our last workshops, people built plant-based nutrition businesses, a way to measure planetary resilience, crochet headwear and a way to connect lonely people living in the same building. An asset manager went on to raise an additional £25M on their fund within nine months. The fourth workshop is a full day at KOKO with Nick Sarafa, learning to build with Claude Code. Bring the thing you keep meaning to start. We’ll bring the pizza.

◆ Fri 13 Nov 2026 ◆ KOKO, London
Makers residency 4th edition
Live event
Makers residency 4th edition
Fri 13 Nov 2026
Find a job · 1 live role

Jobs in AI.
Apply now.

Apply once › get screened › meet the company

See all roles

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.