Hugging Face Journal Club: Direct On-Policy Distillation

From the creator

The Hugging Face research team discusses the paper "Weak-to-Strong Generalization via Direct On-Policy Distillation" which proposes a cheap way to transfer the benefits of reinforcement learning from small models to much larger ones. Instead of imitating the smaller model directly, it measures how RL changed the small model’s policy and uses that change as a dense reward signal to train the larger model. The result is a form of weak-to-strong generalization that can match or outperform direct RL on the larger model at a fraction of the compute cost. Paper link: https://huggingface.co/papers/2607.05394

Choose to Build with AI
Matched to Machine Learning Project

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.