Don't hand AI a blank check.
Four levels of hands-on lessons on reward hacking, deceptive alignment, interpretability, and verifiable rewards - taught like a game, built for people who ship models.
Overall progress0 / 16 lessons
Level 1
Outer Alignment & Specification Gaming
Why a reward is never quite what you meant
Locked
Inner Alignment & Deception
What the model actually wants inside
Locked
Mechanistic Interpretability & SAEs
Reading the model's mind, neuron by neuron
Locked
Verifiable Rewards & RLVR
Rewarding what you can actually check
Progress is saved in your browser.