A Dynamical Model of AI Governability
Jul 13, 2026
A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.
Read post →
Early Indicators of Reward Hacking via Reasoning Interpolation
Apr 15, 2026
Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training
Read post →
Reward Hacking Research Update
Oct 7, 2025
Interim report on ongoing work on reward hacking
Read post →