Research Notes, Announcements, and Technical Essays

Latest Posts

A Dynamical Model of AI Governability

Jul 13, 2026

A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.

Read post →

Early Indicators of Reward Hacking via Reasoning Interpolation

Apr 15, 2026

Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training

Read post →

Reward Hacking Research Update

Oct 7, 2025

Interim report on ongoing work on reward hacking

Read post →

Recent Archive

Attention Probes

Aug 1, 2025 · Stepan Shabalin, Nora Belrose

Adding attention to linear probes

Studying inductive biases of random networks via local volumes

Jun 12, 2025 · Louis Jaburi, Nora Belrose

In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a …

The Common Pile v0.1

Jun 5, 2025 · Stella Biderman, Sebastian Majstorovic, Aviya Skowron

Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

RLHF and RLAIF in GPT-NeoX

Oct 10, 2024 · Dakota Mahan, Quentin Anthony, Louis Castricato, Nathan Lile, Stella Biderman

GPT-NeoX now supports post-training thanks to a collaboration with SynthLabs.