Research Notes, Announcements, and Technical Essays

Latest Posts

What We Learned Trying to Catch AI Liars: An Aletheia's Quest Retrospective

Aug 25, 2026

What we learned while building black-box and white-box detectors for AI deception during Aletheia's Quest.

Read post →

FineBooks: are open OCR models good enough to unlock historical knowledge?

Aug 10, 2026

Testing open OCR models on historical books with the FineBooks BHL OCR Leaderboard.

Read post →

A Dynamical Model of AI Governability

Jul 13, 2026

A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.

Read post →

Recent Archive

Attention Probes

Aug 1, 2025 · Stepan Shabalin, Nora Belrose

Adding attention to linear probes

Studying inductive biases of random networks via local volumes

Jun 12, 2025 · Louis Jaburi, Nora Belrose

In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a …

The Common Pile v0.1

Jun 5, 2025 · Stella Biderman, Sebastian Majstorovic, Aviya Skowron

Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

RLHF and RLAIF in GPT-NeoX

Oct 10, 2024 · Dakota Mahan, Quentin Anthony, Louis Castricato, Nathan Lile, Stella Biderman

GPT-NeoX now supports post-training thanks to a collaboration with SynthLabs.