What We Learned Trying to Catch AI Liars: An Aletheia's Quest Retrospective
What we learned while building black-box and white-box detectors for AI deception during Aletheia's Quest.
Read post →What we learned while building black-box and white-box detectors for AI deception during Aletheia's Quest.
Read post →Testing open OCR models on historical books with the FineBooks BHL OCR Leaderboard.
Read post →A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the basin boundary lies, what current evidence says about which side we are on, and what would tell us we are on the good path.
Read post →Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training
Interim report on ongoing work on reward hacking
Announcing Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
Adding attention to linear probes
Research update on applying local volume measurement to downstream tasks
In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a …
Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Using Product Key Memories to encode sparse coder features
Comparing feature overlap and interpretability across TopK sparse autoencoders trained with different random seeds.
Using interpretations of SAE latents to simulate activations.
A demonstration of third-party evaluation of private language-model training data using PySyft.
Interim report on ongoing work on mechanistic anomaly detection
GPT-NeoX now supports post-training thanks to a collaboration with SynthLabs.
Exploring the implementation details of muTransfer
Interim report on ongoing work on mechanistic anomaly detection