Research Notes

The Common Pile v0.1

Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Multiple Choice Normalization in LM Evaluation

There are multiple ways of evaluating multiple-choice tasks on autoregressive language models like GPT-3/Neo/J. This post lays out the current prevalent normalization methods.