We're building a technology that will be able to manage a weapons production line and plot a coup, if it were inclined. We will increasingly depend on it to manage its own production and security. Some people think we'll see and fix problems as they arise; others think the important problems will stay hidden until we lose control. Can modelling the development shed light on the situation?
People in AI safety put the probability of AI takeover — "P(doom)" — anywhere from under 1% to over 99%, and the only thing anyone seems to agree on is that arguing about P(doom) is not very productive. But we have to decide how to manage AI development, and whether or not we'll stay in control matters, whether we can agree on it or not.
We've built a model of the contest for control of AI: between a cooperative team that tries to find and fix subversive behaviour and an uncooperative team that tries to evade detection and grow. We can map theories and accumulating evidence to outcomes, and focus our questions on the parameters that matter most. We want to grow this into a research programme that sharpens consensus around AI risk, identifies the levers and open questions that matter most, and gives the people steering AI development continuous tracking of the trajectory's health.
Our model considers a "race" between cooperative and uncooperative AI labour pools, plus a human pool which starts large and shrinks in importance. Humans and cooperative AI labour produce more cooperative labour, as well as accidentally producing some amount of uncooperative labour, which we term "leakage". They also produce "monitoring technology" which enables the observation of uncooperative AI behaviour. Uncooperative AI behaviour produces more uncooperative AI behaviour, as well as "evasion" which hides uncooperative behaviour from monitoring. The following diagram shows a simplified version of the model.
Our model features large uncertainty over several key parameters, including how effectively uncooperative behaviour propagates and at what rate the outputs of cooperative AI development accidentally leak into the uncooperative pool. Nonetheless, we calibrate the parameters according to our best inferences from publicly available data, and it delivers predictions.
Run forward from today's best-guess calibration, with business as usual:
Our model makes predictions, but if we want to have impact there are technical and social challenges to overcome: better measuring and projecting key parameters, and building credibility with the communities who need to consider takeover risks. What's the payoff from addressing them?
Simple models can be predictively effective: exponential models have been tremendously successful in forecasting many changes in AI. Epidemic curves, climate energy balances and bank stress tests all distil complex systems into a few decision-relevant variables, tell us what to measure, and when to worry.
Unless it can hide perfectly, we will get some information — and a model can help us understand how much we should attend to the information we get. We can also make other observations — like re-auditing old models with new tools — that can clarify misbehaviour rates in a way we can't by looking solely at misbehaviour on the frontier.
No — but we have some idea, and we expect to learn more as we develop the technology. We can apply new lessons to old models to check whether we have been heading in the right direction, and we can estimate how long the important lessons take to learn — so we don't lose control before we're able to see a problem.
There's a risk — but it is much harder for a deceiver to break the regularities in leakage and propagation rates that govern the dynamics than to break the regularities in the observed-misbehaviour series. Deception corrupts the measurement channel first; the mechanics are sturdier.
Yes:
All numbers on this page come from the full post's central calibration and are simplified in the telling; the post and the basin explorer are canonical, including every hedge this page skips. The early-warning framing is developed further in the companion LessWrong post.