EleutherAI · a visual explainer

When AI builds AI, will we maintain control?

We're building a technology that will be able to manage a weapons production line and plot a coup, if it were inclined. We will increasingly depend on it to manage its own production and security. Some people think we'll see and fix problems as they arise; others think the important problems will stay hidden until we lose control. Can modelling the development shed light on the situation?

The question

People in AI safety put the probability of AI takeover — "P(doom)" — anywhere from under 1% to over 99%, and the only thing anyone seems to agree on is that arguing about P(doom) is not very productive. But we have to decide how to manage AI development, and whether or not we'll stay in control matters, whether we can agree on it or not.

We've built a model of the contest for control of AI: between a cooperative team that tries to find and fix subversive behaviour and an uncooperative team that tries to evade detection and grow. We can map theories and accumulating evidence to outcomes, and focus our questions on the parameters that matter most. We want to grow this into a research programme that sharpens consensus around AI risk, identifies the levers and open questions that matter most, and gives the people steering AI development continuous tracking of the trajectory's health.

The model

Our model considers a "race" between cooperative and uncooperative AI labour pools, plus a human pool which starts large and shrinks in importance. Humans and cooperative AI labour produce more cooperative labour, as well as accidentally producing some amount of uncooperative labour, which we term "leakage". They also produce "monitoring technology" which enables the observation of uncooperative AI behaviour. Uncooperative AI behaviour produces more uncooperative AI behaviour, as well as "evasion" which hides uncooperative behaviour from monitoring. The following diagram shows a simplified version of the model.

A · total AI labour — grows monitoring m fixed cooperative uncooperative evasion e fog over the hard-to-see O = 0.50 hover a dashed outline to see that part's impact coverage compounds — bigger lamp, more fixed → O rises self-propagation reaches deeper into the fog, and feeds it → O falls automation grows everything — and scale itself thickens the fog
Hover to see what each component affects. A simplified illustration of our model (with human labour excluded). O — the observed share of uncooperative behaviour — represents detect-and-fix coverage: those misbehaviours we have documented, which is the subset that we can work on fixing. The fog is "evasion" — a combination of active attempts to avoid observation and the passively accumulating difficulty of observing increasingly complex AI systems.

What the model predicts

Our model features large uncertainty over several key parameters, including how effectively uncooperative behaviour propagates and at what rate the outputs of cooperative AI development accidentally leak into the uncooperative pool. Nonetheless, we calibrate the parameters according to our best inferences from publicly available data, and it delivers predictions.

Run forward from today's best-guess calibration, with business as usual:

~25%
of AI labour ends up uncooperative on the central path — and small perturbations tip into a full takeover.
New problems on their way
Current uncooperative behaviour is almost all leakage, and its production rate will rise when (and if) propagation becomes another important source.
Observations: useful, but limited
It is good news if we see less misbehaviour as AI improves, but far from a guarantee of safety.
Sampled model trajectories: bands showing the spread of good and bad futures for the uncooperative share of AI labour, against automation growth
Plausible futures from the full post: each band is the 85% spread over a Monte-Carlo ensemble drawing the uncertain parameters. Good outcomes keep the uncooperative share low in perpetuity; bad ones bend upward. According to our model, both are on the table.

What problems does this solve?

Our model makes predictions, but if we want to have impact there are technical and social challenges to overcome: better measuring and projecting key parameters, and building credibility with the communities who need to consider takeover risks. What's the payoff from addressing them?

Consensus Alignment on settled questions, calibrated uncertainty on unsettled ones Disagreements often stall on duelling intuitions — advanced AI will be better at defeating safeguards, but will also help build better safeguards; we can learn from experience, but scheming models will deny us that experience. An explicit model does two things here: first, it pushes discussion towards specific parameters (and mechanisms) with existing operationalisations and estimates. Focusing on smaller questions raises the likelihood of resolution, and asks for better estimates in place of vague counter-intuitions. Secondly, it maps every estimate to an outcome, so duelling intuitions produce concrete differences (or lack thereof) in projected outcomes.
Evidence A place to put the news Every month brings a fresh report of scheming, sandbagging, or reward-gaming. Some get large reactions, some don't. Careful observers are often cautious about drawing conclusions. Are these signals truly alarming, or consistent with a healthy development programme? A model can continually ingest fresh evidence, update parameters and deliver impact estimates.
Interventions A first-pass policy impact analysis Scale up safety investment d/acc-style, slow down and reflect, stay laissez-faire, lean on liability — today these are compared by intuition. These can be parametrised and their impact estimated, in relation to other interventions and in relation to the estimated task of delivering good outcomes with high confidence. Even simple models can separate the levers that move the outcome from the ones that don't.
Regulatory targets An instrument you could steer by How do we know if our policy approach can handle danger signals before it's too late? We can model that too, and give concrete guidance for what we need to achieve for robustness. That's the early-warning proposal.

Questions you should ask

Isn't loss of control too complicated to model?

Simple models can be predictively effective: exponential models have been tremendously successful in forecasting many changes in AI. Epidemic curves, climate energy balances and bank stress tests all distil complex systems into a few decision-relevant variables, tell us what to measure, and when to worry.

If uncooperative AI is trying to hide from us, won't we be unable to observe it before it's too late?

Unless it can hide perfectly, we will get some information — and a model can help us understand how much we should attend to the information we get. We can also make other observations — like re-auditing old models with new tools — that can clarify misbehaviour rates in a way we can't by looking solely at misbehaviour on the frontier.

Do we even know how we need models to behave for good outcomes?

No — but we have some idea, and we expect to learn more as we develop the technology. We can apply new lessons to old models to check whether we have been heading in the right direction, and we can estimate how long the important lessons take to learn — so we don't lose control before we're able to see a problem.

Doesn't deceptive behaviour break the very regularities that make simple models effective?

There's a risk — but it is much harder for a deceiver to break the regularities in leakage and propagation rates that govern the dynamics than to break the regularities in the observed-misbehaviour series. Deception corrupts the measurement channel first; the mechanics are sturdier.

Does the model make empirical predictions?

Yes:

Concrete recommendations

Critical question What happens when misbehaviour no longer needs us? Uncooperative models and traits today aren't viable without human help. That will change. How that shift moves leakage and self-propagation is the single most important thing to understand — the better we understand it, the better we can predict.
Retrospective audits Re-score old models with new audit harnesses. Rising catch-rates on frozen models estimate what today's audits are missing; converging catch-rates indicate we've nearly found it all.
Archive everything Preserve model weights, harnesses, and development context so those audits stay possible.

All numbers on this page come from the full post's central calibration and are simplified in the telling; the post and the basin explorer are canonical, including every hedge this page skips. The early-warning framing is developed further in the companion LessWrong post.

Go deeper

The full analysis The complete model: equations, parameter calibrations from public evidence, predictions and postdictions, and the AI-2027 comparison.
The basin explorer Every parameter on a slider, one-click presets for the calibrations in the post, and an outcome map of which worlds end where. Disagree with a number? Move it.
The early-warning question Can this become a real monitoring regime? Zones, response lags, and what it would take — discussion welcome.