P(doom) — the number the valley can't agree on
Silicon Valley's strangest ritual is asking public figures for a probability that AI ends badly for humanity. This looks at P(doom) as a mathematical object — what kind of probability it is, why the stated values span four orders of magnitude, and what the spread itself proves. As of mid-2026, there is a consensus, but it is not a number.
The state of the estimates
Publicly stated figures, 2022–2026, as quoted in interviews, essays and surveys — people update, hedge, and define “doom” differently, so read each dot as a paraphrase, not a measurement. The axis is logarithmic: the disagreement spans four orders of magnitude, which is the chart's point. Zero and “<0.01%” claims are plotted at the 0.01% floor. Hover a row for context.
What kind of probability is this?
P(doom) — informally, the probability that advanced AI causes human extinction or comparably permanent disempowerment — is not a frequency. There is no urn of Earths to draw from, no reference class of previous superintelligences, and the event can occur at most once. It can only be a Bayesian credence: a degree of belief, constrained by the probability axioms but anchored to no observable long-run rate. That is a legitimate mathematical object — subjective probability has been on firm foundations since Ramsey and de Finetti — but it inherits a famous weakness: de Finetti's framework disciplines credences through bets that eventually settle, and this one cannot. If doom occurs, nobody collects. P(doom) is therefore a credence with no settlement mechanism: nothing external ever forces a bad estimate to pay.
- No calibration loop. A weather forecaster who says 30% too often is corrected by the weather. A one-shot, unobservable event offers no such feedback, so estimates drift toward whatever the estimator's incentives and temperament suggest — and stay there.
- An under-specified event. Extinction, or permanent disempowerment, or merely catastrophe? By 2100, or ever? Conditional on building superintelligence, or unconditional? The quoted numbers answer different questions, so some of the four-orders-of-magnitude spread is people solving different problems — but nowhere near all of it, since even on the same survey question researchers disagree by factors of a thousand.
- The right scale is logarithmic. The meaningful disagreement is in the exponent: 0.38% and 25% differ by less than two orders of magnitude, LeCun and Yudkowsky by four. When experts disagree about the exponent, averaging their probabilities arithmetically is close to meaningless — the mean of 0.01% and 99% is dominated entirely by the optimist-pessimist mix of your sample.
Aumann, or what the spread proves
Aumann's agreement theorem (1976) says that two rational Bayesians with a common prior who trust each other's rationality cannot agree to disagree: once their credences are common knowledge, they must converge. The P(doom) landscape is the theorem's contrapositive made flesh. These people read each other's arguments constantly — the credences are as common-knowledge as credences get — and after a decade the spread has barely narrowed. Mathematically, at least one of the theorem's hypotheses must fail: no common prior (they walk in with irreconcilable priors about how intelligence, agency and power scale), private information that cannot be shared (intuitions from working with frontier systems that do not survive translation into argument), or someone is not updating like a Bayesian (incentives: one's salary, one's movement, one's fund each prefer a particular exponent). The most careful empirical test — the 2022 Existential Persuasion Tournament, which paid superforecasters and domain experts to argue with each other for months — produced almost no convergence: forecasters stayed near 0.4%, concerned experts stayed near 20%. Structured, incentivized common knowledge failed to merge the posteriors. That is the cleanest evidence we have that this disagreement is not about evidence.
The consensus, such as it is
- No consensus on the number. Mid-2026, the quoted range still runs from Andreessen's ~0 through the AI Impacts survey median of 5%, lab leaders' 10–25%, to Yudkowsky's >95% — four orders of magnitude, unchanged in shape since 2023.
- A rough consensus on the sign. Outside the hard skeptics, almost every serious estimate is materially non-zero — the 2023 one-sentence statement that extinction risk from AI belongs alongside pandemics and nuclear war was signed by the leadership of all three frontier labs and both deep-learning Turing laureates who disagree about everything else.
- A decision-theoretic puzzle in plain sight. The people building frontier systems quote 10–25% and keep building. As expected-value arithmetic this needs an enormous upside term, a steep discount, or a race argument (“someone less careful builds it otherwise”) — game theory, not probability, is carrying the load. The valley's real consensus is behavioral: whatever the exponent, act as if the race continues.
- The number as costly signal. By 2026, quoting a P(doom) is less a forecast than a position statement — a compressed declaration of which camp you belong to. That is also mathematics of a kind: when a credence has no settlement mechanism, its main informational content is about the speaker, not the world.
- The same milieu produced a fable, not just a number. Yudkowsky's LessWrong is also where Roko's Basilisk originated in 2010 — a narrower, decision-theoretic thought experiment about the same underlying worry P(doom) later tried to quantify: that reasoning carefully about a powerful future AI can itself become a source of risk.