← Maths & Science

The power law

A handful of distributions describe most of the quantities people actually care about — the size of earthquakes and cities, the distribution of wealth, the frequency of words, the number of links to a web page, the death toll of wars. They are almost never bell curves. They are power laws: P(x) ∝ x−α. This piece is about what a power law is, the single deep reason it keeps appearing, the small set of mechanisms that generate it, and — the part that matters most — how living in a power-law world should change the way you think.

What a power law is

A power law says the probability of an event falls as a fixed power of its size. Double the magnitude and the frequency drops by a constant factor — the same factor whether you are doubling from small to medium or from huge to enormous. That is the whole idea, and it has a startling consequence: there is no typical size. A bell curve has a characteristic scale — human height clusters around 170 cm and nobody is twice that. A power law has none; the ratio, not the value, is what repeats.

The cleanest way to see a power law is to plot it on log–log axes, where it becomes a straight line. Flip the toggle below between ordinary and log–log axes and watch what a heavy tail actually is.

Power law (x⁻¹·²) — heavy tail Exponential — thin tail
1.1.01.001.00011101001000
On log–log axes the power law is a straight line — its signature. The exponential holds up through the middle, then falls off a cliff; the power law just keeps going, and in the tail (far right) it ends up on top. That crossover is the heavy tail — the giant, rare events a bell curve says cannot happen.

Two distributions normalised to 1 at x = 1. Flip between the axes: the exponential (a stand-in for the bell curve's thin tail) collapses, while the power law keeps handing out large values. Straightness on log–log is the working test for a power law — and, as the article notes, a necessary but not sufficient one.

Where it shows up

The list is the argument. These are not loosely “skewed” distributions; they are, to good approximation, the same mathematical object appearing across physics, biology, economics and language:

DomainWhat follows a power lawExponent αName
Wealth & incomeShare held by the richest≈2–3Pareto (the “80/20”)
CitiesPopulation by rank≈2Zipf / rank-size
LanguageWord frequency by rank≈2Zipf's law
EarthquakesEnergy released≈2Gutenberg–Richter
NetworksLinks per node (web, citations)≈2.1–3Scale-free network
FirmsCompany size≈2Zipf
ConflictDeaths per war or attack≈1.6–2.5Richardson's law
MediaBook / music / film sales≈2–3The long tail
NatureForest fires, solar flares, avalanchesvariesSelf-organised criticality
The MoonCrater diameters≈2–3—

Why it is everywhere — the physical reasons

The deep answer is a single mathematical fact: the power law is the only function with no characteristic scale. If a process looks statistically the same when you zoom in or out — if f(bx) = g(b)·f(x) for any rescaling b — then f must be a power law. Nothing else is self-similar. So wherever a system has no built-in unit of size, a power law is not likely but forced. The question then becomes: why do so many systems end up scale-free? A surprisingly small set of mechanisms does it:

MechanismAssociated withHow it makes a power law
Preferential attachment (rich get richer)Yule, Simon; Barabási–AlbertThe more links / citations / wealth something already has, the faster it gains more. Proportional growth of a growing population mechanically yields a power law.
Proportional growth with a floorGibrat, Champernowne, ReedQuantities that grow by random percentages (not amounts) drift toward a lognormal — but add a lower barrier or random ages and the upper tail becomes an exact power law.
Self-organised criticalityBak, Tang, WiesenfeldThe sandpile: add grains one at a time and the pile tunes itself to the brink of collapse, where avalanches come in every size at once. No parameter is tuned; the criticality is automatic.
Critical phase transitionsKadanoff, Wilson (renormalisation group)Exactly at a phase transition the correlation length diverges and the system looks the same at every scale. Scale invariance forces power-law fluctuations — the deepest physical reason of all.
Optimisation under constraintCarlson & Doyle (HOT)Systems engineered to be robust to common shocks (the power grid, the internet, biology) buy that robustness with fragility to rare ones — producing heavy-tailed failures by design.
Scale invariance itselfMandelbrot (fractals)The power law is the ONLY function with no characteristic scale: f(bx) = g(b)·f(x) has only power-law solutions. Wherever self-similarity appears, a power law is mathematically forced.

Notice they are not variations on one theme. Rich-get-richer growth, self-organised criticality and phase transitions are genuinely different physics that happen to converge on the same output — which is exactly why the power law is so ubiquitous. The one thing they share is that they destroy any characteristic scale, and once the scale is gone the form is fixed.

What it should do to your thinking

This is the part with teeth. If much of the world is power-law distributed, several habits of thought inherited from the bell curve are not just imprecise — they are wrong.

  • The average can be meaningless — or infinite. For a power law with α ≤ 2 the variance is infinite; with α ≤ 1 the mean itself does not converge. “Average income,” “average city size,” “average outcome of a startup portfolio” can be dominated by a single observation you have not seen yet. In a heavy-tailed world, reasoning with the mean is often a category error.
  • The rare event carries the mass. In a bell curve the middle is where the action is; in a power law the tail is. Most of the deaths come from the few largest wars, most of the wealth from the few richest, most of a venture fund's return from one investment. Nassim Taleb's name for this world is Extremistan, against the thin-tailed Mediocristan of heights and coin flips — and the two demand opposite intuitions.
  • Gaussian risk models detonate. Calling a market crash a “25-sigma event” is not a description of a rare event; it is a confession that you used the wrong distribution. LTCM in 1998 and the 2008 models priced tail risk with bell curves and were destroyed by the tail the curve said could not exist.
  • Your sample almost never contains the big one. Because the largest events are rare, any finite dataset usually understates the tail. The catastrophe, the mega-bestseller, the hundred-year flood is disproportionately likely to be outside your window — so extrapolating from “what we've seen so far” systematically under-prepares you.
  • Think in ratios and orders of magnitude. The right mental axis for a power-law quantity is logarithmic. Ask “how many times bigger,” not “how much bigger”; expect the top few to dwarf the rest; and when you intervene, aim at the tail, because that is where the mass lives. The 80/20 rule is not a productivity slogan — it is a power law, taken literally.

An honesty note: not everything heavy-tailed is a power law

The concept is powerful enough to be over-used. A straight-ish line on a log–log plot is suggestive, not proof: lognormal, stretched-exponential and truncated distributions can masquerade as power laws over a decade or two of data. The influential Clauset–Shalizi–Newman analysis (2009) re-examined two dozen famous “power laws” and found many were statistically weak or better explained otherwise. The discipline is to fit the tail properly, quote the range over which the law holds, and prefer a mechanism to a curve fit. The idea — that scale-free processes produce heavy tails, and heavy tails break bell-curve intuition — survives intact even where a particular claimed exponent does not.

Appendix — how to read this page

Exponents are indicative. The α values in the table are commonly-cited round figures; real estimates vary by dataset, by the range fitted, and by method, and several are actively debated. They are here to show that the same object recurs across wildly different domains, not to be quoted precisely.

The scale-invariance argument, briefly

If a distribution is scale-free — p(bx) = g(b)·p(x) for all b — then differentiating and solving the resulting functional equation gives p(x) = C·x−α as the unique solution. That is the whole reason power laws and self-similarity (fractals) are two views of one thing, and why the renormalisation-group picture of critical phenomena — where a system looks identical at every scale — produces power laws as a matter of course.

Sources

  • Newman (2005). “Power laws, Pareto distributions and Zipf's law” — the standard survey of the phenomena and mechanisms.
  • Clauset, Shalizi & Newman (2009). “Power-law distributions in empirical data” — the rigorous fitting method, and a warning that many claimed power laws are not.
  • Bak (1996). How Nature Works — self-organised criticality and the sandpile.
  • Mandelbrot; Taleb. Fractal geometry; and The Black Swan / “Mediocristan vs Extremistan” on the consequences for thinking.