The power law
A handful of distributions describe most of the quantities people actually care about — the size of earthquakes and cities, the distribution of wealth, the frequency of words, the number of links to a web page, the death toll of wars. They are almost never bell curves. They are power laws: P(x) ∝ x−α. This piece is about what a power law is, the single deep reason it keeps appearing, the small set of mechanisms that generate it, and — the part that matters most — how living in a power-law world should change the way you think.
What a power law is
A power law says the probability of an event falls as a fixed power of its size. Double the magnitude and the frequency drops by a constant factor — the same factor whether you are doubling from small to medium or from huge to enormous. That is the whole idea, and it has a startling consequence: there is no typical size. A bell curve has a characteristic scale — human height clusters around 170 cm and nobody is twice that. A power law has none; the ratio, not the value, is what repeats.
The cleanest way to see a power law is to plot it on log–log axes, where it becomes a straight line. Flip the toggle below between ordinary and log–log axes and watch what a heavy tail actually is.
Two distributions normalised to 1 at x = 1. Flip between the axes: the exponential (a stand-in for the bell curve's thin tail) collapses, while the power law keeps handing out large values. Straightness on log–log is the working test for a power law — and, as the article notes, a necessary but not sufficient one.
Where it shows up
The list is the argument. These are not loosely “skewed” distributions; they are, to good approximation, the same mathematical object appearing across physics, biology, economics and language:
| Domain | What follows a power law | Exponent α | Name |
|---|---|---|---|
| Wealth & income | Share held by the richest | ≈2–3 | Pareto (the “80/20”) |
| Cities | Population by rank | ≈2 | Zipf / rank-size |
| Language | Word frequency by rank | ≈2 | Zipf's law |
| Earthquakes | Energy released | ≈2 | Gutenberg–Richter |
| Networks | Links per node (web, citations) | ≈2.1–3 | Scale-free network |
| Firms | Company size | ≈2 | Zipf |
| Conflict | Deaths per war or attack | ≈1.6–2.5 | Richardson's law |
| Media | Book / music / film sales | ≈2–3 | The long tail |
| Nature | Forest fires, solar flares, avalanches | varies | Self-organised criticality |
| The Moon | Crater diameters | ≈2–3 | — |
Why it is everywhere — the physical reasons
The deep answer is a single mathematical fact: the power law is the only function with no characteristic scale. If a process looks statistically the same when you zoom in or out — if f(bx) = g(b)·f(x) for any rescaling b — then f must be a power law. Nothing else is self-similar. So wherever a system has no built-in unit of size, a power law is not likely but forced. The question then becomes: why do so many systems end up scale-free? A surprisingly small set of mechanisms does it:
| Mechanism | Associated with | How it makes a power law |
|---|---|---|
| Preferential attachment (rich get richer) | Yule, Simon; Barabási–Albert | The more links / citations / wealth something already has, the faster it gains more. Proportional growth of a growing population mechanically yields a power law. |
| Proportional growth with a floor | Gibrat, Champernowne, Reed | Quantities that grow by random percentages (not amounts) drift toward a lognormal — but add a lower barrier or random ages and the upper tail becomes an exact power law. |
| Self-organised criticality | Bak, Tang, Wiesenfeld | The sandpile: add grains one at a time and the pile tunes itself to the brink of collapse, where avalanches come in every size at once. No parameter is tuned; the criticality is automatic. |
| Critical phase transitions | Kadanoff, Wilson (renormalisation group) | Exactly at a phase transition the correlation length diverges and the system looks the same at every scale. Scale invariance forces power-law fluctuations — the deepest physical reason of all. |
| Optimisation under constraint | Carlson & Doyle (HOT) | Systems engineered to be robust to common shocks (the power grid, the internet, biology) buy that robustness with fragility to rare ones — producing heavy-tailed failures by design. |
| Scale invariance itself | Mandelbrot (fractals) | The power law is the ONLY function with no characteristic scale: f(bx) = g(b)·f(x) has only power-law solutions. Wherever self-similarity appears, a power law is mathematically forced. |
Notice they are not variations on one theme. Rich-get-richer growth, self-organised criticality and phase transitions are genuinely different physics that happen to converge on the same output — which is exactly why the power law is so ubiquitous. The one thing they share is that they destroy any characteristic scale, and once the scale is gone the form is fixed.
What it should do to your thinking
This is the part with teeth. If much of the world is power-law distributed, several habits of thought inherited from the bell curve are not just imprecise — they are wrong.
- The average can be meaningless — or infinite. For a power law with α ≤ 2 the variance is infinite; with α ≤ 1 the mean itself does not converge. “Average income,” “average city size,” “average outcome of a startup portfolio” can be dominated by a single observation you have not seen yet. In a heavy-tailed world, reasoning with the mean is often a category error.
- The rare event carries the mass. In a bell curve the middle is where the action is; in a power law the tail is. Most of the deaths come from the few largest wars, most of the wealth from the few richest, most of a venture fund's return from one investment. Nassim Taleb's name for this world is Extremistan, against the thin-tailed Mediocristan of heights and coin flips — and the two demand opposite intuitions.
- Gaussian risk models detonate. Calling a market crash a “25-sigma event” is not a description of a rare event; it is a confession that you used the wrong distribution. LTCM in 1998 and the 2008 models priced tail risk with bell curves and were destroyed by the tail the curve said could not exist.
- Your sample almost never contains the big one. Because the largest events are rare, any finite dataset usually understates the tail. The catastrophe, the mega-bestseller, the hundred-year flood is disproportionately likely to be outside your window — so extrapolating from “what we've seen so far” systematically under-prepares you.
- Think in ratios and orders of magnitude. The right mental axis for a power-law quantity is logarithmic. Ask “how many times bigger,” not “how much bigger”; expect the top few to dwarf the rest; and when you intervene, aim at the tail, because that is where the mass lives. The 80/20 rule is not a productivity slogan — it is a power law, taken literally.
An honesty note: not everything heavy-tailed is a power law
The concept is powerful enough to be over-used. A straight-ish line on a log–log plot is suggestive, not proof: lognormal, stretched-exponential and truncated distributions can masquerade as power laws over a decade or two of data. The influential Clauset–Shalizi–Newman analysis (2009) re-examined two dozen famous “power laws” and found many were statistically weak or better explained otherwise. The discipline is to fit the tail properly, quote the range over which the law holds, and prefer a mechanism to a curve fit. The idea — that scale-free processes produce heavy tails, and heavy tails break bell-curve intuition — survives intact even where a particular claimed exponent does not.
Appendix — how to read this page
The scale-invariance argument, briefly
If a distribution is scale-free — p(bx) = g(b)·p(x) for all b — then differentiating and solving the resulting functional equation gives p(x) = C·x−α as the unique solution. That is the whole reason power laws and self-similarity (fractals) are two views of one thing, and why the renormalisation-group picture of critical phenomena — where a system looks identical at every scale — produces power laws as a matter of course.
Sources
- Newman (2005). “Power laws, Pareto distributions and Zipf's law” — the standard survey of the phenomena and mechanisms.
- Clauset, Shalizi & Newman (2009). “Power-law distributions in empirical data” — the rigorous fitting method, and a warning that many claimed power laws are not.
- Bak (1996). How Nature Works — self-organised criticality and the sandpile.
- Mandelbrot; Taleb. Fractal geometry; and The Black Swan / “Mediocristan vs Extremistan” on the consequences for thinking.