← Business

The academic critiques of GDP

GDP is the most consequential number in public life and the most criticised. Most of the criticism is a quotation — Robert Kennedy in 1968, that it measures everything except that which makes life worthwhile — and quotations do not tell you how much is at stake, or which objections a statistician would concede tomorrow and which cut at the foundations. This is a review of the academic literature, organised around a single question: under what conditions can a price-weighted sum of quantities be read as a measure of how well things are going? There are three, each fails in a characteristic way, and almost every critique in ninety years of argument is the failure of exactly one of them.

What the number is

Gross domestic product is the market value of final goods and services produced inside a territory in a period. It is built three ways that must agree — as production (value added by industry), as expenditure (consumption, investment, government, net exports) and as income (wages, profits, taxes less subsidies) — and the agreement between three independent routes to the same total is the main reason the number is trusted at all.

It was not designed to measure welfare. William Petty made the first estimates in the 1660s to work out what the Dutch wars could be financed with. Simon Kuznets built the first official US series in 1934 for a Senate that wanted to know the size of the Depression. Keynes turned the exercise into capacity planning in How to Pay for the War in 1940, and Richard Stone and James Meade gave the accounts the double-entry structure that became the UN System of National Accounts in 1953. The measure is a war-economy instrument: how much output exists, how much of it the state can take, how much civilian consumption must fall. Every critique below should be read against that origin.

The three conditions

Real GDP is a quantity index with prices as weights. Write the change in output as

dY = Σᵢ pᵢ · dqᵢ

and compare it with the change in social welfare, which for a utility function U over the same goods is

dU = Σᵢ (∂U/∂qᵢ) · dqᵢ

The two sums are the same object with different weights. They move together if and only if pᵢ is proportional to ∂U/∂qᵢ for every good — that is, if the price of each thing equals what one more unit of it is worth. That is the first condition, and stating it exposes the other two: the derivative ∂U/∂qᵢ belongs to someone, so the step requires a household whose marginal utility can stand for everyone else's; and dqᵢ must be a net flow, since output produced by consuming a stock is not a gain in anything. Underneath all three sits a practical condition — that the index can be built and compared at all — and beside them, a political one about how the number is used.

Prices are marginal social values

Every good in the sum carries a price equal to what one more unit of it is worth to society, so that the weights in the index are the weights welfare would use.

Makes the price vector p proportional to the gradient of the social utility function. Without it the index is still a quantity aggregate, but it aggregates with the wrong weights — or omits the good entirely, at a weight of zero.

One household stands in for all

The economy behaves as if a single representative household consumed the whole product, so that a rise in the total is a rise for the agent whose utility is being tracked.

Licenses the step from an aggregate to a statement about welfare. Without it the index tracks a mean, and the mean of a skewed distribution is nobody's experience.

The capital consumed has been subtracted

What is measured is what can be consumed while leaving the productive base intact — Hicksian income, not gross output.

Turns a flow into a sustainable flow. Without it an economy liquidating its forests, its fisheries or its atmosphere books the proceeds as income and never books the loss.

The index can actually be built

Nominal values can be split into price and quantity, quality change can be held constant, and the same basket can be priced in two countries or two decades.

The practical floor under the other three. Deflators, hedonic adjustment, base years and purchasing-power parities are all choices, and each choice moves the answer.

The measure is not the objective

GDP is used as an observation of production, not as the thing institutions are rewarded for maximising.

Not a property of the index but of its use. Once growth in the measure becomes the objective of policy, the measure starts to be managed rather than observed.

Kuznets got there first

The strongest early critique of GDP is in the document that created it. Kuznets's 1934 report to the Senate warns, in its own text, that the welfare of a nation can scarcely be inferred from a measurement of national income; he spent the following decades arguing that the accounts should exclude activity he considered a cost of doing business rather than a benefit, and lost that argument to the wartime planners who wanted the biggest possible measure of capacity. In 1962 he put the point in the form it is still quoted in: goals for more growth should specify more growth of what, and for what.

The two-sentence version of everything below. GDP was designed as a production measure for a state that needed to know its capacity, and it does that job better than any competing statistic. It is read as a welfare measure by people who know it is not one, and the gap between the two readings is where every serious critique lives.

How big are the gaps?

Before the arguments, the orders of magnitude. The chart sizes each documented gap as a share of GDP, on a log axis, with the range each source gives rather than a false point estimate. The colours separate three different kinds of objection that are usually run together: output the accounts leave out, spending they count that arguably adds nothing to welfare, and numbers inside the total that no transaction ever produced.

0.2%0.5%1%2%5%10%20%50%100%Risk premium booked as bank outputImputations and the financial sec…0.3–0.4%Illegal production (drugs, sex work)The production boundary0.6–1%Consumer surplus from free digital goodsFree and digital goods · editorial0.5–3%Natural-capital depletionGross, not net · editorial1–5%Defensive and regrettable spendingDefensive and regrettable expendi… · editorial3–10%Imputed rent on owner-occupied housingImputations and the financial sec…6–8%All imputations in the accountsImputations and the financial sec…14–15%Household and unpaid workThe production boundary26–40%Base-year error, low-income countriesComparability across countries30–89%
Hover a row for the source. Note the log axis: unpaid household work is not a bigger version of the others, it is two orders of magnitude above the financial-sector adjustment that gets far more attention.

Each bar is a range, in per cent of GDP as a level. Ranges marked editorial are this page's own estimate across published studies; the rest are taken from the source named in the readout. They are not additive — some are output GDP omits, some are spending it includes, and the imputations are already inside the number.

Three things stand out. In a rich country the largest gap by an order of magnitude is unpaid household work, which is the item the public argument spends least time on; the only bar above it belongs to low-income countries, where the level of GDP itself can be wrong by more than everything else combined until the next rebasing. The constructed part of the accounts — roughly one dollar in seven of the US total — is bigger than most of the omissions people complain about. And the financial-sector adjustment that generated the most heat after 2008 is worth about 0.3% of GDP: real, correct, and rounding error next to the boundary questions it sits beside.

The critiques

Eleven families, grouped by the condition each one says fails. They are stated with the best evidence for them and the strongest available reply, because a critique whose rebuttal is not stated is an opinion.

Prices are marginal social values

The production boundary

Work that is not transacted is not counted, so the same activity enters or leaves GDP depending on who performs it and whether money changes hands.

The national accounts count production that crosses a market. Cooking, care, cleaning and child-rearing done at home are production by any technical definition, but they have no price, so they enter the sum at a weight of zero. The textbook consequence — a man who marries his housekeeper lowers GDP — is not a paradox but a straightforward reading of the boundary.

Evidence. The BEA's own household satellite account puts the omission at 26% of US GDP in 2010, down from 39% in 1965 — and the fall is not a fall in household work's value, it is women moving into paid employment, which registers as growth twice over (Bridgman et al., Survey of Current Business, 2012).

The reply. The boundary is not an oversight; it is what makes the number reproducible. Valuing unpaid time requires choosing a wage to value it at, and the choice between a housekeeper's wage and the forgone wage of the person doing the work moves the estimate by tens of per cent. Statistical offices publish it as a satellite account precisely so the choice is visible rather than buried in the headline.

Free and digital goods

A good with a price of zero contributes zero to GDP however much consumer surplus it generates, so a large part of digital welfare is invisible to the accounts.

GDP measures expenditure, and expenditure on a free service is nil. Search, maps, encyclopaedias and messaging enter only through the advertising that funds them, which is a fraction of what users would pay, and the substitution of a free good for a paid one can register as a contraction.

Evidence. In incentive-compatible choice experiments the median user required about $48 to give up Facebook for a month; on the authors' GDP-B accounting, including Facebook alone would have added 0.05 to 0.11 percentage points a year to measured US growth (Brynjolfsson, Collis, Diewert, Eggers & Fox, 2019).

The reply. Consumer surplus has always been outside GDP — it is outside it for bread as much as for search. Measuring it requires stated-preference experiments whose numbers are an order of magnitude less stable than transaction data. And the sceptical result matters: Byrne, Fernald and Reinsdorf (2016) found digital mismeasurement too small, and moving the wrong way, to explain the post-2004 productivity slowdown.

Defensive and regrettable expenditure

Spending that repairs damage, buys security or compensates for congestion adds to GDP exactly as spending that creates enjoyment does.

The accounts record transactions, not their purpose. Commuting, prisons, clean-ups, litigation, flood defences and the medical treatment of pollution-induced illness are all final expenditure. An oil spill raises GDP twice — once when the oil is sold, once when the coast is cleaned.

Evidence. Nordhaus and Tobin's Measure of Economic Welfare, which deducted the regrettables and added leisure and household work, grew at 1.1% a year per head from 1929 to 1965 against 1.7% for net national product — slower, but still growth (Nordhaus & Tobin, 1972).

The reply. Sorting expenditure into defensive and enjoyable requires a theory of what people ought to want. A lock on a door, a commute to a better job and a hospital bed are all inputs to something someone valued; a statistician who deducts them has substituted their judgement for the buyer's, and the deductions in ISEW-style indices are where most of their divergence from GDP comes from.

Unpriced damage

Production that imposes costs on people outside the transaction books the revenue and not the cost, so the measured contribution of damaging industries is systematically too high.

If a factory's emissions shorten lives downwind, the electricity is priced and the shortened lives are not. The index sums priced flows, so the damage enters at zero, and an economy can raise measured output by shifting activity toward the industries whose costs are borne elsewhere.

Evidence. Muller, Mendelsohn and Nordhaus (2011) priced US air-pollution damage industry by industry and found gross external damages exceeding measured value added in the worst sectors — coal-fired electricity generation destroyed more value than it added, on their estimates.

The reply. This is an argument for environmental accounts, not against GDP, and it has largely been won: damage estimates now sit in satellite accounts and, since the 2025 revision of the System of National Accounts, depletion of minerals and fossil fuels is treated as a cost of production rather than as income.

Imputations and the financial sector

A large slice of GDP is not observed at all but constructed by the statistician, and the construction embeds contestable theory — nowhere more than in finance.

Where there is no transaction the accounts sometimes invent one anyway. Owner-occupiers are treated as renting to themselves; banks are credited with output equal to an interest-rate margin (FISIM) that rises when risk premia rise. The boundary is therefore not a principle but a set of case-by-case decisions.

Evidence. Imputations are roughly 15% of US GDP, of which owner-occupied housing is about 6–8 points (Mayerhauser & Reinsdorf, BEA, 2007). Adjusting bank output for the risk premium embedded in FISIM cuts measured US bank output by about 21% and GDP by about 0.3% (Basu, Inklaar & Wang, 2011) — and on the unadjusted method, the UK financial sector recorded its fastest measured growth on record in late 2008.

The reply. The alternative to imputing is worse: without imputed rent, a country of owners looks poorer than an identical country of renters, and the measure would jump whenever tenure shifted. The FISIM critique is correct and is a live methodological debate, not a hidden one.

One household stands in for all

The mean is not the median

GDP per capita is an average, and when growth accrues to the top of the distribution the average can rise for decades while the typical household's income does not move.

Dividing a total by a population destroys the information that matters for almost every question the total is used to answer. The representative-household assumption that licenses reading the aggregate as welfare is precisely the assumption that fails when dispersion widens.

Evidence. Distributional national accounts for the US show average pre-tax national income per adult up about 60% since 1980 while the bottom half of the distribution stagnated near $16,000; the ratio of top-1% to bottom-50% average income went from 27 to 81 (Piketty, Saez & Zucman, QJE 2018).

The reply. This is a critique of using GDP per capita alone, not of GDP. The fix is additive rather than substitutive — publish median equivalised household income and the distribution of growth next to the total, which was the Stiglitz-Sen-Fitoussi Commission's core recommendation and is now routine in OECD statistics.

Welfare has other arguments

Life expectancy, leisure, health, security and the freedom to do what one values enter welfare directly, and none of them is a component of output.

Consumption is one argument of a utility function, not the whole of it. A country that buys its output with longer hours, shorter lives or greater insecurity has paid for it in a currency the accounts do not record, which is Sen's capabilities objection in its economic form.

Evidence. Jones and Klenow's consumption-equivalent welfare measure, which adds leisure, mortality and inequality to consumption, puts French welfare at 92% of the US level while French income per head is 67% of it — lower mortality, lower inequality and more leisure each add roughly 10 points (AER, 2016).

The reply. The same paper is the strongest defence of GDP in the literature: the authors' own summary is that welfare is highly correlated with GDP per person, even though individual deviations are often large. GDP is a coarse proxy for welfare, but an extraordinarily good one, and the deviations are systematic enough to be corrected rather than fatal.

The capital consumed has been subtracted

Gross, not net

The G in GDP is gross: it does not subtract the capital used up in producing the flow, and it recognises no natural capital at all, so liquidation reads as income.

Hicks defined income as the maximum a person can consume in a week and still be as well off at the end of it. GDP is not that. A country that cuts its forests, pumps its aquifers or exhausts a field books the sale and never books the depletion of the stock that produced it.

Evidence. Between 1992 and 2014 produced capital per person roughly doubled while the stock of natural capital per person fell by nearly 40% (Dasgupta Review, 2021). The theory has been settled since Weitzman (1976) showed that it is net product, comprehensively defined, that has a welfare interpretation — and Hartwick (1977) gave the rule for keeping it constant.

The reply. Net figures exist — NDP, adjusted net savings, inclusive wealth — and are published. They are ignored because depreciation is the least reliable estimate in the accounts, and because netting out a number nobody can measure well buys accuracy in principle at the cost of precision in practice.

The index can actually be built

Index numbers and quality change

Splitting a nominal total into real quantity and price requires index-number choices that have no uniquely correct answer, and the choices move growth by enough to matter.

Real GDP is the nominal total deflated by a price index, so every question about what a dollar buys — whether this year's phone is the same good as last year's, which base year weights the basket, how a new good enters — is a question about the growth rate. Gerschenkron showed in 1947 that the base year alone can swing measured industrial growth dramatically.

Evidence. The Boskin Commission (1996) put US CPI overstatement at about 1.1 percentage points a year; Lebow and Rudd (2003) revisited it at roughly 0.9, within a plausible range of 0.3 to 1.4. An error of that size compounds to more than a quarter of the level over three decades.

The reply. Index-number theory is one of the better-developed parts of economics: Diewert's superlative indices, chain weighting and hedonic quality adjustment exist precisely to bound these errors, and statistical agencies adopted them. The residual uncertainty is acknowledged and roughly quantified, which is more than can be said for any alternative measure.

Comparability across countries

Cross-country GDP comparisons rest on purchasing-power parities and on statistical capacity, and in much of the world both are weak enough to swamp the differences being discussed.

Comparing two economies requires pricing a common basket in both, but the basket that is representative in one country is not representative in the other, and the weights chosen determine the answer. Underneath that sits the raw data: a national account is only as good as the surveys, censuses and base-year benchmarks feeding it.

Evidence. Ghana's 2010 rebasing raised measured GDP by about 60%; Nigeria's 2014 rebasing raised it by 89% and made it Africa's largest economy overnight, with no change in anything real (Jerven, Poor Numbers, 2013). Deaton and Heston (2010) document how much PPP-based comparisons move with defensible methodological choices.

The reply. Known, and the response has been more measurement rather than less: successive rounds of the International Comparison Program, published revision histories, and explicit uncertainty intervals. A number with a stated error bar of ±30% still orders the world better than no number.

The measure is not the objective

GDP as a target

The serious problem is not what GDP measures but that governments are graded on it, which converts a statistic into an objective and an objective into a constraint on policy.

Once growth in the measure is the thing being optimised, every gap above becomes an exploitable margin, and the institutions producing the number are under pressure from the institutions being judged by it. This is Goodhart's Law applied to the most consequential statistic in public life.

Evidence. The degrowth and steady-state literature (Daly; Jackson, Prosperity Without Growth, 2009; Raworth, Doughnut Economics, 2017) argues the target is incompatible with biophysical limits; the empirical decoupling debate over whether output can grow while material throughput falls remains unresolved.

The reply. The target critique is about policy, not measurement, and it cuts both ways: no government has ever been graded on a dashboard, because a dashboard cannot arbitrate a trade-off. Abolishing the number does not abolish the pressure it was serving, and growth remains the only mechanism that has reliably moved poor countries' life expectancy and child mortality.

The literature

40 works, with venue and year so they can be checked, tagged by the critique they bear on. Findings are stated as the authors reported them; the sorting into conditions is this page's synthesis, not theirs.

WorkFindingBears on
Verbum Sapienti
William Petty (1665), Political arithmetic, England
First systematic estimate of a nation's income and stock, built to work out how much tax the Dutch wars could raise. National accounting begins as a fiscal instrument, not a welfare one.Origins of the measure
National Income, 1929–1932
Simon Kuznets (1934), US Senate document 124
The report that gave the United States its first official national income series, and which states in its own text that the welfare of a nation can scarcely be inferred from a measurement of national income.Origins of the measure
How to Pay for the War
John Maynard Keynes (1940), Macmillan
Turned national income estimation into a planning tool: how much output exists, how much the state can take, how much civilian consumption must fall. The modern accounts are a wartime capacity-planning device.Origins of the measure
National income white paper and the double-entry accounts
Richard Stone & James Meade (1941), UK Treasury
Gave the accounts their double-entry structure, later the basis of the 1953 System of National Accounts. Stone's framework, not Kuznets's, is what the world adopted.Origins of the measure
System of National Accounts 2025
United Nations Statistical Commission (2025), UN, adopted March 2025
The current standard. Treats depletion of minerals, coal, oil and gas as a cost of production, recognises renewable energy resources as assets, and brings data into the asset boundary — a direct answer to the netting and valuation critiques.Origins of the measure
Economics of Household Production
Margaret Reid (1934), Wiley
Established the third-person criterion for what counts as production — if you could pay someone else to do it, it is production — and so defined the boundary the accounts decline to cross. The production boundary
If Women Counted
Marilyn Waring (1988), Harper & Row
The feminist critique of the production boundary: the accounts systematically value what men were paid for and zero what women did unpaid, with direct consequences for which policies look productive. The production boundary
Accounting for Household Production in the National Accounts, 1965–2010
Bridgman, Dugan, Lal, Osborne & Villones (2012), Survey of Current Business 92(5)
The official US estimate of the gap: including nonmarket household production raises nominal GDP by 39% in 1965 and 26% in 2010. The decline is women's paid-work participation, not a fall in household output's value. The production boundary
GDP-B: Accounting for the Value of New and Free Goods
Brynjolfsson, Collis, Diewert, Eggers & Fox (2019), NBER WP 25695 / AEJ: Macro
Builds a benefits-based index from incentive-compatible valuation experiments. Median user compensation to drop Facebook for a month: about $48; including it alone adds 0.05–0.11pp a year to measured growth. Free and digital goods
Does the United States Have a Productivity Slowdown or a Measurement Problem?
Byrne, Fernald & Reinsdorf (2016), Brookings Papers on Economic Activity
The sceptical result: digital mismeasurement is real but too small, and moving in the wrong direction, to explain the post-2004 productivity slowdown. A critique of GDP that does not rescue the optimists. Free and digital goods
Is Growth Obsolete?
William Nordhaus & James Tobin (1972), NBER, Economic Research: Retrospect and Prospect
The first serious welfare-adjusted aggregate. Deducting regrettables and adding leisure and household work, the Measure of Economic Welfare grew 1.1% a year per head from 1929 to 1965 against 1.7% for NNP — slower, but still growth. Defensive and regrettable expenditure
For the Common Good (the Index of Sustainable Economic Welfare)
Herman Daly & John Cobb (1989), Beacon Press
The template for every later adjusted index: start from consumption, weight by inequality, add household work, deduct defensive expenditure and environmental depletion. Defensive and regrettable expenditure
Beyond GDP: Measuring and achieving global genuine progress
Kubiszewski, Costanza, Franco, Lawn, Talberth, Jackson & Aylmer (2013), Ecological Economics 93
Genuine Progress Indicator estimates for 17 countries covering 53% of world population: global GPI per capita peaks in 1978 and declines after, while GDP per capita more than triples. The single most-cited chart in the beyond-GDP literature. Defensive and regrettable expenditure
Environmental Accounting for Pollution in the United States Economy
Muller, Mendelsohn & Nordhaus (2011), American Economic Review 101(5)
Prices air-pollution damage industry by industry. In the worst sectors, gross external damages exceed measured value added — the activity subtracts from national wealth while adding to GDP. Unpriced damage
Housing Services in the National Economic Accounts
Mayerhauser & Reinsdorf (2007), BEA
Documents the scale of construction inside the accounts: imputations rose from 13.8% to 14.8% of GDP between 1996 and 2006, of which owner-occupied housing is the largest single item. Imputations and the financial sector
The Value of Risk: Measuring the Service Output of US Commercial Banks
Basu, Inklaar & Wang (2011), Economic Inquiry 49(1)
Compensation for bearing systematic risk is not a service. Removing it cuts measured US bank output by about 21% and GDP by about 0.3% over 1997–2007. Imputations and the financial sector
The Contribution of the Financial Sector: Miracle or Mirage?
Andrew Haldane (2010), Bank of England speech
Under the standard FISIM method the UK financial sector recorded exceptional measured growth in late 2008, while it was in fact being rescued. Risk-taking was being booked as output. Imputations and the financial sector
Distributional National Accounts: Methods and Estimates for the United States
Piketty, Saez & Zucman (2018), Quarterly Journal of Economics 133(2)
Decomposes national income growth by percentile, adding to 100% of the total. Average pre-tax income per adult up ~60% since 1980 while the bottom half stagnates; top-1%-to-bottom-50% ratio moves from 27 to 81. The mean is not the median
Report of the Commission on the Measurement of Economic Performance and Social Progress
Stiglitz, Sen & Fitoussi (2009), Commissioned by the French government
Twelve recommendations, the first of which is to shift emphasis from production to income and consumption, from the aggregate to the household, and from the mean to the distribution. The agenda most statistical offices have since been working through. The mean is not the median
Does Economic Growth Improve the Human Lot?
Richard Easterlin (1974), Nations and Households in Economic Growth
Within a country the rich report more happiness than the poor, but a country growing richer over time shows little gain in average reported happiness. The founding paradox of the subjective-wellbeing critique. Welfare has other arguments
Economic Growth and Subjective Well-Being: Reassessing the Easterlin Paradox
Betsey Stevenson & Justin Wolfers (2008), Brookings Papers on Economic Activity
With better data, wellbeing rises with income both within and between countries, with no satiation point in the data. The paradox is, at minimum, not established. Welfare has other arguments
Commodities and Capabilities
Amartya Sen (1985), North-Holland
Reframes the objective as what people are able to do and be, not what they consume. The philosophical basis of the Human Development Index and of every capability-based alternative since. Welfare has other arguments
Beyond GDP? Welfare across Countries and Time
Charles Jones & Peter Klenow (2016), American Economic Review 106(9)
A consumption-equivalent welfare measure combining consumption, leisure, mortality and inequality. Highly correlated with GDP per capita, but Western Europe is much closer to the US than GDP suggests — France reaches 92% of US welfare on 67% of its income per head — while many developing countries are further behind.The defence
Value and Capital (the definition of income)
John Hicks (1939), Oxford University Press
Income is the maximum that can be consumed in a period while leaving the consumer as well off at the end as at the beginning. The standard against which a gross measure necessarily fails. Gross, not net
On the Welfare Significance of National Product in a Dynamic Economy
Martin Weitzman (1976), Quarterly Journal of Economics 90(1)
Net national product, comprehensively defined, is the stationary equivalent of the future consumption stream. The theoretical result that makes net, not gross, the welfare-relevant aggregate. Gross, not net
Intergenerational Equity and the Investing of Rents from Exhaustible Resources
John Hartwick (1977), American Economic Review 67(5)
Consumption can stay constant while a resource is exhausted only if the rents are fully reinvested in reproducible capital. Gives a testable sustainability criterion that GDP cannot express. Gross, not net
Wasting Assets: Natural Resources in the National Income Accounts
Repetto, Magrath, Wells, Beer & Rossini (1989), World Resources Institute
Applied the netting critique to Indonesia, adjusting for petroleum, timber and soil depletion, and found that a substantial part of the country's rapid growth through the 1970s and early 1980s was the unsustainable cashing-in of natural wealth. The study that put resource accounting on the policy agenda. Gross, not net
Genuine Savings Rates in Developing Countries
Kirk Hamilton & Michael Clemens (1999), World Bank Economic Review 13(2)
Operationalised the critique as adjusted net savings — gross saving less capital consumption, less resource depletion and pollution damage, plus education spending. Published annually for most countries since. Gross, not net
The Economics of Biodiversity: The Dasgupta Review
Partha Dasgupta (2021), HM Treasury
Argues for inclusive wealth as the organising measure. Between 1992 and 2014 produced capital per person doubled while natural capital per person fell by nearly 40%, a trajectory GDP records as uninterrupted progress. Gross, not net
Toward a More Accurate Measure of the Cost of Living
Boskin, Dulberger, Gordon, Griliches & Jorgenson (1996), US Senate Finance Committee (the Boskin Commission)
Estimated that the CPI overstated inflation by about 1.1 percentage points a year — substitution, outlet, quality and new-goods bias — implying real growth had been correspondingly understated. Index numbers and quality change
Measurement Error in the Consumer Price Index: Where Do We Stand?
David Lebow & Jeremy Rudd (2003), Journal of Economic Literature 41(1)
Revisited Boskin after methodological reform: bias of about 0.9 percentage points a year, plausible range 0.3 to 1.4. The uncertainty band is wider than most of the growth differences that get argued about. Index numbers and quality change
The Soviet Indices of Industrial Production
Alexander Gerschenkron (1947), Review of Economics and Statistics 29(4)
Showed that the choice of base year alone can transform a measured growth rate, because early-period weights flatter goods whose relative prices later collapse. Index construction is not a technicality. Index numbers and quality change
Exact and Superlative Index Numbers
W. Erwin Diewert (1976), Journal of Econometrics 4(2)
Identifies the index formulas that are exact for flexible aggregator functions, giving statistical agencies a principled basis for chain-weighted measures. The constructive answer to the index-number critique. Index numbers and quality change
Understanding PPPs and PPP-based National Accounts
Angus Deaton & Alan Heston (2010), AEJ: Macroeconomics 2(4)
Documents how far internationally comparable GDP moves with defensible choices about baskets, weights and the treatment of government and housing. Cross-country levels are much less firm than their decimal places suggest. Comparability across countries
Poor Numbers: How We Are Misled by African Development Statistics
Morten Jerven (2013), Cornell University Press
Traces how weak base years and thin surveys produce GDP levels wrong by tens of per cent, then corrected in a single revision — Ghana +60% in 2010, Nigeria +89% in 2014 — with rankings and aid allocations built on the old numbers. Comparability across countries
Prosperity Without Growth
Tim Jackson (2009), UK Sustainable Development Commission
Argues that the growth target is structurally required by current macro institutions and structurally incompatible with ecological limits, so the institutions, not the statistic, are what has to change. GDP as a target
Doughnut Economics
Kate Raworth (2017), Random House
Replaces the single maximand with a band between a social floor and an ecological ceiling. The clearest statement of the dashboard alternative — and of its cost, which is that a band cannot rank two policies. GDP as a target
Troubling Tradeoffs in the Human Development Index
Martin Ravallion (2012), Journal of Development Economics 99(2)
The decisive objection to composite indices: the HDI's weights imply an implicit value of a life-year over 17,000 times higher in the richest country than the poorest. Aggregation choices smuggle in valuations nobody defended.The defence
Beyond GDP: The Quest for a Measure of Social Welfare
Marc Fleurbaey (2009), Journal of Economic Literature 47(4)
Surveys the alternatives and finds them individually defensible and jointly incoherent: each embeds a different social welfare function, and the disagreement between them is ethical rather than statistical.The defence
The Measure of Progress: Counting What Really Matters
Diane Coyle (2025), Princeton University Press
Argues the 1940s framework is now a distorting lens for an economy of intangibles, free goods and depleting natural capital, and proposes comprehensive wealth plus time-use accounting as the successor frame.The defence

The distribution of that bibliography is itself informative: 6 works on capital and sustainability, 4 on index numbers, and 2 on the political critique that gets the most airtime. The parts of the argument that are most settled technically are the parts that are least discussed publicly, and the reverse.

The replacements, and why none has replaced it

Ten serious candidates in fifty years. Each fixes a real defect; none has displaced the aggregate, and the reasons are consistent enough to be worth naming as a pattern.

MeasureFixesWhat it costsWhere it stands
Measure of Economic Welfare (MEW)
Nordhaus & Tobin (1972)
Prices are marginal social valuesRequires the statistician to decide which spending is regrettable, which is a value judgement dressed as an adjustment.Never institutionalised; the template for everything after it.
ISEW / Genuine Progress Indicator
Daly & Cobb; Cobb, Halstead & Rowe (1989)
Prices are marginal social valuesIts divergence from GDP is driven by a handful of large, contested deductions, and different national teams make different ones.Computed for ~20 countries by researchers; adopted officially by no national statistical office.
Human Development Index
Mahbub ul Haq & Amartya Sen, UNDP (1990)
One household stands in for allThe weights are arbitrary and imply valuations of life and schooling nobody would defend if stated explicitly (Ravallion, 2012).The most successful alternative: published annually, universally cited, and still reported alongside GDP rather than instead of it.
Adjusted Net Savings (genuine saving)
Hamilton & Clemens, World Bank (1999)
The capital consumed has been subtractedDepends on depletion and damage valuations that are themselves estimates, and says nothing about the level of welfare, only its direction.Published annually for most countries in the World Development Indicators. Rarely reported in the press.
OECD Better Life Index
OECD, post-Stiglitz-Sen-Fitoussi (2011)
One household stands in for allA dashboard refuses to weight its dimensions, which is intellectually honest and operationally useless for ranking two options.Live, eleven dimensions, user-weighted. No policy is set by it.
Inclusive Wealth Index
UNEP; Dasgupta; Arrow et al. (2012)
The capital consumed has been subtractedRequires shadow prices for natural and human capital that are model-dependent and revised heavily between editions.Periodic UNEP reports; given a major push by the 2021 Dasgupta Review and the 2025 SNA revision.
Consumption-equivalent welfare
Jones & Klenow (2016)
One household stands in for allNeeds an explicit utility function and a value of a statistical life, so it inherits every disagreement about those.The most rigorous academic alternative; used in research, not in policy. Its headline result is that GDP is a good proxy.
GDP-B
Brynjolfsson, Collis, Diewert, Eggers & Fox (2019)
Prices are marginal social valuesRests on stated-preference experiments, which are far less stable than transaction data and are gameable at scale.Experimental. The clearest candidate for measuring what the digital economy actually delivers.
Wellbeing budgets
New Zealand Treasury; Wales; Scotland; Iceland (2019)
The measure is not the objectiveChanges what is reported alongside the budget, not what the budget is constrained by.Real institutional change in a handful of small states; the fiscal rules remain GDP-denominated.
2025 System of National Accounts
UN Statistical Commission (2025)
The capital consumed has been subtractedIncremental by design: it revises the accounts around GDP rather than replacing the headline aggregate.Adopted March 2025. Depletion of subsoil assets becomes a production cost; data enters the asset boundary; implementation runs through the late 2020s.
  • A composite index hides its ethics in its weights. Any measure that adds life expectancy to income has implicitly priced a life-year. Ravallion's calculation on the Human Development Index — an implicit value of an extra year of life over seventeen thousand times higher in the richest country than the poorest — is the decisive objection, because nobody chose that number and nobody would defend it.
  • A dashboard refuses to weight, and so cannot decide. The OECD Better Life Index and the Doughnut are honest about the trade-offs they will not make, which is exactly why no finance ministry is constrained by one. Governments need a scalar to be graded on; refusing to supply one does not remove the demand.
  • Frequency and comparability are underrated. GDP is quarterly, available for almost every country, revised on a published schedule, and produced by a method two hundred agencies have agreed on. Every alternative is annual at best, covers fewer countries, and is produced by whoever happens to fund it.
  • The best alternative concludes that GDP is a good proxy. Jones and Klenow build the most theoretically disciplined welfare measure in the literature and report that it is highly correlated with GDP per person. The deviations are large enough to matter for specific comparisons — France reaches 92% of US welfare on 67% of its income per head — and systematic enough to be corrected, which leaves GDP as the right first approximation rather than the wrong measure.

What the critique is actually right about

Strip out what the accounts already concede and what the alternatives cannot deliver, and three claims survive intact.

  • Gross is the wrong concept and everyone knows it. Weitzman settled the theory in 1976: it is net product, comprehensively defined, that carries a welfare interpretation. An economy consuming its natural capital records the sale and never the loss — produced capital per head doubled between 1992 and 2014 while natural capital per head fell by nearly 40%. The 2025 System of National Accounts finally treats subsoil depletion as a cost of production, which is an admission as much as a reform.
  • A mean is not a measure of a population. Two countries with the same GDP per capita and different distributions are not in the same condition, and after 1980 the American distribution moved enough that the aggregate stopped describing the typical household at all. This one requires no new theory, only publishing the median next to the mean.
  • The target, not the measure, is the problem. A statistic used to grade governments stops behaving like an observation, which is Goodhart's Law applied to the most important number in public life. Nothing about the index causes this, and no replacement would be immune to it — which is why abolishing GDP is the one reform that would change nothing.

How to use the number anyway

For anyone making decisions on this data rather than writing about it, the literature converts into four working rules.

  • Use it as a production index, never as a welfare or a sustainability index. For capacity, demand, tax base and cyclical position it is excellent and there is nothing better. For whether a country is getting better to live in, or can keep doing what it is doing, it is the wrong instrument and there are purpose-built ones — median household income, adjusted net savings, inclusive wealth.
  • Treat cross-country levels as ranges, not numbers. Purchasing-power comparisons move materially with defensible methodological choices; for low-income countries the level itself can be wrong by tens of per cent until the next rebasing, as Ghana and Nigeria demonstrated at +60% and +89%.
  • Watch the composition, not the total. Growth delivered by depleting a stock, by defensive spending, or by shifting unpaid work into the market is arithmetically identical to growth delivered by producing more with the same inputs, and economically nothing like it.
  • Read a growth rate with its measurement error attached. Deflator bias alone has a plausible range of roughly a percentage point a year. Most of the growth differences that are argued over in public are inside the error bars of the instrument measuring them.

Appendix — method and limits

How to read this page. A review of published economics, organised around one framing device — the three conditions — that is this page's own. The framing is a teaching tool, not a claim to novelty: its value is that it sorts eleven arguments that are usually made in a single undifferentiated pile, and it makes clear which ones a statistical office could fix and which are objections to the concept of an aggregate.

On the numbers in the chart

Every range is a level, in per cent of GDP, and the ranges are not additive: some are output the accounts omit, some are spending they include, and the imputations are already inside the total. Ranges marked editorial are estimated across several published studies rather than taken from one; the others come from the source named in the readout. The two per-year magnitudes in the literature — deflator bias of roughly 0.9 to 1.1 percentage points a year, and the 0.05 to 0.11 points a year that free digital goods would add — are deliberately kept out of the chart, because putting a growth rate on an axis of levels is the kind of error the article is about.

What this page leaves out

  • The macroeconomics of whether growth can continue. The decoupling debate — whether output can rise while material throughput falls — is a question about the economy, not about the measure, and it is treated here only where it bears on how the measure is used.
  • The national-accounts mechanics. Chain-linking, seasonal adjustment, the treatment of government output at cost and the productivity measurement of public services each deserve their own treatment and are compressed here into a single condition.
  • GDP in wartime and command economies. Gerschenkron's index-number result appears here as a methodological point; its original subject — the systematic overstatement of Soviet industrial growth — is a subject in itself.
Related on this site: Goodhart's Law — why a measure used as a target stops measuring — Public spending in France and Taxes in France, both of which express everything as a share of the number examined here, and Where is the margin in the world? — the same measurement problem at the level of a firm's profit pool.