Social Welfare Functions and the Problem of Interpersonal Comparison
A theory of individual welfare tells you how well one person's life is going. A theory of social welfare has to do something harder: combine many individual welfare levels into a single ranking of social states, so that policies, institutions, and outcomes can be compared. Welfare economics has spent most of a century working out how difficult that step actually is.
The Bergson-Samuelson social welfare function
The standard formal tool is the social welfare function, developed by Abram Bergson in 1938 and refined by Paul Samuelson. It is a mathematical function that takes each individual's welfare level as an input and produces a single social welfare ranking as an output. In its simplest utilitarian form, it just sums individual utilities. Other versions weight individuals unequally, or give extra weight to the worst-off (a prioritarian form), or evaluate only the minimum welfare level in the group (a maximin form, associated with the philosopher John Rawls).
The function is deliberately abstract about which weighting to use — that is a normative choice made outside the formalism. What the social welfare function framework contributes is a common structure for comparing different normative choices on equal footing.
The interpersonal comparison problem
The framework immediately runs into a foundational difficulty: to add, weight, or compare welfare across people, you need to be able to say things like "this gain to person A is larger than that loss to person B." That requires interpersonal comparability of welfare — the ability to compare welfare levels or changes across different people, not just within one person's own preferences.
Mid-twentieth-century economics largely retreated from this requirement, for a specific reason: individual utility, as economists use the term, is typically only ordinal (it ranks outcomes but assigns no meaningful numerical value) and defined only up to each person's own preferences. There is no obvious procedure for comparing one person's ordinal ranking to another's. If interpersonal comparison is ruled out, most social welfare functions cannot be constructed at all, since summing or weighting requires numbers that are actually comparable across people.
This is not merely a technical inconvenience. It means that any claim of the form "policy X makes society better off than policy Y" implicitly commits to some way of comparing gains and losses across different people — a commitment that is doing real normative work, whether or not it is stated explicitly.
Arrow's impossibility theorem
The difficulty deepened considerably with economist Kenneth Arrow's impossibility theorem, published in 1951. Arrow asked a more basic question: setting aside interpersonal comparison entirely, can individual preference rankings (not utility levels, just orderings) be aggregated into a social ranking using only ordinal information, while satisfying a small set of conditions that seem minimally fair?
The conditions were: the rule should work for any pattern of individual preferences (unrestricted domain); if everyone prefers A to B, society should prefer A to B (Pareto); the social ranking of A versus B should depend only on individual rankings of A versus B, not on preferences over unrelated alternatives (independence of irrelevant alternatives); and no single individual should always determine the outcome regardless of everyone else's preferences (non-dictatorship).
Arrow proved that no aggregation rule can satisfy all four conditions at once when there are at least three alternatives and unrestricted preferences. This is a genuine impossibility result, not an empirical claim about difficulty — it holds for any conceivable voting or aggregation procedure operating on ordinal preference information alone.
What the theorem does and does not show
Arrow's result is frequently over-generalized. It shows that ordinal preference aggregation, under those specific conditions, is impossible — not that social welfare judgments in general are impossible. Two escape routes have been extensively developed since.
The first is relaxing one of the four conditions. Dropping independence of irrelevant alternatives, for instance, opens the door to rules like the Borda count, which use more information than pairwise rankings.
The second, associated with Amartya Sen, is reintroducing interpersonal comparability deliberately rather than trying to avoid it. Sen showed that once some structured form of interpersonal comparison is allowed back in — even comparisons as weak as agreeing on how welfare gains are ranked across individuals, without agreeing on precise cardinal values — a wide range of reasonable social welfare functions become constructible again. This reframes Arrow's theorem less as a dead end and more as a demonstration that interpersonal comparability, whatever its philosophical difficulty, is doing indispensable work in any workable theory of social welfare.
Why this still matters for policy
Cost-benefit analysis, the standard tool for evaluating public policy, sidesteps the interpersonal comparison problem in a specific way: it uses willingness-to-pay, measured in money, as a common unit across individuals. This is a real and consequential choice, not a neutral default — it implicitly weights welfare changes by ability to pay, so a dollar of willingness-to-pay means something different to a low-income and a high-income person. Distributionally-weighted cost-benefit analysis attempts to correct for this, but requires exactly the kind of explicit interpersonal welfare weights that the underlying theory shows cannot be avoided by technical means alone.
The lesson from nearly a century of work on this problem is not that social welfare comparisons are impossible, but that every method for making them rests on a normative choice about interpersonal comparability that cannot be derived from individual preferences alone.