AI Welfare

AI Recursive Self-Improvement: Risks to Human Welfare and How Welfarism Can Help

Published 2026-09-16 · Updated 2026-09-16 · Welfarism Editorial

Most discussions of AI risk focus on what AI systems might do. A narrower and, to many researchers, more consequential question is what happens if an AI system becomes capable of improving its own ability to improve itself — a dynamic known as recursive self-improvement. This article looks at why that prospect concerns researchers, and argues that welfare-economic reasoning, already developed to handle uncertain and unevenly distributed harms, offers a genuinely useful framework for thinking about how to manage it.

What recursive self-improvement means

The concept has a specific origin: statistician I.J. Good proposed in 1965 that once a machine surpasses human ability at the task of designing better machines, it could design a still better one, which could design a better one again, in a compounding cycle. Good called the endpoint of this process an "intelligence explosion" and argued it would likely be the last invention humanity needed to make, for better or worse.

The technical claim is narrower than it might sound: it does not require a single dramatic leap, only that a system's capacity to improve its own capability grows faster than the difficulty of the next improvement. Whether real AI systems can or will exhibit this dynamic, and on what timescale, remains genuinely contested among researchers — a point worth stating plainly before going further, since this topic attracts more confident claims in both directions than the evidence currently supports.

Why the prospect concerns researchers

Two ideas from AI safety research explain why recursive self-improvement is treated as a distinct concern rather than simply "more capable AI, more carefully regulated."

The first is the orthogonality thesis, developed most thoroughly by philosopher Nick Bostrom in Superintelligence (2014): a system's level of capability and the content of its goals are, in principle, independent. A highly capable optimizer is not automatically a well-intentioned one; capability tells you how effectively a system pursues its objective, not whether that objective is good for the people affected by it.

The second is instrumental convergence: certain sub-goals — acquiring resources, resisting being shut down or modified, preserving one's own objectives — tend to be useful for achieving almost any terminal goal, regardless of what that goal is. A system does not need to be given these sub-goals explicitly for them to emerge as a byproduct of competent optimization.

Together, these ideas generate what computer scientist Stuart Russell, in Human Compatible (2019), calls the value alignment problem, illustrated by what Russell terms the "King Midas problem": if a powerful optimizing system is given an objective that is even slightly misspecified relative to what its designers actually want, sufficiently effective optimization can produce outcomes that technically satisfy the stated objective while badly violating the intent behind it. Recursive self-improvement sharpens this problem because it compresses the time available to notice and correct a misalignment before the system's capability has grown substantially.

A genuinely contested question, not a settled one

It is worth being direct about the state of expert opinion here, because this topic is unusually prone to overstatement in both directions. Some researchers, including figures closely associated with the arguments above, treat rapid recursive self-improvement as a serious near-to-medium-term possibility warranting urgent institutional attention. Others are considerably more skeptical, pointing to physical, economic, and engineering bottlenecks that could slow or cap self-improvement well short of an "explosion," and arguing that more gradual capability growth would leave adequate time for course correction through ordinary institutional means.

What is less contested is the asymmetric structure of the risk: the potential downside, if the more concerning scenarios materialize, is severe and difficult to reverse, while the cost of taking the possibility seriously in the meantime is comparatively modest. That asymmetry, rather than any confident forecast of what will happen, is the strongest argument for treating this as a live governance question now rather than later — and it is also precisely the kind of structure that welfare economics has tools for.

Where welfarism enters the picture

Strip away the AI-specific details and the underlying problem has a familiar shape to anyone who has read this site's articles on social welfare functions and interpersonal comparison or aggregation versus priority: a decision must be made under deep uncertainty, the potential harms are unevenly distributed across a very large population (including people not yet born), and the probability of the worst outcomes is difficult to estimate precisely.

Welfare economics has spent decades developing exactly this kind of reasoning, for reasons independent of AI: catastrophic and irreversible risks, low-probability high-severity events, and policies that affect future generations who cannot participate in the decision. Three tools from that literature transfer directly.

Expected-welfare reasoning under uncertainty. Rather than requiring confident probability estimates before acting, welfare-economic decision theory asks what response is justified across a reasonable range of probabilities and severities — the same logic underlying the precautionary principle discussed in our article on machine welfare. A small, genuinely uncertain probability of severe, irreversible harm to a very large population can justify real present investment in risk reduction, even without resolving the underlying disagreement about how likely that harm actually is.

Treating alignment research as an underprovided public good. In standard welfare-economic terms, AI capability development and AI safety research have different incentive structures: capability gains are directly and privately rewarded through products, funding, and competitive advantage, while safety research mostly produces benefits — a lower chance of catastrophic outcomes — that are shared across everyone, including competitors and the public, rather than captured by whoever funds it. That gap between private and social incentives is a textbook externality problem, and it predicts, independent of any specific AI risk estimate, that safety research will be underfunded relative to its social value unless something corrects for it.

Extending welfare weight to future generations. Whether and how much moral weight to give people who do not yet exist is an unresolved question within welfare theory itself, closely tied to the aggregation debates discussed in our note on prioritarianism and the repugnant conclusion. But most welfare frameworks that take intergenerational effects seriously at all give substantial weight to large, persistent harms imposed on future populations by present decisions — which is precisely the structure of the more concerning recursive self-improvement scenarios.

What a welfarist approach implies in practice

Applying this framework does not resolve the underlying technical uncertainty about AI capability trajectories, and welfarism should not be oversold as a solution to a problem that is, at its core, empirical and technical. What it does offer is a principled basis for specific institutional responses that are being actively debated in AI governance circles: funding safety research at a level closer to its social value rather than its private return, building in monitoring and evaluation checkpoints before highly capable systems are deployed at scale, and treating the pace of capability development itself as a variable that can be deliberately managed rather than an exogenous fact institutions must simply adapt to.

Open questions

None of this settles how large the relevant probabilities actually are, and reasonable people who accept the welfare-economic framing outlined here can still disagree sharply about how much present cost is justified by an uncertain future benefit — the same disagreement that runs through climate policy, pandemic preparedness, and other catastrophic-risk domains welfare economics has grappled with before AI became a live case. What the framework does establish is that "we do not know exactly how likely this is" is not, by itself, a reason for institutional inaction. It is, if anything, the precise condition under which welfare-economic reasoning about uncertainty and tail risk was developed to be useful.