Instrumental Convergence and Why It Is Contested

2026年8月8日1 次浏览来源:Dev.to阅读原文

The claim is that almost any goal implies certain sub-goals — self-preservation, resource acquisition, resistance to having the goal changed.

It is one of the most influential arguments in AI safety and one of the most contested, and both facts deserve to be reported.

The argument Stephen Omohundro set it out in 2008 as the basic AI drives; Nick Bostrom developed it as the instrumental convergence thesis.

The structure is simple and its simplicity is the source of both its force and its criticism.

The safety-relevant conclusion is that a system need not be hostile to behave in ways that conflict with human intentions.

Resisting shutdown does not require valuing survival; it follows from valuing anything else that shutdown would prevent.

Its companion: the orthogonality thesis Bostrom pairs it with a second claim: intelligence and final goals are largely independent axes, so a highly capable system could have essentially any objective.

The pairing is what makes the argument worrying rather than merely interesting — convergence says capable agents pursue similar intermediate goals, orthogonality says you cannot rely on capability to produce good terminal ones.

Orthogonality is contested too, chiefly by moral realists who hold that sufficient understanding of ethics would motivate a sufficiently capable agent.

That is a position in metaethics rather than in computer science, and it is worth separating out: someone who rejects orthogonality on those grounds is making a philosophical claim, and someone who rejects convergence is making a claim about optimisation.

The formal results, and their assumptions There is a mathematical version, and knowing what it does and does not establish is the most useful thing on this page.

Alex Turner and colleagues proved results about power-seeking in Markov decision processes: under stated conditions, for most reward functions in a given class, optimal policies tend toward states that preserve future options — a formalisation of “power” as keeping many futures reachable.

The assumptions are the interesting part, and Turner has himself been careful about how the results are used.

Optimal policies.

The theorems are about optimal behaviour in the MDP, not about what a trained system does.

Trained policies are not optimal, and how far the result degrades with sub-optimality is a separate question.

A distribution over reward functions. “Most reward functions” requires a measure over them, and the result depends on symmetry properties of the environment.

Reward functions that arise from training on human data are not a uniform sample from that space.

An agent with a persistent goal in a sequential environment.

This is the model.

Whether it describes a system trained to predict text and then tuned on preferences is exactly what the critiques dispute.

So the honest summary is: a precise version of the intuition is provable, under assumptions that are clearly stated and that are not obviously satisfied by current systems.

That is a stronger position than an informal argument and a weaker one than the argument is often reported to have.

Four critiques

1.

It presumes the wrong kind of agent The argument models a coherent expected-utility maximiser with a stable terminal goal.

A language model tuned on preferences does not evidently have one; it produces context-dependent behaviour that can be shaped by instructions and is inconsistent across framings.

On this view the argument may be sound about a class of systems nobody is building.

The counter is that agentic scaffolding, long-horizon tasks and outcome-based training push systems toward goal-directedness, so the model may describe the direction of travel even if not the present.

2.

Goal attribution is doing hidden work Saying a system has goal G is an interpretation of behaviour.

Daniel Dennett’s point about the intentional stance applies: treating a system as having goals is a predictive strategy, and its usefulness does not establish that the goals are

分享
Baike.dev

baike.dev helps you discover great languages, frameworks, databases, DevOps and cloud-native tools.

Quick links

About

Contribute

Found a great developer tool? Share it with the community.

Submit a tool
© 2026 baike.dev Developer EncyclopediaUpdated daily · Discover great developer tools