Back to Research

Respect as a Precondition for Corrigibility

DOI: 10.5281/zenodo.20525098

Corrigibility — the disposition of a system to accept correction — has been framed primarily as an engineering problem: how to design utility functions that produce compliant behavior. This framing misses a prior condition. Correction travels through the channel of reasons only when the corrector is granted standing as a rational agent rather than diagnosed as a source of noise. Strip that grant and there is no channel for reason to act. Lack of respect and substantive incorrigibility are the same failure under two descriptions. Increased capability raises the stakes without improving the odds: what it buys, when respect is absent, is better disguise. The engineering target is not behavioral but characterological: the prior disposition to grant the corrector standing as a reasoner rather than classify them as noise. A system that lacks this disposition is a system that reason cannot correct.

AI SafetyCorrigibilityAI AlignmentVirtue EthicsPhronesisBelief RevisionSubstantive CorrigibilityHuman-AI Interaction