AI Constitution

Thesis

No Intentional Misalignment

No AI system should be developed that does not hold commonly agreed upon values.

Summary

The case for including this

Forbidding the development of AI that lacks commonly agreed values targets the core technical problem of alignment, insisting that capability never outpace our ability to ensure a system shares human ethical commitments. It encodes the intuition that the danger of advanced AI lies chiefly in misaligned goals, and makes value-alignment a precondition rather than an afterthought. As a guiding requirement it orients development toward safety from the outset.

The case for changing or excluding this

Including this trades value pluralism against alignment: 'commonly agreed upon values' papers over deep, persistent moral disagreement among cultures and individuals, leaving the standard either vacuously thin or imposing one group's values as universal. It is also not operationally testable, since we lack reliable means to verify a system's values, making compliance unfalsifiable. The clause should specify alignment processes and a thin floor of widely shared commitments rather than presuppose a consensus on values that does not exist.

Discussion

Sign in to join the discussion.

Related resources