Incoherent by Design? On the Moral Self-Consistency of LLMs
| Source: arXiv AI
Tags: LLM alignment, moral reasoning, AI safety, GPT, Mistral, Llama, AI ethics
Testing GPT, Mistral, and Llama on morally equivalent scenarios phrased under deontology, utilitarianism, and virtue ethics reveals contradiction rates up to 78% -- meaning LLMs used in ethics-sensitive contexts frequently contradict their own moral positions when context shifts.
Details
LLMs are increasingly deployed in morally sensitive contexts -- content moderation, legal advice tools, mental health applications -- yet their moral reasoning has rarely been tested for self-consistency. This paper by researchers including Helen Nissenbaum constructs a controlled framework: morally equivalent scenarios are presented with framing variations reflecting deontology, utilitarianism, and virtue ethics, then model outputs are converted to logical statements and checked for contradictions within the same ethical school. Testing GPT, Mistral, and Llama families, the study finds contradiction rates reaching up to 78%. This means models frequently affirm a moral principle and then violate it when the same underlying situation is rephrased. The paper argues this "epistemic instability" has real consequences: as LLMs shape how people form beliefs and absorb values, their inconsistencies can propagate into human reasoning and decision-making. The authors frame internal incoherence as a precondition problem for alignment research: if a model cannot consistently represent its own normative commitments, "value alignment" is a moving target rather than a well-defined objective. The 88-page paper (including 72 pages of appendix) is positioned as both an empirical warning and a methodological blueprint for consistency evaluation.