LLM-Derived Preference Judgments Are Not Self-Consistent
| Source: arXiv AI
Tags: LLM evaluation, preference learning, utility functions, AI agents, LLM reliability, decision-making
Across six LLMs tested on flight, apartment, and hotel preference scenarios, LLM-derived cardinal preference judgments show large, persistent inconsistencies — the same model gives contradictory willingness-to-pay estimates that cannot be reproduced by any single utility function.
Details
A growing number of agent systems use LLMs to estimate user preferences numerically — asking how much a person would pay for an item to build a utility function for decision-making. This paper tests whether those judgments are internally consistent: do willingness-to-pay differences between items match stated indifference payments? Across six LLMs and three domains (flights, apartments, hotels), the answer is clearly no. The researchers develop statistical tests and interpretable measures of how far observed LLM responses depart from the best-fitting self-consistent utility function. The inconsistencies are not small rounding errors — they are large and persistent across models and scenarios. This is problematic because the entire pipeline (LLM preference judgment → utility function → action selection) assumes approximate self-consistency. The study includes 16 pages of analysis across three real-world preference domains. No specific model names are listed in the abstract, but six LLMs were tested. Direct implications for teams building recommendation systems, preference elicitation agents, or any pipeline that translates LLM-rated preferences into decisions. It suggests that using raw LLM preference scores for utility estimation is unreliable and alternative approaches are needed.