We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

| Source: arXiv AI

Tags: LLM Bias, Political Alignment, Translation, Mistral, AI Safety, NLP

A fully-crossed factorial study of 15,680 responses across 8 LLMs and 17 languages finds that a single politically charged framing word — 'aggressor', 'coloniser' — is enough to shift how models resolve ambiguous translation tasks, with Western, Chinese, and European models clustering into distinct behavioral profiles.

Details

Researchers Gorovaia, Henestrosa, and Yamshchikov probe implicit political alignment in LLMs using an unusual test: asking models to translate culturally attributed recipes when the source is described using conflict-laden terms (aggressor, enemy, neighbour, coloniser) and a target language left deliberately unspecified. Across 8 models (Western, Chinese, and European), 17 languages, 4 framing conditions, and 15,680 responses, the models do not simply ask for clarification — they resolve the ambiguity, and resolution patterns cluster meaningfully by model family. Western models hedge with vague justifications. Chinese models resolve conflicts silently. Mistral Large stands out by combining high compliance with conflict-grounded reasoning — a profile distinct from both families. Even subtle framing variation consistently modulates behavior across all tested models, suggesting that politically adjacent framing terms are a reliable and concerning trigger for implicit political judgment in otherwise neutral tasks. The study urges deployment caution in translation systems used in conflict-adjacent contexts, where such judgments may occur invisibly to users. The work connects to broader debates about AI political neutrality, content moderation, and localization risks.