Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems

| Source: arXiv AI

Tags: text-to-image, cultural-bias, DALL-E-3, Imagen-4, bias-auditing, WORLDVIEW

A EMNLP 2026 paper identifies the silent prompt-revision layer in DALL-E-3, Imagen-4, and GPT-Image-1.5 as a previously undocumented causal source of cultural stereotyping—non-Western contexts are far more heavily rewritten and flattened than US English prompts across 15 languages.

Details

Commercial text-to-image systems silently rewrite user prompts before generating images, a step that is typically invisible to users and has been ignored by existing bias audits that only examine final outputs. WORLDVIEW is a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings designed to audit this revision layer specifically. The audit applies three-step analysis to the revision layer in DALL-E-3, Imagen-4, and GPT-Image-1.5: how heavily each cultural context is marked, whether it gets flattened into a narrow vocabulary, and whether that vocabulary is stereotypical. Key finding: the US context is the least-marked and least-modified. Non-Western and non-Anglophone contexts are marked far more heavily, flattened into narrow vocabularies regardless of topic diversity, and reduced to recognizable cultural stereotypes. Comparing images from original versus revised prompts on models without a revision layer confirms that the revision layer—not the underlying image generation model—is the causal driver of this stereotyping effect. The paper argues that bias audits must examine deployed systems as deployed, not models in isolation. Accepted to EMNLP 2026.