UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

| Source: arXiv AI

Tags: multimodal, image generation, evaluation, ICML 2026, MLLM, subject-driven generation, customization

UFO, accepted at ICML 2026, introduces a unified evaluation framework for multi-modal image generation that improves correlation with human judgment by 15.25% on average by scoring all conditions simultaneously rather than in isolation.

Details

Evaluating multi-modal image generation — where a model must match both text prompts and visual reference subjects — has lagged behind generation quality. Existing methods evaluate each condition (text alignment, subject fidelity) independently, producing scores that poorly reflect human judgment of the combined output. UFO (Unified Framework for Omni-condition alignment) addresses this with an Atomized Chain-of-Evaluation: conditions are first decomposed into fine-grained Atomic Evaluation Units (AEUs), categorized by modality relevance, then assessed via specific functional calls matched to each AEU type. The sequential chain captures interdependencies between conditions that isolated scoring misses. Across benchmarks, UFO achieves an average 15.25% improvement in correlation with human evaluation preferences over prior methods. The authors also release UFO-Bench, a dedicated benchmark for evaluating customization models under diverse textual and visual condition combinations. This work is particularly relevant to teams building or fine-tuning subject-driven image generation pipelines who need reliable automated evaluation metrics to replace or augment expensive human feedback.