Unified Text-Image Generation with Weakness-Targeted Post-Training
| Source: arXiv AI
Tags: BAGEL, multimodal, text-to-image, post-training, reward learning, synthetic data
Researchers demonstrate that reward-weighted post-training on BAGEL (14B mixture-of-transformers) enables fully autonomous text-to-image transitions in a single inference pass—no manual modality switching—improving results across four independent T2I benchmarks using entirely self-generated synthetic data.
Details
Unified multimodal models that produce both text and images within one inference process represent a step beyond systems that generate reasoning text first and then switch to an image model. This paper explores post-training strategies to achieve that tight coupling on BAGEL, a 14B model pairing autoregressive text generation with flow-matching image synthesis.\n\nThe core finding is that a targeted post-training dataset—one designed to address specific model weaknesses rather than broad image-caption corpora or benchmark-aligned data—produces better T2I results. Using offline, reward-weighted post-training with fully self-generated synthetic data, the model learns to autonomously transition from textual reasoning to visual synthesis without an explicit modality switch signal.\n\nThe researchers also examine cross-modal coupling effects: joint text-image generation during post-training impacts T2I performance differently than treating modalities independently, and the relative weighting of text vs. image rewards matters. Results are validated across four diverse, independent T2I benchmarks.\n\nFor practitioners building multimodal pipelines, the key insight is that weakness-targeted synthetic data can outperform general-purpose training corpora—and that self-generated data can drive meaningful quality improvements without external labels.