STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
| Source: Apple ML Research
Tags: STARFlow2, Apple, multimodal, normalizing flows, image generation, unified models, Pretzel architecture, TARFlow
Apple's STARFlow2 achieves unified text-image generation using autoregressive normalizing flows — architecturally identical to causal Transformers — eliminating the structural mismatch between LLM text generation and diffusion-based image synthesis in a single causal forward pass.
Details
Unified multimodal models that generate both text and images have been structurally fragmented: combining causal LLM generation with iterative diffusion-based image denoising creates architectural asymmetry, different inference modes, and complications for KV-cache sharing. STARFlow2 resolves this by recognizing that autoregressive normalizing flows are mathematically equivalent to causal Transformers — same causal mask, same KV-cache mechanism, same left-to-right generation. The model is built on what Apple calls the Pretzel architecture: a frozen pretrained VLM stream (for multimodal understanding) vertically interleaved with a TARFlow stream via residual skip connections, both operating under the same causal mask. This preserves the VLM's existing understanding capabilities while enabling continuous image generation — avoiding the quality degradation that typically occurs when adapting VLMs for generation tasks. A deep-shallow flow design and unified FAE (Feature Auto-Encoder) latent space allow both text and visual outputs to enter the KV-cache directly without re-encoding, enabling cache-friendly interleaved generation. Apple reports strong performance across both image generation and multimodal understanding benchmarks. This is a research publication, not yet peer-reviewed through an external venue. If the causal unification holds under independent evaluation, the approach could influence multimodal model architectures more broadly — eliminating the need for different inference engines per modality.