Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

| Source: Apple ML Research

Tags: Apple ML Research, video reasoning, Visual CoT, multimodal, IVT, inference efficiency, post-training

Apple ML Research introduces Internalized Visual Thinking (IVT), a post-training method that bakes visual foresight into model weights during training — eliminating inference-time image generation and cutting latency by over 5× versus Visual Chain-of-Thought.

Details

Apple ML Research published IVT (Internalized Visual Thinking), a post-training framework for multimodal models that internalizes visual reasoning into model weights rather than requiring explicit image generation at inference. Visual CoT generates intermediate future-frame images as reasoning steps, which is accurate but slow — IVT eliminates this step entirely. IVT trains models on unlabeled video by jointly predicting the target text answer and latent representations of future frames. This forces the model to encode motion, object transitions, and spatial relationships into its weights. At inference, the model answers directly without generating or re-encoding any future frames, preserving the full efficiency of standard text generation. Across all six evaluation benchmarks tested, IVT outperforms text-only post-training baselines. Compared with Visual CoT, it achieves comparable or better accuracy with more than 5× reduction in end-to-end latency. The work runs controlled studies across multiple design dimensions: target representations, decoder architectures, prediction horizons, data mixtures, and training curricula. The research challenges a core assumption: that pixel-space future generation is necessary for proactive video reasoning. The practical implications are significant for real-time applications — robotics, video understanding, interactive media — where Visual CoT's overhead is prohibitive.