Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context
| Source: arXiv AI
Tags: time series forecasting, multimodal AI, benchmarks, foundation models, LLM evaluation
MUSE-Bench is a new benchmark for multimodal time series forecasting covering 14 datasets across 8 domains and 6 context types — finding that numerical foundation models still dominate, general-purpose LLMs are poor direct forecasters, and external context consistently helps context-aware models when properly aligned.
Details
Most time series forecasting benchmarks are numerical-only and cannot evaluate how models use contextual information like news, images, metadata, or calendar events that shape real-world dynamics. MUSE-Bench addresses this gap with a unified benchmark covering 14 datasets across 8 domains and 6 context types: metadata, events, holidays, news, images, and numerical covariates. The benchmark evaluates statistical, data-specific, foundation, multimodal, and general-purpose LLM forecasting methods under shared non-overlapping forecast windows and consistent metrics. Three key findings emerge: First, numerical time series foundation models dominate overall rankings. Aurora, the evaluated multimodal foundation model, trails leading numerical TSFMs but outperforms all data-specific models. Second, external context consistently improves the four evaluated context-aware models, but only when it is correctly aligned — temporally misaligned or incorrect context degrades performance. Third, general-purpose LLMs perform poorly as direct forecasters, and LLM-guided refinement does not yield consistent improvements. This benchmark is a useful reference for teams building forecasting systems and deciding whether to invest in multimodal context integration.