TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
| Source: arXiv AI
Tags: timeseries forecasting, multimodal LLMs, reinforcement learning, RLVR, healthcare AI
TimeThink trains timeseries multimodal LLMs using reinforcement learning with verifiable rewards on synthetically generated compositional question-answer pairs—achieving significant improvements over strong baselines on real-world benchmarks despite training only on synthetic data, including healthcare timeseries.
Details
Timeseries multimodal LLMs (TS-MLLMs) have begun applying LLM reasoning to temporal data, but existing approaches struggle with compositional questions combining multiple primitives (trend + seasonality + anomaly) and provide only implicit reasoning that is hard to verify in high-stakes settings like healthcare. Regmi et al. from Dartmouth College address this with TimeThink. The core insight: timeseries primitives are domain-independent and deterministically composable, enabling objective ground truth with reasoning traces—unlike real-world data where ground truth is often ambiguous. TimeThink's training pipeline has three parts: a synthetic data generator producing atomic and composite question-answer pairs with full reasoning traces; an RLVR (reinforcement learning with verifiable rewards) training strategy that rewards explicit reasoning over implicit; and a model that learns compositional logic rather than imitating template traces. Experiments show TimeThink—trained only on synthetic data—significantly outperforms strong baselines on both synthetic and real-world benchmarks, including healthcare timeseries applications. The RLVR approach is notable for not requiring labeled real-world data, which is scarce and expensive in clinical settings. Code is released alongside the paper, enabling practitioners to extend the generation framework to new domains.