Your Model’s MSE Is Lying to You

| Source: Towards Data Science

Tags: MSE, probabilistic forecasting, uncertainty quantification, time-series, Gaussian distribution, sensor data, heteroskedasticity

Two models with identical MSE can have ≈0% vs ≈40% probability of triggering the same alarm threshold — MSE captures mean prediction accuracy but misses conditional uncertainty entirely, making it an unreliable sole metric for any forecasting system where decisions hinge on risk quantification.

Details

MSE (Mean Squared Error) is the default metric for ranking forecasting models, but this Towards Data Science tutorial shows it can be actively misleading. Two models watching the same time-series signal — seismic readings, electrical load, bridge strain — can post identical MSE scores while having radically different uncertainty profiles. One model sees a tight Gaussian N(0.5, 0.01²) where the alarm threshold is 50 standard deviations away (probability ≈0%). The other sees N(0.5, 2.0²), where the same threshold is just 0.25 standard deviations away — roughly 40% of observations will cross it. The core problem: MSE optimizes for the mean prediction and penalizes errors symmetrically regardless of where they fall in the signal's variance landscape. When signal variance is heteroskedastic — tighter in some regimes, wider in others — MSE averages over this structure, hiding the exact information needed for threshold-based decisions like alarms and triggers. The article frames this as the case for probabilistic forecasting: replacing single point estimates with full predictive distributions so downstream consumers can compute exceedance probabilities and understand model confidence. This is part 1 of a series; part 2 will cover multi-step forecast rollouts. For practitioners building monitoring systems, anomaly detectors, or any model where decisions depend on thresholds, this is a concrete reminder to look beyond MSE and consider calibrated probabilistic outputs.