SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
| Source: arXiv AI
Tags: video highlight detection, information bottleneck, computer vision, temporal modeling, Gromov-Wasserstein
SGWIB introduces a structure-aware information bottleneck for video highlight detection that preserves temporal relationships between segments using a Sliced Gromov-Monge Gap regularizer, improving Kendall's tau by up to 0.031 over prior single-modal methods on MrHiSum.
Details
Video highlight detection requires identifying the most informative segments in a video, which demands understanding both what each segment shows and how segments relate to each other temporally. Standard information bottleneck approaches compress representations but ignore inter-segment relationships.\n\nSGWIB addresses this with the Sliced Gromov-Monge Gap (SGMG), a structure-aware regularizer that measures relational distortion introduced by the bottleneck mapping. It also adds a contextual disentanglement module that separates highlight-relevant features from sports-specific context, reducing dataset-specific bias.\n\nOn MrHiSum and MoSu datasets, SGWIB achieves the best Kendall's tau, Spearman's rho, mAP@50, and mAP@30 among compared single-modal methods. Gains of 0.031 on both rank correlation metrics are modest but consistent. The paper does not report multimodal results, suggesting room for further improvement.