Predictive audio representations for early detection and tracking of hidden dynamic objects

| Source: arXiv AI

Tags: autonomous vehicles, audio perception, JEPA, occlusion detection, self-supervised learning

A two-stage audio pipeline using JEPA self-supervised pre-training on raw multichannel waveforms simultaneously estimates the count, type, and direction of occluded road vehicles — handling multi-agent Non-Line-Of-Sight scenarios that prior systems could not.

Details

Sound travels around obstacles in ways that cameras cannot, making audio a natural complement to optical sensors for detecting hidden traffic agents. Prior work in this area was limited to single-vehicle scenarios and handled only type classification or direction estimation separately — not both simultaneously. This paper introduces a two-stage pipeline. The first stage pre-trains an encoder via JEPA (Joint-Embedding Predictive Architecture) applied directly to multichannel raw waveforms, learning to predict latent representations of future audio segments from past context without labels. The second stage fine-tunes with a bidirectional LSTM using three classification heads for vehicle count, type, and direction of arrival simultaneously. No suitable public NLOS dataset with multiple simultaneous vehicles existed, so the authors collected one. The system outperforms prior art and shows reasonable transfer to an unseen driving environment — suggesting the JEPA-learned representations generalize beyond the training domain. The approach is niche but technically clean: applying JEPA to audio is a genuine methodological contribution, and the multi-task framing reduces the need for separate specialized detectors. Practical deployment in autonomous vehicles would require integration with existing sensor stacks.