Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos

| Source: arXiv AI

Tags: egocentric video, social interaction, multimodal AI, computer vision, Ego4D, augmented reality

A multimodal model combining audio, visual, and speech-semantic cues achieves 75.5% frame-level F1 on online talk-to-me detection in egocentric video — reformulating what was a binary offline task into a richer online prediction over four speaking states.

Details

Egocentric video systems need to understand who is speaking to the camera wearer in real time — a capability required for AR glasses, social robots, and wearable assistants. Prior work treated this as a binary offline clip-level classification, which misses the complexity of real social interaction.\n\nDu et al. reformulate talk-to-me (TTM) detection as an online, frame-level prediction task with four classes: background, talking-to-me, talking-to-others, and self-talking. They introduce a new Online TTM Dataset with 406 egocentric clips, approximately 900,000 annotated frames, built by extending the Ego4D social interaction benchmark.\n\nTheir multimodal model integrates audio, visual, and speech-semantic cues, reaching 75.5% F1 on TTM versus strong single-modality and fusion baselines. The paper's systematic analysis shows how each non-TTM speaking state degrades TTM recognition differently — useful for building robust social AI systems.\n\nAccepted at ACM Multimedia 2026. The 900K-frame annotated dataset is the lasting contribution; it enables future work on fine-grained egocentric social interaction.