Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

| Source: arXiv AI

Tags: multimodal AI, law enforcement, body-worn cameras, NLP, AI accountability, computer vision

Researchers propose a multimodal AI framework — combining video, audio, and NLP — to analyze Rochester Police Department body-worn camera footage, automatically classifying behavioral dynamics like de-escalation and disrespect to support law enforcement accountability reviews.

Details

Body-worn cameras generate enormous volumes of footage with limited capacity for systematic review. This paper proposes an interdisciplinary framework combining speaker separation, speech transcription, NLP, computer vision, and LLMs to automatically analyze police-civilian encounters from Rochester PD footage. The framework detects and classifies behavioral patterns including respect, disrespect, escalation, and de-escalation. LLMs generate structured, interpretable summaries of individual encounters. A custom evaluation pipeline assesses transcription quality and behavior classification accuracy in these high-stakes real-world scenarios. This is v4 of the paper (originally April 2025, updated September 2026). The authors frame the system as a tool for review, training, and accountability — not autonomous enforcement. No specific accuracy metrics for behavior classification are reported, limiting evaluation of real-world reliability. The societal stakes are high: automated behavioral analysis of police interactions raises civil liberties, bias, and due process concerns. The framing as a reviewer aid rather than decision-maker is appropriate, but deployment in real accountability processes would require significantly more validation.