A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification
| Source: arXiv AI
Tags: person-reidentification, computer-vision, multimodal, surveillance-AI, early-exit
A cost-aware early-exit cascade for multi-camera person re-identification resolves 60.7% of DukeMTMC and 68.4% of Market-1501 queries using only cheap visual features — avoiding expensive VLM semantic analysis on the majority of queries while maintaining competitive retrieval accuracy.
Details
Person re-identification (ReID) across multiple cameras is fundamental to video surveillance and smart city applications, but deploying accurate systems at scale requires balancing accuracy against inference cost. This paper proposes a generative AI-integrated multimodal framework with an explicit cost-aware decision rule: first inspect the top-k retrieval results from cheap global visual embeddings, and only escalate to VLM-generated semantic descriptions or optional facial embeddings for ambiguous queries. On Market-1501, 68.4% of queries are resolved at the cheap stage without semantic reasoning; on DukeMTMC-reID, 60.7% exit early. The full system integrates segmented person regions, fine-grained VLM-generated text descriptions, and optional face features with adaptive reliability weighting. Presented at the IEEE 2026 Conference on Responsible Artificial Intelligence (IRAI). The approach is practical for deployment budgets — average inference cost scales with the fraction of hard queries rather than total query volume. Privacy implications of persistent person tracking across cameras are not analyzed in the abstract. Full mAP and Rank-1 numbers for the complete cascade are not detailed in the abstract.