Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification
For researchers in intelligent surveillance, this survey provides a structured overview of emerging multi-modal ReID paradigms, but it is a literature review without novel experimental results.
This survey reviews the transition from single-modal to multi-modal person re-identification (ReID), covering cross-modal tasks like visible-infrared, text-image, sketch-based, and NLOS ReID, and proposes a Transformer-based baseline for visible-infrared ReID. It summarizes datasets, challenges, and future directions.
Person re-identification (ReID) serves as a critical component in intelligent surveillance systems, aiming to match identities across disjoint camera networks. While traditional methods primarily rely on single-modal RGB imagery, they are often constrained by environmental challenges such as low illumination and occlusion. To overcome these limitations, the field is rapidly evolving toward cross-modal and multi-modal paradigms. This survey presents a comprehensive overview of this transition, systematically reviewing key cross-modal tasks including visible-infrared (VI-ReID), text-image (TI-ReID), sketch-based (Sketch-ReID), and the emerging Non-Line-of-Sight (NLOS) ReID, which extends perception beyond direct visibility. Furthermore, we examine tri-spectral and multi-modal fusion ReID, discussing how complementary information from diverse sensors enhances robustness. Beyond summarizing datasets, challenges, and methodologies, we propose a Transformer-based baseline framework for visible-infrared ReID, designed to effectively capture modality-invariant features. Finally, based on the current landscape, we outline several promising directions for future research.