SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation
For sign language translation researchers, this work addresses the challenge of cross-lingual transfer without gloss annotations, enabling scalable and effective multilingual SLT.
SIGNET achieves state-of-the-art sign language translation across four benchmarks (How2Sign, Phoenix14T, CSL-Daily, MeineDGS) by transferring motion-level knowledge from multiple pretrained backbones via an attention-based hand-prior aggregation mechanism, also surpassing prior methods on WLASL for recognition.
Sign language translation (SLT) remains challenging due to its high spatio-temporal complexity, long sequences, and the need to model multiple articulators without relying on gloss annotations. Existing approaches are typically tailored to individual datasets or languages and struggle to scale, while overlooking the relationships between sign languages that could inform more effective cross-lingual transfer. We present \textbf{SIGNET}, a framework that enables motion-level knowledge transfer for cross-language sign language translation. Our key insight is that, although sign languages differ in grammar and lexicon, pretrained models capture motion-level visual patterns that can be reused across datasets and languages. \textbf{SIGNET} integrates multiple pretrained sign language backbones through an attention-based, hand-prior aggregation mechanism that guides a gated fusion network in dynamically selecting the most relevant experts. Comprehensive experiments on four benchmarks (How2Sign, Phoenix14T, CSL-Daily, and MeineDGS) demonstrate state-of-the-art translation performance, and \textbf{SIGNET} also surpasses prior methods on WLASL for sign language recognition.