CVJun 5

From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards

arXiv:2606.069666.4
Predicted impact top 58% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses cross-domain shift in ID card PAD, a privacy-sensitive problem, but the results are incremental as the model fails in zero-shot settings and relies on existing methods.

The authors propose a compact multimodal model combining visual and textual data for Presentation Attack Detection on ID cards, finding that while it generalizes well after fine-tuning, it fails in zero-shot settings, highlighting the need for real-world data and more realistic synthetic datasets.

Cross-domain shifts challenge Presentation Attack Detection (PAD) on ID Cards, given the restricted data available due to privacy concerns. This work proposes a compact multimodal model, based on new generative and discriminative blocks, which combines visual and textual data for PAD on genuine and synthetic ID images. While multimodal models exhibit strong generalisation after supervised fine-tuning, they fail in zero-shot settings. Our findings underscore that model capacity and real-world data are essential for reliable PAD, while existing synthetic datasets may not reflect real-world challenges. We argue for a re-evaluation of synthetic data as a benchmark and emphasise the need for more realistic, diverse datasets to advance PAD research.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes