5.8CVMay 8Code
CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy AnalysisHaroon Wahab, Irfan Mehmood, Hassan Ugail
Wireless capsule endoscopy (WCE) enables non-invasive visual assessment of the small bowel, but its clinical utility is constrained by the large volume of frames generated per examination and the difficulty of recognising subtle abnormalities under highly variable imaging conditions. Existing learning-based approaches for WCE are predominantly vision-only, often confined to narrow pathology sets, and show limited transfer across datasets and centres. To address these limitations, this study introduces CapCLIP, a domain-specific vision-language representation learning framework for WCE. CapCLIP aligns capsule endoscopy frames with clinically grounded textual descriptions derived from standardised nomenclature and pathology-aware caption templates, thereby learning embeddings that are both semantically informed and transferable. The proposed framework is evaluated against relevant open-source vision and vision-language foundation models under strict zero-shot conditions using unseen WCE datasets. Evaluation covers three downstream tasks: K-nearest neighbour classification, CLIP-style image-text classification, and text-to-image retrieval. Across these settings, CapCLIP consistently outperforms the compared baselines, with particularly strong gains in zero-shot image-text classification and cross-modal retrieval on out-of-distribution datasets. The results indicate that language-guided representation learning can improve both generalisation and semantic interpretability in WCE analysis. These findings position CapCLIP as a step toward foundation models tailored to capsule endoscopy and support the use of language-grounded WCE analysis.
4.9CVJun 30
Symmetry-Structured Neural Completion of Islamic Geometric Patterns from Sparse Control GeometryHassan Ugail, Irfan Mehmood
Islamic geometric patterns are governed by exact rotational symmetry and strict construction rules. This paper treats these rules as formal geometric knowledge and embeds them in a neural completion framework, rather than leaving them to be learned statistically from data. Given sparse control geometry and a target symmetry order, the system completes the pattern as a vector graph by predicting edges and refinements of bounded curves over a candidate lattice whose edges are organised into rotational orbits under the cyclic group. Symmetry is enforced either by constraining predictions within these orbits or by projecting them onto them during inference. The orbit-tied variant provides a constructive guarantee: for any input and any orbit-level selection rule, it produces exact N-fold symmetry, preserves anchor points, and keeps all refinements within prescribed bounds. These properties are verified numerically. The study focuses on rotational symmetry, and all quantitative results are obtained from procedurally generated graphs inspired by Islamic geometric design rather than from a historical corpus. On clean inputs, enforcing exact validity produces no measurable loss in fidelity. When control geometry is missing, an unstructured decoder loses fidelity and breaks symmetry; retraining on corrupted inputs recovers much of the fidelity but not exact validity. Symmetry-structured inference, by contrast, keeps violations at zero throughout. The results show that augmentation and symmetry structure address distinct failure modes: augmentation improves fidelity under corruption, while symmetry structure guarantees validity. The framework therefore provides a knowledge-constrained, guarantee-backed approach to neural completion for scalable vector ornaments whose validity depends on exact geometric structure.
3.3MMOct 8, 2015
Ontology-based Secure Retrieval of Semantically Significant Visual ContentsKhan Muhammad, Irfan Mehmood, Mi Young Lee et al.
Image classification is an enthusiastic research field where large amount of image data is classified into various classes based on their visual contents. Researchers have presented various low-level features-based techniques for classifying images into different categories. However, efficient and effective classification and retrieval is still a challenging problem due to complex nature of visual contents. In addition, the traditional information retrieval techniques are vulnerable to security risks, making it easy for attackers to retrieve personal visual contents such as patients records and law enforcement agencies databases. Therefore, we propose a novel ontology-based framework using image steganography for secure image classification and information retrieval. The proposed framework uses domain-specific ontology for mapping the low-level image features to high-level concepts of ontologies which consequently results in efficient classification. Furthermore, the proposed method utilizes image steganography for hiding the image semantics as a secret message inside them, making the information retrieval process secure from third parties. The proposed framework minimizes the computational complexity of traditional techniques, increasing its suitability for secure and real-time visual contents retrieval from personalized image databases. Experimental results confirm the efficiency, effectiveness, and security of the proposed framework as compared with other state-of-the-art systems.
10.9IRFeb 25, 2015
Describing Colors, Textures and Shapes for Content Based Image Retrieval - A SurveyJamil Ahmad, Muhammad Sajjad, Irfan Mehmood et al.
Visual media has always been the most enjoyed way of communication. From the advent of television to the modern day hand held computers, we have witnessed the exponential growth of images around us. Undoubtedly it's a fact that they carry a lot of information in them which needs be utilized in an effective manner. Hence intense need has been felt to efficiently index and store large image collections for effective and on- demand retrieval. For this purpose low-level features extracted from the image contents like color, texture and shape has been used. Content based image retrieval systems employing these features has proven very successful. Image retrieval has promising applications in numerous fields and hence has motivated researchers all over the world. New and improved ways to represent visual content are being developed each day. Tremendous amount of research has been carried out in the last decade. In this paper we will present a detailed overview of some of the powerful color, texture and shape descriptors for content based image retrieval. A comparative analysis will also be carried out for providing an insight into outstanding challenges in this field.