3.3HCJun 6, 2020
Multimodal Systems: Taxonomy, Methods, and ChallengesMuhammad Zeeshan Baig, Manolya Kavakli
Naturally, humans use multiple modalities to convey information. The modalities are processed both sequentially and in parallel for communication in the human brain, this changes when humans interact with computers. Empowering computers with the capability to process input multimodally is a major domain of investigation in Human-Computer Interaction (HCI). The advancement in technology (powerful mobile devices, advanced sensors, new ways of output, etc.) has opened up new gateways for researchers to design systems that allow multimodal interaction. It is a matter of time when the multimodal inputs will overtake the traditional ways of interactions. The paper provides an introduction to the domain of multimodal systems, explains a brief history, describes advantages of multimodal systems over unimodal systems, and discusses various modalities. The input modeling, fusion, and data collection were discussed. Finally, the challenges in the multimodal systems research were listed. The analysis of the literature showed that multimodal interface systems improve the task completion rate and reduce the errors compared to unimodal systems. The commonly used inputs for multimodal interaction are speech and gestures. In the case of multimodal inputs, late integration of input modalities is preferred by researchers because it allows easy update of modalities and corresponding vocabularies.
3.1HCNov 7, 2019
An Agent-Based Intelligent HCI Information System in Mixed RealityHamed Alqahtani, Charles Z. Liu, Manolya Kavakli-Thorne et al.
This paper presents a design of agent-based intelligent HCI (iHCI) system using collaborative information for MR to improve user experience and information security based on context-aware computing. In order to implement target awareness system, we propose the use of non-parameter stochastic adaptive learning and a kernel learning strategy for improving the adaptivity of the recognition. The proposed design involves the use of a context-aware computing strategy to recognize patterns for simulating human awareness and processing of stereo pattern analysis. It provides a flexible customization method for scene creation and manipulation. It also enables several types of awareness related to the interactive target, user-experience, system performance, confidentiality, and agent identification by applying several strategies, such as context pattern analysis, scalable learning, data-aware confidential computing.
0.9CVApr 12, 2019
An Introduction to Person Re-identification with Generative Adversarial NetworksHamed Alqahtani, Manolya Kavakli-Thorne, Charles Z. Liu
Person re-identification is a basic subject in the field of computer vision. The traditional methods have several limitations in solving the problems of person illumination like occlusion, pose variation and feature variation under complex background. Fortunately, deep learning paradigm opens new ways of the person re-identification research and becomes a hot spot in this field. Generative Adversarial Nets (GANs) in the past few years attracted lots of attention in solving these problems. This paper reviews the GAN based methods for person re-identification focuses on the related papers about different GAN based frameworks and discusses their advantages and disadvantages. Finally, it proposes the direction of future research, especially the prospect of person re-identification methods based on GANs.
2.1RODec 2, 2015
Continuous and Simultaneous Gesture and Posture Recognition for Commanding a Robotic Wheelchair; Towards Spotting the Signal PatternsAli Boyali, Naohisa Hashimoto, Manolya Kavakli
Spotting signal patterns with varying lengths has been still an open problem in the literature. In this study, we describe a signal pattern recognition approach for continuous and simultaneous classification of a tracked hand's posture and gestures and map them to steering commands for control of a robotic wheelchair. The developed methodology not only affords 100\% recognition accuracy on a streaming signal for continuous recognition, but also brings about a new perspective for building a training dictionary which eliminates human intervention to spot the gesture or postures on a training signal. In the training phase we employ a state of art subspace clustering method to find the most representative state samples. The recognition and training framework reveal boundaries of the patterns on the streaming signal with a successive decision tree structure intrinsically. We make use of the Collaborative ans Block Sparse Representation based classification methods for continuous gesture and posture recognition.