A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
This work addresses the gap in on-device applications for screen understanding and multilingual support, but it appears incremental as it applies existing large language models to new tasks.
The authors tackled the problem of enabling new on-device NLP applications by proposing three experiences: visual question answering, automated form filling, and smart replies for linguistic code-switching, aiming to bridge research with real-world impact.
Recent improvement in large language models, open doors for certain new experiences for on-device applications which were not possible before. In this work, we propose 3 such new experiences in 2 categories. First we discuss experiences which can be powered in screen understanding i.e. understanding whats on user screen namely - (1) visual question answering, and (2) automated form filling based on previous screen. The second category of experience which can be extended are smart replies to support for multilingual speakers with code-switching. Code-switching occurs when a speaker alternates between two or more languages. To the best of our knowledge, this is first such work to propose these tasks and solutions to each of them, to bridge the gap between latest research and real world impact of the research in on-device applications.