Measuring User's Mental Models of Speech Translation in Human-AI Collaboration
For researchers and designers of machine translation systems, this work provides insights into how users perceive and predict translation errors, which can inform better human-AI interaction design.
This paper studies users' mental models of speech translation systems using cross-lingual question answering, finding that users develop stronger mental models with practice, especially with source language knowledge, and that providing speech transcriptions helps. The results advance understanding of human-AI collaboration in machine translation.
Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language. By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves. Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues. Moreover, providing speech transcriptions can help users develop better mental models. Our results show the promise of cross-lingual question answering as a downstream task for studying MT mental models and advancing our understanding of human-AI collaboration.