The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models
For researchers and developers of social robots, this work demonstrates a practical integration of VLMs into HRI with modest latency trade-offs, though it is an incremental application of existing models.
The paper explores integrating a Vision Language Model (Mistral AI) with a Pepper robot to enhance human-robot dialogue by adding visual context, finding that it increases response time moderately while improving situational understanding and complying with European data protection regulations.
Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.