"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?
This addresses the problem of understanding multimodal communication transfer in AI for applications like deception detection, but it is incremental as it builds on existing LLM capabilities.
The paper investigates whether speech+text LLMs and conversation-trained LLMs can transfer multimodal skills to detect covert deceptive communication, finding that both types have an advantage over unimodal LLMs without special prompting.
Human communication is a multifaceted and multimodal skill. Communication requires an understanding of both the surface-level textual content and the connotative intent of a piece of communication. In humans, learning to go beyond the surface level starts by learning communicative intent in speech. Once humans acquire these skills in spoken communication, they transfer those skills to written communication. In this paper, we assess the ability of speech+text models and text models trained with special emphasis on human-to-human conversations to make this multimodal transfer of skill. We specifically test these models on their ability to detect covert deceptive communication. We find that with no special prompting speech+text LLMs have an advantage over unimodal LLMs in performing this task. Likewise, we find that human-to-human conversation-trained LLMs are also advantaged in this skill.