HC AI LGFeb 12, 2024

Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking

Anjali Khurana, Hari Subramonyam, Parmit K Chilana

arXiv:2402.08030v125.183 citationsh-index: 24IUI

Originality Synthesis-oriented

AI Analysis

This highlights a critical usability problem for users relying on LLM assistants in software domains, showing that current prompt-based interactions are ineffective and may mislead non-experts.

The study investigated the effectiveness of LLM-based assistants for software help-seeking, finding that despite optimizations like SoftAIBot, users struggled to understand prompt-response relationships and often followed incorrect suggestions, leading to low task completion rates with no significant improvement from prompt guidelines or domain context.

Large Language Model (LLM) assistants, such as ChatGPT, have emerged as potential alternatives to search methods for helping users navigate complex, feature-rich software. LLMs use vast training data from domain-specific texts, software manuals, and code repositories to mimic human-like interactions, offering tailored assistance, including step-by-step instructions. In this work, we investigated LLM-generated software guidance through a within-subject experiment with 16 participants and follow-up interviews. We compared a baseline LLM assistant with an LLM optimized for particular software contexts, SoftAIBot, which also offered guidelines for constructing appropriate prompts. We assessed task completion, perceived accuracy, relevance, and trust. Surprisingly, although SoftAIBot outperformed the baseline LLM, our results revealed no significant difference in LLM usage and user perceptions with or without prompt guidelines and the integration of domain context. Most users struggled to understand how the prompt's text related to the LLM's responses and often followed the LLM's suggestions verbatim, even if they were incorrect. This resulted in difficulties when using the LLM's advice for software tasks, leading to low task completion rates. Our detailed analysis also revealed that users remained unaware of inaccuracies in the LLM's responses, indicating a gap between their lack of software expertise and their ability to evaluate the LLM's assistance. With the growing push for designing domain-specific LLM assistants, we emphasize the importance of incorporating explainable, context-aware cues into LLMs to help users understand prompt-based interactions, identify biases, and maximize the utility of LLM assistants.

View on arXiv PDF

Similar