HCJun 14

Are LLM-based Chatbots Good Enough to Support Computer Science Students in Multiple-Choice Exercises?

arXiv:2606.159197.8
Predicted impact top 39% in HC · last 90 daysOriginality Synthesis-oriented
AI Analysis

For computer science educators, this work shows that while advanced LLMs can answer MCQs accurately, their use as learning aids may not enhance student outcomes.

The study evaluated LLM-based chatbots on university-level multiple-choice questions and found that GPT-4o and GPT-5 achieved the best performance, but providing ChatGPT answers with explanations did not improve student performance.

Chatbots based on large language models (LLMs) are increasingly adopted for information retrieval, text generation, and writing assistance. In educational settings, their use is also rapidly increasing. Students leverage these systems to complete tasks, access information, and support learning. However, the role of LLM-based chatbots in supporting learning and assessment in university-level computer science education is still underexplored. To address this gap, we investigate the performance of several LLM-based chatbots in solving multiple-choice questions (MCQs) at the university level and evaluate their capabilities to assist student learning. We developed 70 MCQs for a university lecture on interactive visual data analysis and evaluated the chatbots' performance using different prompt designs. We further compared the results with students' performance. Finally, we conducted a user study in two lectures (interactive visual data analysis, computer vision) to investigate how chatbot-generated answers and explanations affect students' performance. The chatbot performance showed significant differences between smaller models and GPT-4o and GPT-5 models, which achieved the best results. The results of the user study show that presenting ChatGPT answers together with an explanation does not improve students' performance in general.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes