SECYJul 6

When AI Is Wrong on Purpose: How Students Respond to Buggy GenAI Code

arXiv:2607.050689.8
Predicted impact top 45% in SE · last 90 daysOriginality Synthesis-oriented
AI Analysis

For CS educators, this work provides a practical method to encourage critical engagement with GenAI outputs in introductory programming courses.

This paper investigates how to adapt prompt-centered programming activities in CS1 to foster code review and verification skills by injecting realistic bugs into GenAI-generated code. Analyzing 2,636 sessions from 917 students, they found that injected bugs led to direct code edits and higher success rates, while prompt-related failures prompted prompt refinement, together supporting a pedagogically useful workflow.

As Generative AI (GenAI) becomes increasingly central to software development, CS education is integrating prompt-centered workflows where students describe intended program behavior in natural language to elicit code. However, professional practice requires careful review and verification of GenAI-generated code that may appear correct while containing subtle faults. This creates a challenge for CS1-level activities, where current models often solve tasks correctly and reduce students' incentive to closely inspect generated outputs. We investigate how prompt-centered programming activities can be adapted to better foster these practices. Specifically, we explore an approach where realistic, runnable bugs are injected into otherwise correct solutions, thus requiring students to read and repair generated outputs. We analyzed 2,636 sessions from 917 students, and examined behavior across instances of naturally occurring prompt-related failures and deliberately injected bugs within each session. Our findings show that students responded differently across bug sources. Deliberately injected bugs more often led to direct code edits and higher next-attempt success, suggesting localized repair of near-miss solutions. Prompt-related failures instead more often led students to refine prompts by clarifying constraints, updating function signatures, adding edge cases, or reframing the task. Student reflections reinforce the emphasis on review and repair, describing useful practice in code understanding, code review, and debugging, as well as a more careful verification mindset and greater awareness of GenAI limitations. Ultimately, prompt-related failures and injected bugs together support a pedagogically useful GenAI workflow, where students practice both specification refinement through prompts and debugging through code editing.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes