How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts
For researchers and practitioners in AI-assisted language education, this work highlights the need for learner-centered evaluation frameworks, as expert ratings alone may misrepresent feedback effectiveness.
This study evaluated AI-generated written corrective feedback (WCF) in EFL writing by collecting over 20,000 drafts from nearly 2,000 students. Results showed low alignment between expert teacher ratings and student feedback, indicating that traditional expert evaluation alone may not capture WCF's usability from the learner's perspective.
This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.