AICLCYHCJun 15

Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact

arXiv:2606.162069.2
Predicted impact top 71% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of LLM-based educational tools, this work highlights the need to evaluate tutoring systems on learning support rather than just task accuracy.

The paper shows that LLM tutoring benchmarks do not adequately distinguish between solving ability and pedagogical support, finding only a 0.421 correlation between the two across eight models. It proposes a diagnostic based on the gap between solving-oriented and pedagogy-oriented scores.

Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we study whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. We propose a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using public MathTutorBench leaderboard results, we show that these dimensions are only partially aligned: across eight publicly reported models, the correlation between solving and pedagogy composites is 0.421, and several models shift meaningfully in rank when evaluation moves from solving to pedagogy. We then analyze the public TutorBench sample and show that agency-relevant behaviors are explicitly encoded in benchmark rubrics, especially in active-learning settings that reward guiding questions, calibrated hints, and non-disclosive scaffolding. Together, these findings suggest that educational-impact evaluation should not treat task success as a sufficient proxy for learning support. We argue that public tutoring benchmarks can better support positive-impact evaluation by reporting solving-oriented and pedagogy-oriented scores separately and by making disclosure-sensitive, student-agency-preserving criteria more explicit.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes