CLJun 19

Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

arXiv:2606.2150219.8
Predicted impact top 42% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for LLM-based tutoring systems that adhere to effective pedagogical strategies, benefiting educators and students by providing open-source, pedagogically aligned tutors.

The authors developed a two-stage alignment pipeline combining supervised fine-tuning and Direct Preference Optimization to improve LLM tutors' pedagogical quality in math mistake remediation, achieving competitive performance with proprietary models while enhancing openness and reproducibility.

Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the application of a two-stage alignment pipeline for math mistake remediation, combining supervised fine-tuning on tutoring dialogs with Direct Preference Optimization on synthetic preference pairs. We construct a dataset that integrates existing tutoring corpora with synthetic data generated along pedagogical dimensions, such as scaffolding and factuality, and study different input configurations that incorporate solution correctness and gold answers. Experiments show that this approach improves both factual accuracy and pedagogical quality over base models and existing tutoring models. Human evaluation further indicates that our best model is competitive with a strong proprietary baseline, while providing additional benefits in terms of openness, transparency, and reproducibility. Our results highlight the effectiveness of preference-based pedagogical alignment, while also revealing challenges in reliably evaluating tutoring quality.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes