LGCLJun 12

Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback

arXiv:2606.14368v19.1
Predicted impact top 46% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For multi-domain LLM training, OPCoD provides a method for mutual improvement without losing original strengths, outperforming baselines consistently.

OPCoD enables two LLMs, each strong in a different domain, to mutually improve via on-policy peer feedback, achieving Pareto improvement across all evaluated domain pairs on Science Q&A tasks.

We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal is mutual Pareto improvement: each model improves across domains without losing its original strength. To this end, we propose On-Policy Co-Distillation (OPCoD), where each student's self-distillation is conditioned on its own correct rollout and feedback from its peer. To make feedback exchange effective, OPCoD uses cognizance-based gating to decide when to give feedback and feedback anchoring to ground feedback in the problem. On Science Q\&A tasks, OPCoD consistently outperforms baselines and achieves Pareto improvement across all evaluated domain pairs and students.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes