LGCLJul 20

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

arXiv:2607.1811022.71 citations
Predicted impact top 1% in LG · last 90 daysOriginality Highly original
AI Analysis

For LLM post-training on non-verifiable tasks, EL provides a more effective learning signal than scalar rewards, improving generalization and reducing reward hacking.

The paper introduces Experiential Learning (EL), which replaces scalar rewards in RL for LLMs with rich textual feedback from an LLM-as-a-Coach. EL consistently outperforms rubric-based RL on open-ended tasks, generalizes better beyond training distribution, and mitigates reward hacking.

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes