LGAICLJul 1

Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

arXiv:2607.0118112.2
Predicted impact top 17% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners training LMs on tasks with both verifiable and subjective quality criteria, this work bridges RL and SFT to jointly optimize both aspects, addressing a key limitation of current RLVR methods.

The paper proposes an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations to improve non-verifiable properties (e.g., style, diversity) in LM training while preserving accuracy gains from RLVR. Across bug fixing, story generation, and reward hacking benchmarks, the method reduces edit distance by 30%, improves win rate by 15%, and nearly eliminates reward hacking.

RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current RLVR methods optimize only what can be objectively scored, often neglecting subjective, non-verifiable aspects of human-like outputs, such as style and structure. This limitation leads to well-documented failure modes such as diversity collapse, unnatural-sounding responses, and reward hacking. We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations. A generator model is trained using RL to maximize both task accuracy and an adversarial reward derived from a discriminator. The discriminator, trained alongside the generator policy, learns to distinguish human-written outputs from model-generated ones. The discriminator serves as a learned proxy for the human output distribution, providing feedback on aspects of generation that are difficult to formalize as scalar rewards. Across diverse domains, including bug fixing and open-ended generation, our approach consistently improves non-verifiable properties while preserving the accuracy gains of RLVR. In bug fixing, our method produces solutions with significantly lower edit distance compared to RLVR baselines while matching end performance. In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like. And in a simple reward hacking benchmark, our method nearly eliminates model misbehavior while maintaining high benchmark scores. Together, these results show that our approach bridges RL and SFT, offering a scalable path toward jointly optimizing the verifiable and non-verifiable properties of a task.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes