CLAIJun 28

To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregation

arXiv:2606.2948120.4
Predicted impact top 22% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For LLM reasoning, this tackles a critical bottleneck where models fabricate reasoning due to data overlap, offering a method to extract authentic reasoning skills.

HIPPO addresses the problem of LLMs exploiting shortcuts from Pre-RL data overlap by using hint-injected aggregation and a pairwise reward model to distinguish genuine reasoning from rationalization, achieving substantial improvements over baselines and generalizing to out-of-distribution tasks.

While reinforcement learning (RL) significantly enhances LLM reasoning, its efficacy is severely undermined by Pre-RL data overlap, where RL datasets overlap with pretraining or SFT corpora, causing models to exploit shortcuts by memorizing correct answers and fabricating post-hoc reasoning. To address this, we introduce HIPPO, a novel RL framework that integrates hint-injected aggregation with a tailored pairwise reward model. By utilizing hint injection to deliberately trigger overlap-induced behaviors, the resulting traces naturally serve as explicit anchors for pairwise comparison. This provides highly discriminable preference signals, enabling a lightweight judge model to reliably distinguish genuine reasoning deduction from shortcut-driven rationalization, while the pairwise formulation ensures stable and robust optimization compared to standard PRMs. Extensive experiments demonstrate that HIPPO yields substantial improvements over standard baselines and generalizes effectively to out-of-distribution general tasks, showing it extracts authentic, transferable reasoning skills rather than superficial shortcut patterns.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes