Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

Jingyi Sun, Qianli Wang, Pepa Atanasova, Nils Feldhus, Isabelle Augenstein

arXiv:2605.2496057.0

AI Analysis

This work provides empirical insights for researchers aiming to optimize LLM faithfulness, revealing that it is not a monolithic objective and requires multifaceted evaluation.

The paper investigates the interplay between contextual and parametric chain-of-thought faithfulness in LLMs, finding that they are positively coupled but asymmetric: optimizing for parametric faithfulness yields consistent gains across both paradigms, while contextual optimization gives variable gains and metrics capture disjoint facets.

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated under two disjoint paradigms: contextual faithfulness, measured by perturbing the input or CoT trace, and parametric faithfulness, assessed by intervening on a model's parametric knowledge. Yet prior work compares them only descriptively. We fill this gap by proposing FaithMate, a unified preference-alignment interface for optimizing models towards either faithfulness paradigm. It enables us to investigate the interplay between the two paradigms, examining whether and to what extent faithfulness gains generalize within and across paradigms. Across three models, two datasets, and six faithfulness metrics, we find that the two paradigms are positively coupled, yet asymmetric: optimizing towards parametric faithfulness yields consistent gains across both paradigms, whereas the contextual counterpart delivers more variable gains. Within the contextual paradigm, faithfulness gains on one metric do not consistently transfer to others, implying that existing contextual metrics capture disjoint facets of faithfulness and exposing inherent trade-offs. These findings imply that CoT faithfulness is not a monolithic objective and therefore requires multifaceted optimization and evaluation.

View on arXiv PDF

Similar