10.1CLMar 17Code
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM SycophancyZhaoxin Feng, Zheng Chen, Jianfei Ma et al.
Alignment techniques often inadvertently induce sycophancy in LLMs. While prior studies studied this behaviour in direct-answer settings, the role of Chain-of-Thought (CoT) reasoning remains under-explored: does it serve as a logical constraint that mitigates sycophancy, or a tool for post-hoc rationalization that masks it? We evaluate a range of models across objective and subjective tasks to investigate the issue. Results show that reasoning generally reduces sycophancy in final decisions but also masks sycophancy in some samples, where models construct deceptive justifications through logical inconsistencies, calculation errors, and one-sided arguments etc. Furthermore, LLMs are more prone to sycophancy in subjective tasks and under authority-bias. Our mechanistic analysis on three open-source models reveals that the tendency of sycophancy is dynamic during the reasoning process rather than being pre-determined at the input stage.
5.3LGJan 27, 2023
Generalized Munchausen Reinforcement Learning using Tsallis KL DivergenceLingwei Zhu, Zheng Chen, Matthew Schlegel et al.
Many policy optimization approaches in reinforcement learning incorporate a Kullback-Leilbler (KL) divergence to the previous policy, to prevent the policy from changing too quickly. This idea was initially proposed in a seminal paper on Conservative Policy Iteration, with approximations given by algorithms like TRPO and Munchausen Value Iteration (MVI). We continue this line of work by investigating a generalized KL divergence -- called the Tsallis KL divergence -- which use the $q$-logarithm in the definition. The approach is a strict generalization, as $q = 1$ corresponds to the standard KL divergence; $q > 1$ provides a range of new options. We characterize the types of policies learned under the Tsallis KL, and motivate when $q >1$ could be beneficial. To obtain a practical algorithm that incorporates Tsallis KL regularization, we extend MVI, which is one of the simplest approaches to incorporate KL regularization. We show that this generalized MVI($q$) obtains significant improvements over the standard MVI($q = 1$) across 35 Atari games.
1.2NADec 21, 2017
Multiscale convergence properties for spectral approximations of a model kinetic equationZheng Chen, Cory D. Hauck
In this work, we prove rigorous convergence properties for a semi-discrete, moment-based approximation of a model kinetic equation in one dimension. This approximation is equivalent to a standard spectral method in the velocity variable of the kinetic distribution and, as such, is accompanied by standard algebraic estimates of the form $N^{-q}$, where $N$ is the number of modes and $q>0$ depends on the regularity of the solution. However, in the multiscale setting, the error estimate can be expressed in terms of the scaling parameter $ε$, which measures the ratio of the mean-free-path to the characteristic domain length. We show that, for isotropic initial conditions, the error in the spectral approximation is $\mathcal{O}(ε^{N+1})$. More surprisingly, the coefficients of the expansion satisfy super convergence properties. In particular, the error of the $\ell^{th}$ coefficient of the expansion scales like $\mathcal{O}(ε^{2N})$ when $\ell =0$ and $\mathcal{O}(ε^{2N+2-\ell})$ for all $1\leq \ell \leq N$. This result is significant, because the low-order coefficients correspond to physically relevant quantities of the underlying system. All the above estimates involve constants depending on $N$, the time $t$, and the initial condition. We investigate specifically the dependence on $N$, in order to assess whether increasing $N$ actually yields an additional factor of $ε$ in the error. Numerical tests will also be presented to support the theoretical results.
1.8LGMay 16, 2022
Enforcing KL Regularization in General Tsallis Entropy Reinforcement Learning via Advantage LearningLingwei Zhu, Zheng Chen, Eiji Uchibe et al.
Maximum Tsallis entropy (MTE) framework in reinforcement learning has gained popularity recently by virtue of its flexible modeling choices including the widely used Shannon entropy and sparse entropy. However, non-Shannon entropies suffer from approximation error and subsequent underperformance either due to its sensitivity or the lack of closed-form policy expression. To improve the tradeoff between flexibility and empirical performance, we propose to strengthen their error-robustness by enforcing implicit Kullback-Leibler (KL) regularization in MTE motivated by Munchausen DQN (MDQN). We do so by drawing connection between MDQN and advantage learning, by which MDQN is shown to fail on generalizing to the MTE framework. The proposed method Tsallis Advantage Learning (TAL) is verified on extensive experiments to not only significantly improve upon Tsallis-DQN for various non-closed-form Tsallis entropies, but also exhibits comparable performance to state-of-the-art maximum Shannon entropy algorithms.
4.1LGAug 6, 2025
Symmetric Behavior Regularization via Taylor Expansion of SymmetryLingwei Zhu, Zheng Chen, Han Wang et al.
This paper introduces symmetric divergences to behavior regularization policy optimization (BRPO) to establish a novel offline RL framework. Existing methods focus on asymmetric divergences such as KL to obtain analytic regularized policies and a practical minimization objective. We show that symmetric divergences do not permit an analytic policy as regularization and can incur numerical issues as loss. We tackle these challenges by the Taylor series of $f$-divergence. Specifically, we prove that an analytic policy can be obtained with a finite series. For loss, we observe that symmetric divergences can be decomposed into an asymmetry and a conditional symmetry term, Taylor-expanding the latter alleviates numerical issues. Summing together, we propose Symmetric $f$ Actor-Critic (S$f$-AC), the first practical BRPO algorithm with symmetric divergences. Experimental results on distribution approximation and MuJoCo verify that S$f$-AC performs competitively.
2.7CLFeb 10, 2025
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy StimulationsRui Chen, Tailai Peng, Xinran Xie et al.
Significant improvements have been observed in the zero-shot capabilities of the Large Language Models (LLMs). Due to their high sensitivity to input, research has increasingly focused on enhancing LLMs' performance via direct and simple prompt engineering rather than intricate domain adaptation. Studies suggest that LLMs exhibit emotional intelligence, and both positive and negative emotions can potentially enhance task performances. However, prior interaction prompts have predominantly concentrated on a single stimulus type, neglecting to compare different stimulus effects, examine the influence of varying task difficulties, or explore underlying mechanisms. This paper, inspired by the positive correlation between self-efficacy and task performance within the social cognitive theory, introduces Verbal Efficacy Stimulations (VES). Our VES comprises three types of verbal prompts: encouraging, provocative, and critical, addressing six aspects such as helpfulness and competence. And we further categorize task difficulty, aiming to extensively investigate how distinct VES influence the self-efficacy and task achievements of language models at varied levels of difficulty. The experimental results show that the three types of VES improve the performance of LLMs on most tasks, and the most effective VES varies for different models. In extensive experiments, we have obtained some findings consistent with psychological theories, providing novel insights for future research.
3.0ROOct 29, 2021
Efficient Map Prediction via Low-Rank Matrix CompletionZheng Chen, Shi Bai, Lantao Liu
In many autonomous mapping tasks, the maps cannot be accurately constructed due to various reasons such as sparse, noisy, and partial sensor measurements. We propose a novel map prediction method built upon the recent success of Low-Rank Matrix Completion. The proposed map prediction is able to achieve both map interpolation and extrapolation on raw poor-quality maps with missing or noisy observations. We validate with extensive simulated experiments that the approach can achieve real-time computation for large maps, and the performance is superior to the state-of-the-art map prediction approach - Bayesian Hilbert Mapping in terms of mapping accuracy and computation time. Then we demonstrate that with the proposed real-time map prediction framework, the coverage convergence rate (per action step) for a set of representative coverage planning methods commonly used for environmental modeling and monitoring tasks can be significantly improved.