Ming Jiang

CL
h-index20
3papers
45citations
Novelty52%
AI Score38

3 Papers

10.0CLOct 29, 2024Code
$f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Jiaqi Han, Mingjian Jiang, Yuxuan Song et al.

Preference optimization has made significant progress recently, with numerous methods developed to align language models with human preferences. This paper introduces $f$-divergence Preference Optimization ($f$-PO), a novel framework that generalizes and extends existing approaches. $f$-PO minimizes $f$-divergences between the optimized policy and the optimal policy, encompassing a broad family of alignment methods using various divergences. Our approach unifies previous algorithms like DPO and EXO, while offering new variants through different choices of $f$-divergences. We provide theoretical analysis of $f$-PO's properties and conduct extensive experiments on state-of-the-art language models using benchmark datasets. Results demonstrate $f$-PO's effectiveness across various tasks, achieving superior performance compared to existing methods on popular benchmarks such as AlpacaEval 2, Arena-Hard, MT-Bench, and Open LLM Leaderboard v2. Additionally, we present ablation studies exploring the impact of different $f$-divergences, offering insights into the trade-offs between regularization and performance in offline preference optimization. Our work contributes both practical algorithms and theoretical understanding to the field of language model alignment. Code is available at https://github.com/MinkaiXu/fPO.

11.9CLDec 28, 2024Code
No Preference Left Behind: Group Distributional Preference Optimization

Binwei Yao, Zefan Cai, Yun-Shiuan Chuang et al.

Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they struggle to capture the distributional pluralistic preferences within a group. These methods often skew toward dominant preferences, overlooking the diversity of opinions, especially when conflicting preferences arise. To address this issue, we propose Group Distributional Preference Optimization (GDPO), a novel framework that aligns language models with the distribution of preferences within a group by incorporating the concept of beliefs that shape individual preferences. GDPO calibrates a language model using statistical estimation of the group's belief distribution and aligns the model with belief-conditioned preferences, offering a more inclusive alignment framework than traditional methods. In experiments using both synthetic controllable opinion generation and real-world movie review datasets, we show that DPO fails to align with the targeted belief distributions, while GDPO consistently reduces this alignment gap during training. Moreover, our evaluation metrics demonstrate that GDPO outperforms existing approaches in aligning with group distributional preferences, marking a significant advance in pluralistic alignment.

1.5LGSep 4, 2018
A Neural Network Aided Approach for LDPC Coded DCO-OFDM with Clipping Distortion

Yuan He, Ming Jiang, Chunming Zhao

In this paper, a neural network-aided bit-interleaved coded modulation (NN-BICM) receiver is designed to mitigate the nonlinear clipping distortion in the LDPC coded direct currentbiased optical orthogonal frequency division multiplexing (DCOOFDM) systems. Taking the cross-entropy as loss function, a feed forward network is trained by backpropagation algorithm to output the condition probability through the softmax activation function, thereby assisting in a modified log-likelihood ratio (LLR) improvement. To reduce the complexity, this feed-forward network simplifies the input layer with a single symbol and the corresponding Gaussian variance instead of focusing on the inter-carrier interference between multiple subcarriers. On the basis of the neural network-aided BICM with Gray labelling, we propose a novel stacked network architecture of the bitinterleaved coded modulation with iterative decoding (NN-BICMID). Its performance has been improved further by calculating the condition probability with the aid of a priori probability that derived from the extrinsic LLRs in the LDPC decoder at the last iteration, at the expense of customizing neural network detectors at each iteration time separately. Utilizing the optimal DC bias as the midpoint of the dynamic region, the simulation results demonstrate that both the NN-BICM and NN-BICM-ID schemes achieve noticeable performance gains than other counterparts, in which the NN-BICM-ID clearly outperforms NN-BICM with various modulation and coding schemes.