22.7CLFeb 15, 2024
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language ModelsSaeed Khaki, JinJin Li, Lan Ma et al.
Reinforcement learning from human feedback (RLHF) has been extensively employed to align large language models with user intent. However, proximal policy optimization (PPO) based RLHF is occasionally unstable requiring significant hyperparameter finetuning, and computationally expensive to maximize the estimated reward during alignment. Recently, direct preference optimization (DPO) is proposed to address those challenges. However, DPO relies on contrastive responses generated from human annotator and alternative LLM, instead of the policy model, limiting the effectiveness of the RLHF. In this paper, we addresses both challenges by systematically combining rejection sampling (RS) and DPO. Our proposed method, RS-DPO, initiates with the development of a supervised fine-tuned policy model (SFT). A varied set of k responses per prompt are sampled directly from the SFT model. RS-DPO identifies pairs of contrastive samples based on their reward distribution. Finally, we apply DPO with the contrastive samples to align the model to human preference. Our experiments indicate that our proposed method effectively fine-tunes LLMs with limited resource environments, leading to improved alignment with user intent. Furthermore, it outperforms existing methods, including RS, PPO, and DPO.
6.2CVMay 21, 2025
BadSR: Stealthy Label Backdoor Attacks on Image Super-ResolutionJi Guo, Xiaolei Wen, Wenbo Jiang et al.
With the widespread application of super-resolution (SR) in various fields, researchers have begun to investigate its security. Previous studies have demonstrated that SR models can also be subjected to backdoor attacks through data poisoning, affecting downstream tasks. A backdoor SR model generates an attacker-predefined target image when given a triggered image while producing a normal high-resolution (HR) output for clean images. However, prior backdoor attacks on SR models have primarily focused on the stealthiness of poisoned low-resolution (LR) images while ignoring the stealthiness of poisoned HR images, making it easy for users to detect anomalous data. To address this problem, we propose BadSR, which improves the stealthiness of poisoned HR images. The key idea of BadSR is to approximate the clean HR image and the pre-defined target image in the feature space while ensuring that modifications to the clean HR image remain within a constrained range. The poisoned HR images generated by BadSR can be integrated with existing triggers. To further improve the effectiveness of BadSR, we design an adversarially optimized trigger and a backdoor gradient-driven poisoned sample selection method based on a genetic algorithm. The experimental results show that BadSR achieves a high attack success rate in various models and data sets, significantly affecting downstream tasks.
1.2GNNov 29, 2021
The language of pre-topology in knowledge spacesFucai Lin, Xiyan Cao, Jinjin Li
We systematically study some basic properties of the theory of pre-topological spaces, such as, pre-base, subspace, axioms of separation, connectedness, etc. Pre-topology is also known as knowledge space in the theory of knowledge structures. We discuss the language of axioms of separation of pre-topology in the theory of knowledge spaces, the relation of Alexandroff spaces and quasi ordinal spaces, and the applications of the density of pre-topological spaces in primary items for knowledge spaces. In particular, we give a characterization of a skill multimap such that the delineate knowledge structure is a knowledge space, which gives an answer to a problem in \cite{falmagne2011learning} or \cite{XGLJ} whenever each item with finitely many competencies; moreover, we give an algorithm to find the set of atom primary items for any finite knowledge spaces.