8.8NEApr 6
Diffusion-based Evolutionary Optimization for 3D Multi-Objective Molecular GenerationRuiqing Sun, Dawei Feng, Sen Yang et al.
In 3D molecular discovery, optimizing conflicting physicochemical properties while strictly adhering to complex structural constraints constitutes a Constrained Multi-Objective Optimization Problem (CMOP). Solving this remains highly challenging: applying traditional Evolutionary Algorithm (EA) operators directly to 3D coordinates destroys chemical validity, whereas valid 3D diffusion models act as rigid generators unable to adapt to novel objectives without retraining. Moreover, employing traditional EA frameworks causes a severe loss of structural diversity, ultimately impairing algorithmic convergence. To overcome these challenges, we propose the Evolutionary-Guided Diffusion (EGD) operator, which executes crossover and mutation exclusively within the continuous noise space at an appropriate noise intensity. EGD enables topological hybridization while leveraging a pre-trained denoising network to project intermediate states back onto the valid chemical manifold. To tackle Multi-Objective Problems (MOPs), we introduce a Structure-Aware Environmental Selection (SAES) mechanism that explicitly enforces geometric diversity. Building upon this, to specifically solve CMOPs, we develop the Diffusion-based Evolutionary Molecular Optimization (DEMO) framework, utilizing a tri-population architecture with distinct responsibilities to safely navigate disjoint feasible regions. Extensive experiments across single-property targeting, unconstrained MOPs, multi-fragment constrained generation, and 3D protein-ligand docking demonstrate that DEMO comprehensively outperforms train-free guidance methods and EA baselines. Without any model retraining, DEMO successfully discovers highly diverse, chemically valid Pareto frontiers, establishing a robust paradigm for complex 3D molecular optimization.
5.8LGMay 21, 2022
Nuclear Norm Maximization Based Curiosity-Driven LearningChao Chen, Zijian Gao, Kele Xu et al.
To handle the sparsity of the extrinsic rewards in reinforcement learning, researchers have proposed intrinsic reward which enables the agent to learn the skills that might come in handy for pursuing the rewards in the future, such as encouraging the agent to visit novel states. However, the intrinsic reward can be noisy due to the undesirable environment's stochasticity and directly applying the noisy value predictions to supervise the policy is detrimental to improve the learning performance and efficiency. Moreover, many previous studies employ $\ell^2$ norm or variance to measure the exploration novelty, which will amplify the noise due to the square operation. In this paper, we address aforementioned challenges by proposing a novel curiosity leveraging the nuclear norm maximization (NNM), which can quantify the novelty of exploring the environment more accurately while providing high-tolerance to the noise and outliers. We conduct extensive experiments across a variety of benchmark environments and the results suggest that NNM can provide state-of-the-art performance compared with previous curiosity methods. On 26 Atari games subset, when trained with only intrinsic reward, NNM achieves a human-normalized score of 1.09, which doubles that of competitive intrinsic rewards-based approaches. Our code will be released publicly to enhance the reproducibility.
1.8LGAug 24, 2022
Self-Supervised Exploration via Temporal Inconsistency in Reinforcement LearningZijian Gao, Kele Xu, Yuanzhao Zhai et al.
Under sparse extrinsic reward settings, reinforcement learning has remained challenging, despite surging interests in this field. Previous attempts suggest that intrinsic reward can alleviate the issue caused by sparsity. In this article, we present a novel intrinsic reward that is inspired by human learning, as humans evaluate curiosity by comparing current observations with historical knowledge. Our method involves training a self-supervised prediction model, saving snapshots of the model parameters, and using nuclear norm to evaluate the temporal inconsistency between the predictions of different snapshots as intrinsic rewards. We also propose a variational weighting mechanism to assign weight to different snapshots in an adaptive manner. Our experimental results on various benchmark environments demonstrate the efficacy of our method, which outperforms other intrinsic reward-based methods without additional training costs and with higher noise tolerance. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
2.5AIAug 24, 2022
Dynamic Memory-based Curiosity: A Bootstrap Approach for ExplorationZijian Gao, YiYing Li, Kele Xu et al.
The sparsity of extrinsic rewards poses a serious challenge for reinforcement learning (RL). Currently, many efforts have been made on curiosity which can provide a representative intrinsic reward for effective exploration. However, the challenge is still far from being solved. In this paper, we present a novel curiosity for RL, named DyMeCu, which stands for Dynamic Memory-based Curiosity. Inspired by human curiosity and information theory, DyMeCu consists of a dynamic memory and dual online learners. The curiosity arouses if memorized information can not deal with the current state, and the information gap between dual learners can be formulated as the intrinsic reward for agents, and then such state information can be consolidated into the dynamic memory. Compared with previous curiosity methods, DyMeCu can better mimic human curiosity with dynamic memory, and the memory module can be dynamically grown based on a bootstrap paradigm with dual learners. On multiple benchmarks including DeepMind Control Suite and Atari Suite, large-scale empirical experiments are conducted and the results demonstrate that DyMeCu outperforms competitive curiosity-based methods with or without extrinsic rewards. We will release the code to enhance reproducibility.
4.6LGMay 23, 2024
Online Self-Preferring Language ModelsYuanzhao Zhai, Zhuo Zhang, Kele Xu et al.
Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward model to provide feedback for on-policy sampling responses. Recently, offline methods that directly fit responses with binary preferences in the dataset have emerged as alternatives. However, existing methods do not explicitly model preference strength information, which is crucial for distinguishing different response pairs. To overcome this limitation, we propose Online Self-Preferring (OSP) language models to learn from self-generated response pairs and self-judged preference strengths. For each prompt and corresponding self-generated responses, we introduce a ranked pairing method to construct multiple response pairs with preference strength information. We then propose the soft-preference cross-entropy loss to leverage such information. Empirically, we demonstrate that leveraging preference strength is crucial for avoiding overfitting and enhancing alignment performance. OSP achieves state-of-the-art alignment performance across various metrics in two widely used human preference datasets. OSP is parameter-efficient and more robust than the dominant online method, RLHF when limited offline data are available and generalizing to out-of-domain tasks. Moreover, OSP language models established by LLMs with proficiency in self-preferring can efficiently self-improve without external supervision.