Qisheng Wang

h-index5
2papers
481citations

2 Papers

11.3SPJul 31, 2019
PrecoderNet: Hybrid Beamforming for Millimeter Wave Systems with Deep Reinforcement Learning

Qisheng Wang, Keming Feng, Xiao Li et al.

In this letter, we investigate the hybrid beamforming for millimeter wave massive multiple-input multiple-output (MIMO) system based on deep reinforcement learning (DRL). Imperfect channel state information (CSI) is assumed to be available at the base station (BS). To achieve high spectral efficiency with low time consumption, we propose a novel DRL-based method called PrecoderNet to design the digital precoder and analog combiner. The DRL agent takes the digital beamformer and analog combiner of the previous learning iteration as state, and these matrices of current learning iteration as action. Simulation results demonstrate that the PrecoderNet performs well in spectral efficiency, bit error rate (BER), as well as time consumption, and is robust to the CSI imperfection.

1.0LGJul 18, 2019
Prioritized Guidance for Efficient Multi-Agent Reinforcement Learning Exploration

Qisheng Wang, Qichao Wang

Exploration efficiency is a challenging problem in multi-agent reinforcement learning (MARL), as the policy learned by confederate MARL depends on the collaborative approach among multiple agents. Another important problem is the less informative reward restricts the learning speed of MARL compared with the informative label in supervised learning. In this work, we leverage on a novel communication method to guide MARL to accelerate exploration and propose a predictive network to forecast the reward of current state-action pair and use the guidance learned by the predictive network to modify the reward function. An improved prioritized experience replay is employed to better take advantage of the different knowledge learned by different agents which utilizes Time-difference (TD) error more effectively. Experimental results demonstrates that the proposed algorithm outperforms existing methods in cooperative multi-agent environments. We remark that this algorithm can be extended to supervised learning to speed up its training.