AIJan 22, 2021

Theory of Mind for Deep Reinforcement Learning in Hanabi

Andrew Fuchs, Michael Walton, Theresa Chadwick, Doug Lange

arXiv:2101.09328v115.318 citationsHas Code

Originality Incremental advance

AI Analysis

This work addresses the challenge of implicit communication and theory of mind reasoning in cooperative AI for games like Hanabi, representing an incremental advance in applying such reasoning to reinforcement learning.

The authors tackled the problem of enabling deep reinforcement learning agents to cooperate effectively in the partially observable game Hanabi by developing a theory of mind mechanism, resulting in improved performance against a state-of-the-art baseline agent.

The partially observable card game Hanabi has recently been proposed as a new AI challenge problem due to its dependence on implicit communication conventions and apparent necessity of theory of mind reasoning for efficient play. In this work, we propose a mechanism for imbuing Reinforcement Learning agents with a theory of mind to discover efficient cooperative strategies in Hanabi. The primary contributions of this work are threefold: First, a formal definition of a computationally tractable mechanism for computing hand probabilities in Hanabi. Second, an extension to conventional Deep Reinforcement Learning that introduces reasoning over finitely nested theory of mind belief hierarchies. Finally, an intrinsic reward mechanism enabled by theory of mind that incentivizes agents to share strategically relevant private knowledge with their teammates. We demonstrate the utility of our algorithm against Rainbow, a state-of-the-art Reinforcement Learning agent.

View on arXiv PDF Code

Similar