Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection
For researchers in spacecraft inspection and multi-agent control, this work offers a more flexible reward function, but the results are incremental as it builds on existing MARL methods.
This work develops a generalized reward function for Multi-Agent Reinforcement Learning (MARL) control of inspection spacecraft, enabling agents to autonomously decide when and where to collect images. The approach provides insights into best practices for MARL-based inspection tasks.
A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MARL). While MARL has already been employed for this purpose in previous work, the reward functions used focus on reaching a finite set of predetermined inspection points around the target. In this work, we study and develop a generalized reward function for the MARL inspection task informed by the analysis of 3D reconstructions of inspected objects in orbit. Because the reward function is generalized such that any number of images at arbitrary locations may evaluated, we also allow trained agents to have complete control over when images are collected. With this approach, we gather insights into best practices for not only the specific MARL inspection task, but also gain key takeaways informative to the broader inspection task outside of a MARL context.