Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

arXiv:2607.226678.2
Predicted impact top 69% in AI · last 90 daysOriginality Synthesis-oriented
AI Analysis

For maritime surveillance systems, this work offers a computationally efficient sensor selection method that maintains tracking accuracy, though it is an incremental improvement over existing information-theoretic approaches.

This paper introduces a reinforcement learning framework for selecting a single sensor at each decision step in heterogeneous maritime sensor networks, achieving tracking performance close to always-on sensing while reducing computational cost. The learned policy activates only one sensor per time step, avoiding expensive online entropy calculations.

This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors deployed in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes