Joint Optimization of Cooperation Efficiency and Communication Covertness for Target Detection with AUVs
Xueyao Zhang, Bo Yang, Zhiwen Yu, Xuelin Cao, Wei Xiang, Bin Guo, Liang Wang, Billy Pik Lik Lau, George C. Alexandropoulos, Jun Luo, Mérouane Debbah, Zhu Han, Chau Yuen
TL;DR
This work tackles covert cooperative target detection with multiple AUVs by formulating a joint trajectory and power optimization problem under a covert communication constraint expressed as $D(\mathcal{H}_{0}||\mathcal{H}_{1})\le 2\varepsilon^{2}$. It introduces a hierarchical framework, HMAPPO, that splits decision-making into macro-level task allocation (MDP with PPO) and micro-level trajectory/power control (POMDP with MAPPO) under a CTDE regime. The approach yields efficient collaboration (high $\eta$ and $\zeta$) while maintaining low detectability (low KL divergence) and energy-aware operation, demonstrated via a high-fidelity 3D underwater simulator showing rapid convergence and favorable scalability. The results highlight a critical trade-off region: as covertness tightens, cooperation can degrade, revealing a threshold beyond which ultra-low power communications impair task execution, but the framework effectively navigates this balance through hierarchical learning and centralized guidance. Overall, the paper provides a practical, scalable method for covert multi-AUV coordination in dynamic underwater environments with clear implications for secure underwater sensing and surveillance tasks.
Abstract
This paper investigates underwater cooperative target detection using autonomous underwater vehicles (AUVs), with a focus on the critical trade-off between cooperation efficiency and communication covertness. To tackle this challenge, we first formulate a joint trajectory and power control optimization problem, and then present an innovative hierarchical action management framework to solve it. According to the hierarchical formulation, at the macro level, the master AUV models the agent selection process as a Markov decision process and deploys the proximal policy optimization algorithm for strategic task allocation. At the micro level, each selected agent's decentralized decision-making is modeled as a partially observable Markov decision process, and a multi-agent proximal policy optimization algorithm is used to dynamically adjust its trajectory and transmission power based on its local observations. Under the centralized training and decentralized execution paradigm, our target detection framework enables adaptive covert cooperation while satisfying both energy and mobility constraints. By comprehensively modeling the considered system, the involved signals and tasks, as well as energy consumption, theoretical insights and practical solutions for the efficient and secure operation of multiple AUVs are provided, offering significant implications for the execution of underwater covert communication tasks.
