Collaborative Ground-Space Communications via Evolutionary Multi-objective Deep Reinforcement Learning

Jiahui Li; Geng Sun; Qingqing Wu; Dusit Niyato; Jiawen Kang; Abbas Jamalipour; Victor C. M. Leung

Collaborative Ground-Space Communications via Evolutionary Multi-objective Deep Reinforcement Learning

Jiahui Li, Geng Sun, Qingqing Wu, Dusit Niyato, Jiawen Kang, Abbas Jamalipour, Victor C. M. Leung

TL;DR

The paper tackles enabling direct ground-space uplinks for energy-constrained terrestrial terminals by leveraging Distributed Collaborative Beamforming (DCB) to form a virtual antenna array with LEO satellites. It reformulates the long-term, multi-objective optimization into an action-space-reduced, universal MOMDP and introduces an evolutionary multi-objective DRL framework (EMODRL-ED3QN) that yields multiple Pareto-style policies. The approach combines a convex one-slot power-weighting problem to fix action dimensions, a masked D3QN agent to learn policies, and an evolutionary coordination mechanism to broaden policy diversity while maintaining portability across scenarios. Simulations with dense LEO constellations demonstrate that EMODRL-ED3QN outperforms baselines, enabling previously marginal terminals to achieve efficient uplinks and providing near-optimal rates with low satellite-switching frequency. The results underscore the practicality and adaptability of portable, multi-objective learning for dynamic ground-space communication systems.

Abstract

In this paper, we propose a distributed collaborative beamforming (DCB)-based uplink communication paradigm for enabling ground-space direct communications. Specifically, DCB treats the terminals that are unable to establish efficient direct connections with the low Earth orbit (LEO) satellites as distributed antennas, forming a virtual antenna array to enhance the terminal-to-satellite uplink achievable rates and durations. However, such systems need multiple trade-off policies that variously balance the terminal-satellite uplink achievable rate, energy consumption of terminals, and satellite switching frequency to satisfy the scenario requirement changes. Thus, we perform a multi-objective optimization analysis and formulate a long-term optimization problem. To address availability in different terminal cluster scales, we reformulate this problem into an action space-reduced and universal multi-objective Markov decision process. Then, we propose an evolutionary multi-objective deep reinforcement learning algorithm to obtain the desirable policies, in which the low-value actions are masked to speed up the training process. As such, the applicability of a one-time trained model can cover more changing terminal-satellite uplink scenarios. Simulation results show that the proposed algorithm outmatches various baselines, and draw some useful insights. Specifically, it is found that DCB enables terminals that cannot reach the uplink achievable threshold to achieve efficient direct uplink transmission, which thus reveals that DCB is an effective solution for enabling direct ground-space communications. Moreover, it reveals that the proposed algorithm achieves multiple policies favoring different objectives and achieving near-optimal uplink achievable rates with low switching frequency.

Collaborative Ground-Space Communications via Evolutionary Multi-objective Deep Reinforcement Learning

TL;DR

Abstract

Paper Structure (29 sections, 1 theorem, 18 equations, 8 figures, 1 table, 4 algorithms)

This paper contains 29 sections, 1 theorem, 18 equations, 8 figures, 1 table, 4 algorithms.

Introduction
Related Works
System Models and Preliminaries
Network Segments
LEO Satellite Orbit
Virtual Antenna Array Model
Satellite Switching Model
Problem Formulation and Analyses
Problem Statement
Problem Analyses
Multi-objective DRL-based Method
MOMDP Simplification and Formulation
Action Transition
MOMDP Formulation
EMODRL-based Solution
...and 14 more sections

Key Result

Lemma 1

In the considered scenarios and feasible set of $P_i$ ($i \in \mathcal{I}$), the problem $(\mathrm{P2})$ is convex.

Figures (8)

Figure 1: Due to the low uplink gain and transmit power of the terminals, the single terminal to LEO uplink only continues short time. Benefiting from the transmission gain of DCB, the virtual antenna array will achieve extended connection duration.
Figure 2: A terminal cluster to LEO satellites communication system. All the terminals can directly connect with LEO satellites that are with fixed earth orbits. Terminals will form a virtual antenna array and select a suitable LEO to perform uplink data transmission.
Figure 3: Framework of EMODRL-ED3QN for Multi-objective optimization in Collaborative Ground-Space Communications.
Figure 4: Uplink achievable rates obtained by an EMODRL-ED3QN policy, ARGP, and non-DCB strategy.
Figure 5: Pareto policy distributions obtained by different algorithms. Each point represents a Pareto policy obtained by an algorithm, and its three coordinate values represent the optimization objective values achieved by this policy. We mark the direction of the Pareto front (i.e., ideal Pareto policy set), and the policy closer to the Pareto front will achieve better performance.
...and 3 more figures

Theorems & Definitions (1)

Lemma 1

Collaborative Ground-Space Communications via Evolutionary Multi-objective Deep Reinforcement Learning

TL;DR

Abstract

Collaborative Ground-Space Communications via Evolutionary Multi-objective Deep Reinforcement Learning

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (8)

Theorems & Definitions (1)