Goal-Oriented Bundle Management and Flow Control for Deep-Space Communications via Reinforcement Learning
77th International Astronautical Congress (IAC 2026), Antalya, Türkiye
Abstract
Deep-space communications operate under intermittent connectivity, long propagation delays, and relay storage constraints. To span vast distances, data are often forwarded through relay nodes, where incoming traffic can vary over time. In such networks, the Bundle Protocol encapsulates data into store-carry-forward units called bundles, which are stored at intermediate nodes and forwarded only when contact opportunities become available. Therefore, relay-side buffer and flow management become critical challenges, since limited storage and scarce contact windows force the relay to decide which bundles to transmit or discard. This decision process is inherently goal-oriented: the relay should not merely forward data, but should allocate limited transmission opportunities to the bundles that are most valuable to the downstream mission objective. Existing schedulers, such as FIFO, are heuristic and do not explicitly optimize task utility; moreover, they often lead to unfair service across heterogeneous planetary sources. This paper presents a closed-loop modeling and learning framework for relay-side bundle and flow management in a contact-driven deep-space network, where updates generated by lunar and Martian sources are forwarded via a relay to Earth. These bundles are stored in a shared relay buffer and carry metadata including source and destination identities, generation time, and time-to-live. On Earth, data utility depends on freshness, which we quantify via the excess age ∆e_i(t) = [∆_i(t) − d_i(t)]+, where ∆_i(t) denotes the Age of Information (AoI) of source i, i.e., the time elapsed since the freshest update currently available on Earth was generated, and d_i(t) is a source-dependent delay baseline. Based on this metric, we define a goal-oriented system loss that penalizes excess age in a source-dependent and nonlinear manner, allowing the relay to prioritize information according to mission relevance. We formulate relay control as a constrained semi-Markov decision process (CSMDP) and solve it using masked Lagrangian Proximal Policy Optimization (PPO), where feasibility-aware action masking restricts decisions to transmissions that can finish within the remaining contact window. Simulation results in DTNSim indicate that the proposed approach significantly improves age-based utility and achieves substantially better inter-flow fairness compared to baseline schedulers. By explicitly prioritizing information according to downstream task relevance and using this signal to guide transmission and dropping decisions, the proposed framework improves goal-oriented information delivery in deep-space networks and, to our knowledge, is among the first reinforcement learning frameworks to jointly optimize relay-side bundle scheduling and dropping, as well as flow control under contact-driven dynamics with explicit fairness considerations across heterogeneous sources.
Selected figures
Cite this work
S. Baghaee, A. Li, S. Balasubramanian and E. Uysal, "Goal-Oriented Bundle Management and Flow Control for Deep-Space Communications via Reinforcement Learning," in 77th International Astronautical Congress (IAC 2026), Antalya, Türkiye, IAC-26,B2,5,8,x113206, 2026.
@inproceedings{baghaee2026goaloriented,
author = {Baghaee, Sajjad and Li, Aimin and Balasubramanian, Shanmugarajan and Uysal, Elif},
title = {Goal-Oriented Bundle Management and Flow Control for Deep-Space Communications via Reinforcement Learning},
booktitle = {77th International Astronautical Congress (IAC 2026), Antalya, Türkiye},
number = {IAC-26,B2,5,8,x113206},
publisher = {International Astronautical Federation (IAF)},
year = {2026},
url = {https://iafastro.directory/iac/paper/id/113206/summary/}
}
Related work
Copyright © 2026 by the International Astronautical Federation (IAF). All rights reserved. The copy offered here is the authors' own.