ROBUST UAV SWARM DECISION-MAKING UNDER UNCERTAINTY VIA COMBINATION OF MULTI-AGENT REINFORCEMENT LEARNING AND HYBRID AUTOMATA

The author is a member of the editorial board of the publication, therefore he did not participate in the review and decision-making regarding the publication of this article.

Authors

  • Oleksii BYCHKOV, DSc (Engin.), Prof. Taras Shevchenko National University of Kyiv, Kyiv, Ukraine Author
  • Mykhailo VOITOVYCH, PhD Student Taras Shevchenko National University of Kyiv, Kyiv, Ukraine Author

DOI:

https://doi.org/10.17721/AIT.2025.2.05

Keywords:

Multi-agent reinforcement learning; UAV swarm; Hybrid automata; Uncertain environments; Electronic warfare; Multi-Agent Deep Deterministic Policy Gradient; Autonomous decision-making

Abstract

The article is presented in the OnlineFirst format, which involves early posting of materials on the journal website before their final printing in paper form.

The article has passed the review procedure and has been accepted for publication.

The article is available for search and citation by its DOI, which remains unchanged. After the issue is printed in the printing house, a new citation can be made using information about the volume number and page number, using the same DOI reference. 

Background. Unmanned Aerial Vehicle (UAV) swarms are increasingly deployed in complex, uncertain environments such as electronic warfare (EW) scenarios, where traditional rule-based control can struggle. Related works have applied multi-agent reinforcement learning (MARL) to enable UAVs to learn cooperative behaviors in such scenarios, and have explored splitting complex missions (e.g. search & track) into phases to simplify control. However, purely learning-based approaches may lack robustness or require prohibitively long training to handle all edge cases.

Methods. Our approach integrates the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm with a supervisory hybrid automaton that encodes high-level behavior modes. The hybrid automaton defines discrete modes for different mission phases (e.g. exploration vs. target tracking) and emergency conditions (e.g. evasion under jamming), with formal guard conditions triggering mode transitions. Within each mode, UAV agents execute continuous control actions learned via MADDPG’s centralized-training, decentralized-execution paradigm. We develop a simulation environment based on Gym-PyBullet-Drones for centralized training of a UAV swarm, including realistic physics and an EW interference model.

Results. The simulation experiments demonstrated that the proposed hybrid RL-automaton approach outperformed a standard end-to-end MADDPG baseline across all tested conditions. The hybrid swarm achieved consistently higher mission success rates, especially under adversarial jamming, where it maintained over 90% success compared to significant drops for the baseline. Learning curves showed faster convergence and higher final returns for the hybrid method, reflecting the benefit of structured mode-conditioned training. In dynamic target scenarios, the hybrid swarm maintained longer continuous tracking times through coordinated tactics, while the baseline frequently lost contact. These results highlight that integrating a formal automaton with multi-agent RL yields more robust, efficient, and scalable UAV swarm decision-making under uncertainty. These findings also align with recent studies showing that dividing missions into phases and training mode-specialized behaviors can significantly improve performance over monolithic policies.

Conclusions. The proposed hybrid model demonstrates that integrating formal automata-based mode switching with deep RL can yield more reliable and efficient decision-making for UAV swarms in uncertain, adversarial environments. Discrete mode supervision injects domain knowledge (e.g. when to switch tasks or enter a safe mode) that guides learning and ensures safety constraints, while MADDPG optimizes the continuous control within each mode. This hybrid approach offers improved adaptability to evolving threats and environmental changes, as evidenced by robust performance under EW conditions. Future work will extend this framework to more complex swarm scenarios using heterogeneous UAV swarms and explore automatic mode discovery (e.g. via hierarchical RL) to further reduce reliance on expert-defined automata.

Downloads

Download data is not yet available.

Author Biographies

  • Oleksii BYCHKOV, DSc (Engin.), Prof., Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

    ORCID ID:0000-0002-9378-9535

  • Mykhailo VOITOVYCH, PhD Student, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

    ORCID ID: 0009-0003-0405-0304

References

Casau, P., Cabecinhas, D., & Silvestre, C. (2011). Autonomous transition flight for a vertical take-off and landing aircraft. In 2011 50th IEEE Conference on Decision and Control and European Control Conference (CDC-ECC) (pp. 3974–3979). IEEE. https://doi.org/10.1109/CDC.2011.6160819

Chi, P., Wei, J., Wu, K., Di, B., & Wang, Y. (2023). A bio-inspired decision-making method of UAV swarm for attack-defense confrontation via multi-agent reinforcement learning. Biomimetics, 8(2), 222. https://doi.org/10.3390/biomimetics8020222

Ekechi, C. C., Elfouly, T., Alouani, A., & Khattab, T. (2025). A survey on UAV control with multi-agent reinforcement learning. Drones, 9(7), 484. https://doi.org/10.3390/drones9070484

Feng, Z., Na, X., Hai, S., Sun, Q., & Shi, J. (2025). Deep reinforcement learning for UAV target search and continuous tracking in complex environments with Gaussian process regression and prior policy embedding. Electronics, 14(7), 1330. https://doi.org/10.3390/electronics14071330

Liu, M., Wei, J., & Liu, K. (2024). A two-stage target search and tracking method for UAV based on deep reinforcement learning. Drones, 8(10), 544. https://doi.org/10.3390/drones8100544

Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., & Mordatch, I. (2017). Multi-agent actor-critic for mixed cooperative-competitive environments (arXiv:1706.02275). arXiv. https://doi.org/10.48550/arXiv.1706.02275

Panerati, J., Zheng, H., Zhou, S., Xu, J., Prorok, A., & Schoellig, A. P. (2021). Learning to fly — a Gym environment with PyBullet physics for reinforcement learning of multi-agent quadcopter control (arXiv:2103.02142). arXiv. https://doi.org/10.48550/arXiv.2103.02142

Sierra-García, J. E., & Santos, M. (2021). Intelligent control of a UAV with a cable-suspended load using a neural network estimator. Expert Systems with Applications, 185, 115380. https://doi.org/10.1016/j.eswa.2021.115380

Xu, J., & Liang, G. (2025). Enhanced Q-learning and deep reinforcement learning for unmanned combat intelligence planning in adversarial environments. Scientific Reports, 15(1), 28364. https://doi.org/10.1038/s41598-025-13752-3

Zhang, K., Yang, Z., & Başar, T. (2021). Multi-agent reinforcement learning: A selective overview of theories and algorithms. In K. G. Vamvoudakis, Y. Wan, F. L. Lewis, & D. Cansever (Eds.), Handbook of reinforcement learning and control (Studies in Systems, Decision and Control, Vol. 325, pp. 321–384). Springer. https://doi.org/10.1007/978-3-030-60990-0_12

Zhao, B., Huo, M., Li, Z., Feng, W., Feng, W.-T., Yu, Z., Qi, N., & Wang, S. (2025). Graph-based multi-agent reinforcement learning for collaborative search and tracking of multiple UAVs. Chinese Journal of Aeronautics, 38(3), 103214. https://doi.org/10.1016/j.cja.2024.08.045

Downloads

Published

2026-05-06

Issue

Section

Artificial and computational intelligence

How to Cite

ROBUST UAV SWARM DECISION-MAKING UNDER UNCERTAINTY VIA COMBINATION OF MULTI-AGENT REINFORCEMENT LEARNING AND HYBRID AUTOMATA: The author is a member of the editorial board of the publication, therefore he did not participate in the review and decision-making regarding the publication of this article. (2026). Advanced Information Technology, 1(2(5). https://doi.org/10.17721/AIT.2025.2.05

Most read articles by the same author(s)