A Quantum-Classical Framework for UAV-based Multi-Agent Reinforcement Learning

Document Type : Original Article

Authors

School of Mathematics and Computer Science, Iran University of Science and Technology, Tehran, Iran

Abstract
This paper presents a quantum-classical reinforcement learning framework designed to improve exploration, stability, and
coordination in multi-UAV systems operating under partial observability and non-stationary dynamics. The method integrates
centralized training with decentralized execution and introduces a quantum-inspired optimization layer implemented via simulated variational circuits. This hybrid structure reshapes the exploration landscape through correlated sampling and energybased objective refinement, allowing UAV agents to avoid premature convergence and maintain robust performance under noise and perturbations. A formal mathematical model of the actor, critic, and quantum-inspired parameterization is provided, along with a complexity analysis covering runtime and sample efficiency. Extensive experiments are conducted in cooperative UAV surveillance, multi-aircraft task allocation, and adversarial pursuit–evasion scenarios. Additional physical-world validation is performed on a 3-UAV testbed. Results demonstrate improved convergence stability, reduced sensitivity to hyperparameters, and consistent gains over classical baselines without overstating claims of quantum advantage. All code and experimental scripts are publicly available for full reproducibility.

Keywords

Subjects


Articles in Press, Accepted Manuscript
Available Online from 29 September 2026

  • Receive Date 23 November 2025
  • Revise Date 26 June 2026
  • Accept Date 30 July 2026