Hierarchical reinforcement learning with kill chain-informed multi-objective optimization to enhance resilience in autonomous unmanned swarm.

Gou, Yingdong; Wei, Siwen; Xu, Kai; Liu, Jiancheng; Li, Ke; Li, Bo; Han, Zaikun; Lai, Xin et al. · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

The resilience of autonomous unmanned swarms (AUS) serves as a cornerstone for guaranteeing continuous mission execution in the face of adversarial interferences. Contemporary methodologies often suffer from limited generalization efficacy and volatile convergence behavior, primarily due to the intricate, high-dimensional, and dynamically evolving landscape of multi-agent systems. In the context of AUS, conventional single-objective reinforcement learning (RL) paradigms amalgamate conflicting objectives into a unified scalar reward, thereby concealing critical trade-offs and undermining the swarm's capacity for adaptive response under adversarial stressors. To transcend these constraints, we propose HRL-KCIMOO, a hierarchical reinforcement learning framework that synergistically integrates kill chain-informed knowledge pre-training with dynamic multi-objective optimization. A graph attention encoder is pre-trained via contrastive representation learning, graph topology reconstruction, and centrality-aware ranking tasks, endowing each node with embeddings that intrinsically encapsulate causal linkages between adversarial maneuvers and corresponding defensive countermeasures. Subsequently, a high-level actor-critic architecture, augmented with short-term memory through LSTM modules, generates a temporally adaptive weight vector to dynamically reconcile objectives concerning rapid resilience restoration, sustained operational continuity, and long-term systemic robustness. In parallel, decentralized low-level agents execute context-sensitive cooperative behaviors or structural reconfigurations in accordance with the computed objective weights. Comprehensive empirical analyses across diverse adversarial landscapes demonstrate that HRL-KCIMOO consistently surpasses established benchmarks. Notably, even under severe conditions of swarm attrition, the proposed approach sustains significantly superior mission success rates compared to conventional methodologies.

Medical subject headings