Securing UAV Swarms with Vision Transformers: A Byzantine-Robust Federated Learning Framework for Cross-Modal Intrusion Detection
Küçük Resim Yok
Tarih
2026
Yazarlar
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Mdpi
Erişim Hakkı
info:eu-repo/semantics/openAccess
Özet
Highlights What are the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The combination of Vision Transformers, GAF encoding, and Byzantine-robust FL offers a scalable, privacy-preserving solution suitable for real-world UAV swarms operating under adversarial conditions. What are the implications of the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The newly introduced ReGCA aggregation method significantly improves federated robustness, maintaining 89.6% accuracy even with 40% Byzantine clients, more than 44 percentage points higher than FedAvg.Highlights What are the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The combination of Vision Transformers, GAF encoding, and Byzantine-robust FL offers a scalable, privacy-preserving solution suitable for real-world UAV swarms operating under adversarial conditions. What are the implications of the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The newly introduced ReGCA aggregation method significantly improves federated robustness, maintaining 89.6% accuracy even with 40% Byzantine clients, more than 44 percentage points higher than FedAvg.Abstract The increasing deployment of uncrewed aerial vehicles (UAVs) in cyber-physical and safety-critical missions has amplified the need for intrusion detection systems that are accurate, privacy-preserving, and resilient to adversarial manipulation. In this paper, we propose CM-BRF-ViT, a Cross-Modal Byzantine-Robust Federated Vision Transformer framework for UAV intrusion detection that jointly addresses heterogeneous attack modeling, distributed learning security, and adaptive decision fusion. The proposed framework integrates Gramian Angular Field (GAF) transformations with Vision Transformer (ViT) architectures to effectively convert tabular network and cyber-physical features into discriminative visual representations suitable for attention-based learning. To enable privacy-preserving collaboration across distributed UAV nodes, CM-BRF-ViT operates within a federated learning paradigm and introduces Reference-GAF Consistency Aggregation (ReGCA). This novel Byzantine-robust aggregation mechanism jointly measures prediction consistency and feature-level semantic consistency using a trusted reference set and MAD-based robust weighting. Unlike conventional defenses that rely solely on parameter-space filtering, ReGCA supervises model updates at both behavioral and representation levels, significantly enhancing robustness against malicious clients. In addition, a learnable cross-modal fusion head is developed to adaptively combine attack probabilities derived from cyber and cyber-physical modalities, allowing the framework to exploit complementary threat signatures across layers. Extensive experiments conducted on the UAVIDS-2025 and Cyber-Physical datasets demonstrate that the proposed method achieves 97.1% detection accuracy for UAV network traffic and 78.5% for cyber-physical data, with a fused detection AUC of 0.993. Under adversarial settings, CM-BRF-ViT preserves 89. 6% accuracy with up to 40% Byzantine clients, outperforming FedAvg by more than 44 percentage points. Ablation studies further confirm that ReGCA, cross-modal fusion, and ViT-based representation learning contribute complementary performance gains over baseline federated and centralized approaches. These results demonstrate that CM-BRF-ViT provides a robust, adaptive, and privacy-aware intrusion detection solution for UAV systems, making it well-suited for deployment in adversarial and resource-constrained aerial networks.
Açıklama
Anahtar Kelimeler
Uav Swarm Security, Federated Learning, Robust Aggregation, Vision Transformer, Blockchain, Intrusion Detection, Cyberattack Detection, Gramian Angular Field, Byzantine-Robust
Kaynak
Drones
WoS Q Değeri
Q1
Scopus Q Değeri
Q1
Cilt
10
Sayı
2












