Securing UAV Swarms with Vision Transformers: A Byzantine-Robust Federated Learning Framework for Cross-Modal Intrusion Detection

dc.contributor.authorSahin, Canan Batur
dc.date.accessioned2026-06-19T06:37:46Z
dc.date.available2026-06-19T06:37:46Z
dc.date.issued2026
dc.departmentMalatya Turgut Özal Üniversitesi
dc.description.abstractHighlights What are the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The combination of Vision Transformers, GAF encoding, and Byzantine-robust FL offers a scalable, privacy-preserving solution suitable for real-world UAV swarms operating under adversarial conditions. What are the implications of the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The newly introduced ReGCA aggregation method significantly improves federated robustness, maintaining 89.6% accuracy even with 40% Byzantine clients, more than 44 percentage points higher than FedAvg.Highlights What are the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The combination of Vision Transformers, GAF encoding, and Byzantine-robust FL offers a scalable, privacy-preserving solution suitable for real-world UAV swarms operating under adversarial conditions. What are the implications of the main findings? The fusion of cyber and cyber-physical modalities enables high-confidence UAV intrusion detection, providing reliable decision-making for safety-critical aerial missions. The newly introduced ReGCA aggregation method significantly improves federated robustness, maintaining 89.6% accuracy even with 40% Byzantine clients, more than 44 percentage points higher than FedAvg.Abstract The increasing deployment of uncrewed aerial vehicles (UAVs) in cyber-physical and safety-critical missions has amplified the need for intrusion detection systems that are accurate, privacy-preserving, and resilient to adversarial manipulation. In this paper, we propose CM-BRF-ViT, a Cross-Modal Byzantine-Robust Federated Vision Transformer framework for UAV intrusion detection that jointly addresses heterogeneous attack modeling, distributed learning security, and adaptive decision fusion. The proposed framework integrates Gramian Angular Field (GAF) transformations with Vision Transformer (ViT) architectures to effectively convert tabular network and cyber-physical features into discriminative visual representations suitable for attention-based learning. To enable privacy-preserving collaboration across distributed UAV nodes, CM-BRF-ViT operates within a federated learning paradigm and introduces Reference-GAF Consistency Aggregation (ReGCA). This novel Byzantine-robust aggregation mechanism jointly measures prediction consistency and feature-level semantic consistency using a trusted reference set and MAD-based robust weighting. Unlike conventional defenses that rely solely on parameter-space filtering, ReGCA supervises model updates at both behavioral and representation levels, significantly enhancing robustness against malicious clients. In addition, a learnable cross-modal fusion head is developed to adaptively combine attack probabilities derived from cyber and cyber-physical modalities, allowing the framework to exploit complementary threat signatures across layers. Extensive experiments conducted on the UAVIDS-2025 and Cyber-Physical datasets demonstrate that the proposed method achieves 97.1% detection accuracy for UAV network traffic and 78.5% for cyber-physical data, with a fused detection AUC of 0.993. Under adversarial settings, CM-BRF-ViT preserves 89. 6% accuracy with up to 40% Byzantine clients, outperforming FedAvg by more than 44 percentage points. Ablation studies further confirm that ReGCA, cross-modal fusion, and ViT-based representation learning contribute complementary performance gains over baseline federated and centralized approaches. These results demonstrate that CM-BRF-ViT provides a robust, adaptive, and privacy-aware intrusion detection solution for UAV systems, making it well-suited for deployment in adversarial and resource-constrained aerial networks.
dc.identifier.doi10.3390/drones10020125
dc.identifier.issn2504-446X
dc.identifier.issue2
dc.identifier.scopus2-s2.0-105031109883
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/drones10020125
dc.identifier.urihttps://hdl.handle.net/20.500.12899/5214
dc.identifier.volume10
dc.identifier.wosWOS:001701463000001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.institutionauthorSahin, Canan Batur
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofDrones
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20260612
dc.subjectUav Swarm Security
dc.subjectFederated Learning
dc.subjectRobust Aggregation
dc.subjectVision Transformer
dc.subjectBlockchain
dc.subjectIntrusion Detection
dc.subjectCyberattack Detection
dc.subjectGramian Angular Field
dc.subjectByzantine-Robust
dc.titleSecuring UAV Swarms with Vision Transformers: A Byzantine-Robust Federated Learning Framework for Cross-Modal Intrusion Detection
dc.typeArticle

Dosyalar