[2607.15628]

MSTF-Net: A UAV-Oriented Multi-Spectral Video Segmentation Method via Modality-Robust, Scale-Adaptive, and Consistent Fusion


Multi-spectral video segmentation is essential for robust scene understanding in unmanned aerial vehicle (UAV) applications such as city planning, land use monitoring, traffic monitoring, and crowd estimation. While the fusion of RGB and thermal modalities offers complementary information for perception under varying lighting and visibility conditions, two fundamental challenges remain: (1) the modal fusion dilemma, arising from significant discrepancies between RGB and thermal features that obscure complementary cues, and (2) temporal variation, induced by rapid motion and viewpoint changes on UAV platforms, which leads to appearance inconsistency and misalignment across frames. To address these issues, this study proposed MSTF-Net, a modality-robust scale-adaptive fusion framework for multi-spectral video segmentation that effectively models cross-modal fusion and temporal consistency. The Modality Spatial Complementary Suppression and Enhancement (MSCSE) module generates unified instance queries via cross-modal attention and suppresses modality-specific noise using residual-guided discrepancy filtering and consistency constraints. To model temporal dynamics, the Multi-scale Temporal Cross-modality Semantic Consistency (MTCSC) module adaptively adjusts the temporal receptive field based on frame distance, capturing both coarse global context and fine local structure across time. Extensive ablation experiments on public RGB-T datasets demonstrate that MSTF-Net achieves state-of-the-art segmentation performance, especially under challenging conditions such as small targets, occlusion, and modality degradation. Specifically, reached 56.42\% mIoU on the MVSeg dataset and 51.80\% mIoU on the CART dataset, respectively.