DMPFDet: An end-to-end deformable mutual-promotion learning network for multispectral visible-infrared fusion detection.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

Multimodal image fusion and object detection significantly enhance detection accuracy and robustness in complex environments, which is crucial for autonomous driving. Despite the progress in multimodal fusion methods, most of them focus on generating visually appealing images while neglecting their effectiveness for downstream tasks. Meanwhile, existing multimodal detection methods often fail to fully exploit modality-specific features and cross-modal complementary information. Although some recent studies have attempted to integrate fusion and detection, they typically rely on multi-stage training pipelines and overlook the potential of mutual guidance between fused features and detection representations, leading to suboptimal performance. To address these challenges, this study proposes DMPFDet, an End-to-End Deformable Mutual-Promotion Learning Network for Multispectral Visible-Infrared Fusion Detection, which achieves both high-quality fusion and accurate detection within a unified training pipeline. It comprises two main components: a Multimodal Deformable Detection Transformer (MDDT) module for detection and a Cross-Modal Attention Fusion (CMAF) module for fusion. The MDDT is designed based on the RT-DETR architecture and is tailored to efficiently extract both modality-specific features and cross-modal complementary information. The CMAF effectively captures local texture details, global contextual information, as well as channel and spatial information. It is worth noting that the two modules are designed to mutually provide feature-level guidance, enabling joint optimization and reinforcing each other's learning processes. Experimental results on multiple datasets demonstrate the outstanding performance of the proposed method, outperforming state-of-the-art approaches. It not only produces effective fusion results for object detection but also delivers impressive detection outcomes.

Authors

Keywords

No keywords available for this article.