AI & Computingarticle2026-09-08

Research on multimodal dense small object detection algorithms guided by cross-modal information

Open access0 citations

Abstract

To address the challenges of modality heterogeneity, scale inconsistency, and background interference in dense small object detection under multimodal conditions, this paper proposes a novel detection framework based on cross-modality guidance and hierarchical scale refinement. Built upon the RT-DETR backbone, the framework integrates a Cross-Modality Guided Dynamic Fusion (CMG-DF) module, which performs semantic-level recalibration between infrared and visible features via a learnable modality attention mechanism, and a Hierarchical Scale Refinement Network (HSRN), which enhances semantic consistency and boundary continuity across scales through bidirectional residual flow and graph-based relational modeling. To validate the effectiveness of the proposed method, extensive comparison and ablation studies are conducted on two public multimodal benchmarks, SMOD and LLVIP. Experimental results show that the proposed method achieves 92.7% mAP@50 and 69.5% mAP@50:95 on SMOD, as well as 77.3% mAP@50 and 43.1% mAP@50:95 on LLVIP, consistently outperforming existing state-of-the-art multimodal detection algorithms. Qualitative visualizations further confirm the robustness and enhancement capability of the method for small objects under low illumination, occlusion, and complex background conditions, highlighting its strong structural generalization and practical deployment potential.

// Source

Authors: Jinxia Hu, Yumeng Ma, Yue Xing, Yun Zi, Yi Deng, Ming Wang, Heyao Liu, Junliang Du

Institutions: University of Pennsylvania, Shanghai Jiao Tong University, Tulane University, Northeastern University, Georgia Institute of Technology, Arizona State University, Trine University