Acta Geodaetica et Cartographica Sinica ›› 2026, Vol. 55 ›› Issue (8): 1343-1356.doi: 10.11947/j.AGCS.2026.20250297

• Advanced Technologies in Imaging Geodesy and Innovative Applications in Smart Disaster Prevention • Previous Articles    

Multi-modal image matching based on modality reconstruction and feature perturbation learning

Tengfeng Tang1(), Renyuan Liu1, Chang Liu1, Yangang Zhao2, Dan Pan3, Yuanxin Ye1()   

  1. 1.Faculty of Geosciences and Engineering, Southwest Jiaotong University, Chengdu 611756, China
    2.The Second Topographic Surveying Brigade of Ministry Natural Resource, Xi'an 710054, China
    3.Beijing Weina Starry Sky Technology Co., Ltd., Beijing 100089, China
  • Received:2025-07-24 Revised:2025-12-05 Published:2026-09-09
  • Contact: Yuanxin Ye E-mail:ttf@my.swjtu.edu.cn;yeyuanxin@home.swjtu.edu.cn
  • About author:Tang Tengfeng (2000—), male, PhD candidate, majors in multi-source remote sensing image matching. E-mail: ttf@my.swjtu.edu.cn
  • Supported by:
    The National Key Research and Development Program of China(2024YFC3015402);The National Natural Science Foundation of China(42271446);Tianjin Science and Technology Plan Project(24YFYSHZ00080)

Abstract:

Multi-modal image matching is a crucial foundational task in multi-source remote sensing earth observation. Due to differences in imaging principles, spectral characteristics, temporal phases, and other factors, multi-modal images exhibit significant geometric distortions and nonlinear radiometric differences, which make the expression of common features and matching challenging. Existing methods mostly focus on forced alignment of multi-modal images in the feature space, failing to fully explore cross-modal transformation mapping relationships and lacking comprehensive consideration of interference factors in complex matching scenarios, thus limiting their robustness. To address this, we propose a multi-modal image matching method based on modality reconstruction and feature perturbation learning. First, a cross-modal local common feature expression model is constructed. Positive and negative samples are created using local regions of matched and mismatched points to implement feature contrastive learning. On the basis of supervision by radiometric differences in raw data, a perturbation sample augmentation supervision mechanism is introduced to drive the model to learn interference-resistant feature expressions. Then, a modality reconstruction decoding module is designed to reconstruct local features of one modal image into a pseudo-image of another modality. Additional supervision signals are provided by optimizing the correlation between the reconstructed image and the original image. Finally, through the above multi-objective training, the model can effectively extract local common features invariant to radiometric and geometric variations, thereby achieving accurate multi-modal image matching. Experiments on visible-infrared, visible-SAR, and other modal datasets demonstrate that the proposed method can effectively extract common features resistant to radiometric differences and geometric distortions. It outperforms current state-of-the-art methods in metrics such as reprojection error, success rate, and area under the curve. Moreover, it is verified to be suitable for scenarios with 0°~360° rotation angle differences, providing robust technical support for multi-source remote sensing collaborative tasks.

Key words: image matching, multi-modal image, deep learning, multi-source remote sensing

CLC Number: