测绘学报 ›› 2026, Vol. 55 ›› Issue (8): 1343-1356.doi: 10.11947/j.AGCS.2026.20250297

• 影像大地测量前沿技术余智慧防灾创新应用 • 上一篇    

模态重构与特征扰动学习相结合的多模态图像匹配方法

唐腾峰1(), 刘人源1, 刘畅1, 赵彦刚2, 潘丹3, 叶沅鑫1()   

  1. 1.西南交通大学地球科学与工程学院,四川 成都 611756
    2.自然资源部第二地形测量队,陕西 西安 710054
    3.北京微纳星空科技股份有限公司,北京 100089
  • 收稿日期:2025-07-24 修回日期:2025-12-05 发布日期:2026-09-09
  • 通讯作者: 叶沅鑫 E-mail:ttf@my.swjtu.edu.cn;yeyuanxin@home.swjtu.edu.cn
  • 作者简介:唐腾峰(2000—),男,博士生,研究方向为多源遥感图像匹配。E-mail:ttf@my.swjtu.edu.cn
  • 基金资助:
    国家重点研发计划(2024YFC3015402);国家自然科学基金(42271446);天津市科技计划项目(24YFYSHZ00080)

Multi-modal image matching based on modality reconstruction and feature perturbation learning

Tengfeng Tang1(), Renyuan Liu1, Chang Liu1, Yangang Zhao2, Dan Pan3, Yuanxin Ye1()   

  1. 1.Faculty of Geosciences and Engineering, Southwest Jiaotong University, Chengdu 611756, China
    2.The Second Topographic Surveying Brigade of Ministry Natural Resource, Xi'an 710054, China
    3.Beijing Weina Starry Sky Technology Co., Ltd., Beijing 100089, China
  • Received:2025-07-24 Revised:2025-12-05 Published:2026-09-09
  • Contact: Yuanxin Ye E-mail:ttf@my.swjtu.edu.cn;yeyuanxin@home.swjtu.edu.cn
  • About author:Tang Tengfeng (2000—), male, PhD candidate, majors in multi-source remote sensing image matching. E-mail: ttf@my.swjtu.edu.cn
  • Supported by:
    The National Key Research and Development Program of China(2024YFC3015402);The National Natural Science Foundation of China(42271446);Tianjin Science and Technology Plan Project(24YFYSHZ00080)

摘要:

多模态图像匹配是多源遥感对地观测的重要基础工作。由于成像原理、光谱特性和时相等因素的差异,多模态图像间具有显著的几何畸变和非线性辐射差异,导致共性特征表达和匹配困难。现有方法多聚焦于多模态图像在特征空间强制对齐,未充分挖掘跨模态的转换映射关系,且缺乏对复杂匹配场景干扰因素的综合考虑,稳健性受限。为此,本文提出模态重构与特征扰动学习相结合的多模态图像匹配方法。首先,构建跨模态局部共性特征表达模型,通过匹配点与不匹配点的局部区域构造正负样本,实现特征对比学习,并在原始数据辐射差异监督基础上,引入扰动样本增强监督机制,驱动模型学习具备抗干扰的特征表达;然后,设计模态重构解码模块,将一种模态图像的局部特征重构为另一模态的伪图像,通过优化重构图像与原始图像的相关性提供额外监督信号;最后,通过上述多目标训练,模型能够有效提取辐射与几何不变的局部共性特征,进而实现精确的多模态图像匹配。在可见光-红外、可见光-SAR等模态数据集上的试验表明,本文方法能够有效提取抗辐射差异与几何畸变的共性特征,在重投影误差、成功率及曲线下面积等指标上优于当前先进方法,且验证了适用于0°~360°旋转角度差异场景,为多源遥感协同任务提供强稳健性技术支撑。

关键词: 图像匹配, 多模态图像, 深度学习, 多源遥感

Abstract:

Multi-modal image matching is a crucial foundational task in multi-source remote sensing earth observation. Due to differences in imaging principles, spectral characteristics, temporal phases, and other factors, multi-modal images exhibit significant geometric distortions and nonlinear radiometric differences, which make the expression of common features and matching challenging. Existing methods mostly focus on forced alignment of multi-modal images in the feature space, failing to fully explore cross-modal transformation mapping relationships and lacking comprehensive consideration of interference factors in complex matching scenarios, thus limiting their robustness. To address this, we propose a multi-modal image matching method based on modality reconstruction and feature perturbation learning. First, a cross-modal local common feature expression model is constructed. Positive and negative samples are created using local regions of matched and mismatched points to implement feature contrastive learning. On the basis of supervision by radiometric differences in raw data, a perturbation sample augmentation supervision mechanism is introduced to drive the model to learn interference-resistant feature expressions. Then, a modality reconstruction decoding module is designed to reconstruct local features of one modal image into a pseudo-image of another modality. Additional supervision signals are provided by optimizing the correlation between the reconstructed image and the original image. Finally, through the above multi-objective training, the model can effectively extract local common features invariant to radiometric and geometric variations, thereby achieving accurate multi-modal image matching. Experiments on visible-infrared, visible-SAR, and other modal datasets demonstrate that the proposed method can effectively extract common features resistant to radiometric differences and geometric distortions. It outperforms current state-of-the-art methods in metrics such as reprojection error, success rate, and area under the curve. Moreover, it is verified to be suitable for scenarios with 0°~360° rotation angle differences, providing robust technical support for multi-source remote sensing collaborative tasks.

Key words: image matching, multi-modal image, deep learning, multi-source remote sensing

中图分类号: