测绘学报 ›› 2026, Vol. 55 ›› Issue (8): 1425-1438.doi: 10.11947/j.AGCS.2026.20260149

• 摄影测量学与遥感 • 上一篇    

融合相对深度的高分辨率光学卫星遥感影像跨域语义分割框架

孙咏琪1(), 戴晨光1(), 张振超1, 秦进春2, 宿钰1   

  1. 1.信息工程大学地理空间信息学院,河南 郑州 450001
    2.西安测绘研究所,陕西 西安 710054
  • 收稿日期:2026-04-20 修回日期:2026-08-10 发布日期:2026-09-09
  • 通讯作者: 戴晨光 E-mail:sunyq2002@163.com;cgdai2008@163.com
  • 作者简介:孙咏琪(2002—),女,博士生,研究方向为遥感影像智能解译、极地环境智能感知。E-mail:sunyq2002@163.com
  • 基金资助:
    国家自然科学基金(42401502);中国地质大学(武汉)国家地理信息系统工程技术研究中心开放基金(NERCGIS-202402)

A cross-domain semantic segmentation framework fusing relative depth for high-resolution optical satellite remote sensing imagery

Yongqi Sun1(), Chenguang Dai1(), Zhenchao Zhang1, Jinchun Qin2, Yu Su1   

  1. 1.School of Surveying and Mapping, Information Engineering University, Zhengzhou 450001, China
    2.Xi'an Research Institute of Surveying and Mapping, Xi'an 710054, China
  • Received:2026-04-20 Revised:2026-08-10 Published:2026-09-09
  • Contact: Chenguang Dai E-mail:sunyq2002@163.com;cgdai2008@163.com
  • About author:Sun Yongqi (2002—), female, PhD candidate, majors in intelligent interpretation of remote sensing images and polar environment intelligent perception. E-mail: sunyq2002@163.com
  • Supported by:
    The National Natural Science Foundation of China(42401502);Open Fund of National Engineering Research Center of Geographic Information System, China University of Geosciences(NERCGIS-202402)

摘要:

高分辨率光学卫星遥感影像的语义分割是实现大范围多场景地表要素智能识别与动态监测的关键技术,对国土调查、灾害应急、城市规划等应用具有重要意义。然而,传感器、卫星成像视角与区域地理环境变化等因素会引发影像的跨域现象,导致已训练模型在未见域内出现性能显著退化现象。针对当前跨域语义分割方法空间结构信息表征挖掘不足、大量依赖目标域影像完成训练的问题,本文提出一种以相对深度作为空间先验的跨域语义分割框架,利用视觉大模型提取的相对地形起伏作为跨域先验信息,经由多尺度特征自适应融合与边缘监督信号实现跨域逐像素语义类别推断。在公开的卫星遥感影像跨域数据集LoveDA、自建环北极数据集与跨传感器建筑物提取数据配置U2WHU上的试验表明,引入相对深度可有效提升跨域语义分割性能,在未见目标域影像的情况下可接近甚至超过当前先进无监督域适应方法的推理性能,证明了相对深度作为先验信息的必要性、所提框架的有效性及提出框架在全球高分辨率卫星影像智能解译的潜力。

关键词: 高分辨率卫星遥感影像, 跨域, 语义分割, 相对深度, 特征融合

Abstract:

Semantic segmentation of high-resolution optical satellite remote sensing imagery serves as a key technique. It supports the intelligent recognition and dynamic monitoring of large-scale multi-scenario land surface elements. It is of great significance for applications such as land survey, disaster emergency response and urban planning. However, cross-domain discrepancies in imagery can be induced by variations in sensors, satellite imaging angles and regional geographical environments. Such discrepancies lead to significant performance degradation of semantic segmentation models in unseen domains. Existing cross-domain semantic segmentation methods suffer from insufficient exploitation of spatial structure information and heavy reliance on target-domain images for training. To address these issues, this paper proposes a cross-domain semantic segmentation framework with relative depth as spatial prior knowledge. Relative topographic relief extracted by vision foundation models is utilized as cross-domain prior information. Cross-domain pixel-wise semantic category inference is achieved through adaptive multi-scale feature fusion and edge supervision signals. Experiments are conducted on the public satellite remote sensing cross-domain dataset LoveDA, a self-built circumpolar dataset and a cross-sensor building extraction data configuration. Results demonstrate that the introduction of relative depth can effectively improve cross-domain semantic segmentation performance. Without using target-domain images for training, the proposed method can achieve inference performance close to even surpassing that of state-of-the-art unsupervised domain adaptation methods. The necessity of relative depth as prior information, the effectiveness of the proposed framework and its potential in intelligent interpretation of global high-resolution satellite imagery are validated.

Key words: high-resolution satellite remote sensing imagery, cross-domain, semantic segmentation, relative depth, feature fusion

中图分类号: