Acta Geodaetica et Cartographica Sinica ›› 2026, Vol. 55 ›› Issue (8): 1425-1438.doi: 10.11947/j.AGCS.2026.20260149

• Photogrammetry and Remote Sensing • Previous Articles    

A cross-domain semantic segmentation framework fusing relative depth for high-resolution optical satellite remote sensing imagery

Yongqi Sun1(), Chenguang Dai1(), Zhenchao Zhang1, Jinchun Qin2, Yu Su1   

  1. 1.School of Surveying and Mapping, Information Engineering University, Zhengzhou 450001, China
    2.Xi'an Research Institute of Surveying and Mapping, Xi'an 710054, China
  • Received:2026-04-20 Revised:2026-08-10 Published:2026-09-09
  • Contact: Chenguang Dai E-mail:sunyq2002@163.com;cgdai2008@163.com
  • About author:Sun Yongqi (2002—), female, PhD candidate, majors in intelligent interpretation of remote sensing images and polar environment intelligent perception. E-mail: sunyq2002@163.com
  • Supported by:
    The National Natural Science Foundation of China(42401502);Open Fund of National Engineering Research Center of Geographic Information System, China University of Geosciences(NERCGIS-202402)

Abstract:

Semantic segmentation of high-resolution optical satellite remote sensing imagery serves as a key technique. It supports the intelligent recognition and dynamic monitoring of large-scale multi-scenario land surface elements. It is of great significance for applications such as land survey, disaster emergency response and urban planning. However, cross-domain discrepancies in imagery can be induced by variations in sensors, satellite imaging angles and regional geographical environments. Such discrepancies lead to significant performance degradation of semantic segmentation models in unseen domains. Existing cross-domain semantic segmentation methods suffer from insufficient exploitation of spatial structure information and heavy reliance on target-domain images for training. To address these issues, this paper proposes a cross-domain semantic segmentation framework with relative depth as spatial prior knowledge. Relative topographic relief extracted by vision foundation models is utilized as cross-domain prior information. Cross-domain pixel-wise semantic category inference is achieved through adaptive multi-scale feature fusion and edge supervision signals. Experiments are conducted on the public satellite remote sensing cross-domain dataset LoveDA, a self-built circumpolar dataset and a cross-sensor building extraction data configuration. Results demonstrate that the introduction of relative depth can effectively improve cross-domain semantic segmentation performance. Without using target-domain images for training, the proposed method can achieve inference performance close to even surpassing that of state-of-the-art unsupervised domain adaptation methods. The necessity of relative depth as prior information, the effectiveness of the proposed framework and its potential in intelligent interpretation of global high-resolution satellite imagery are validated.

Key words: high-resolution satellite remote sensing imagery, cross-domain, semantic segmentation, relative depth, feature fusion

CLC Number: