Acta Geodaetica et Cartographica Sinica ›› 2026, Vol. 55 ›› Issue (8): 1425-1438.doi: 10.11947/j.AGCS.2026.20260149
• Photogrammetry and Remote Sensing • Previous Articles
Yongqi Sun1(
), Chenguang Dai1(
), Zhenchao Zhang1, Jinchun Qin2, Yu Su1
Received:2026-04-20
Revised:2026-08-10
Published:2026-09-09
Contact:
Chenguang Dai
E-mail:sunyq2002@163.com;cgdai2008@163.com
About author:Sun Yongqi (2002—), female, PhD candidate, majors in intelligent interpretation of remote sensing images and polar environment intelligent perception. E-mail: sunyq2002@163.com
Supported by:CLC Number:
Yongqi Sun, Chenguang Dai, Zhenchao Zhang, Jinchun Qin, Yu Su. A cross-domain semantic segmentation framework fusing relative depth for high-resolution optical satellite remote sensing imagery[J]. Acta Geodaetica et Cartographica Sinica, 2026, 55(8): 1425-1438.
Tab. 1
Cross-domain RS image semantic segmentation comparison results from urban to rural of LoveDA dataset"
| 方法 | 类型 | 模态 | IoU | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 背景 | 建筑物 | 道路 | 水系 | 裸地 | 森林 | 农田 | ||||
| CLAN[ | 对抗训练 | 光学 | 22.93 | 44.78 | 25.99 | 46.81 | 10.54 | 37.21 | 24.45 | 30.39 |
| AdaptSegNet[ | 对抗训练 | 26.89 | 40.53 | 30.65 | 50.09 | 16.97 | 32.51 | 28.25 | 32.27 | |
| IAST[ | 自训练 | 29.97 | 49.48 | 28.29 | 64.49 | 2.13 | 33.36 | 61.37 | 38.44 | |
| SegFormer[ | 域泛化 | 26.60 | 55.80 | 35.62 | 65.44 | 17.13 | 40.13 | 39.65 | 40.05 | |
| DCA[ | 自训练 | 36.38 | 55.89 | 40.46 | 62.03 | 22.01 | 38.92 | 60.52 | 45.17 | |
| DAFormer[ | 自训练 | 37.39 | 52.84 | 41.99 | 72.05 | 11.46 | 46.79 | 61.27 | 46.25 | |
| HRDA[ | 自训练 | 36.80 | 64.14 | 40.92 | 71.80 | 14.36 | 47.17 | 68.37 | 49.08 | |
| MIC[ | 自训练 | 39.23 | 61.92 | 43.85 | 72.38 | 13.92 | 47.67 | 68.87 | 49.69 | |
| ST-DASegNet(SegFormer)[ | 对抗训练+自训练 | 36.78 | 59.83 | 43.77 | 73.83 | 19.38 | 49.96 | 67.01 | 50.08 | |
| ConvNeXt-B(基线模型) | 域泛化 | 光学+相对深度 | 35.51 | 56.14 | 40.88 | 70.05 | 13.99 | 48.42 | 64.91 | 47.13 |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度) | 域泛化 | 36.79 | 59.27 | 46.97 | 71.31 | 19.87 | 41.62 | 67.05 | 48.98 | |
| ConvNeXt-XL(光学)+ConvNeXt-B(相对深度) | 域泛化 | 37.23 | 63.05 | 42.69 | 71.71 | 19.91 | 45.03 | 68.36 | 49.71 | |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度)+DF+BAAL | 域泛化 | 38.77 | 65.72 | 39.32 | 72.38 | 16.70 | 48.36 | 68.06 | 49.90 | |
Tab. 2
Cross-domain RS image semantic segmentation comparison results from rural to urban of LoveDA dataset"
| 方法 | 类型 | 模态 | IoU | mIoU | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 背景 | 建筑物 | 道路 | 水系 | 裸地 | 森林 | 农田 | ||||
| CLAN | 对抗训练 | 光学 | 43.41 | 25.42 | 13.75 | 79.25 | 13.71 | 30.44 | 25.80 | 33.11 |
| AdaptSegNet | 对抗训练 | 42.35 | 23.73 | 15.61 | 81.95 | 13.62 | 28.70 | 22.05 | 32.68 | |
| IAST | 自训练 | 48.57 | 31.51 | 28.73 | 86.01 | 20.29 | 31.77 | 36.50 | 40.48 | |
| SegFormer | 域泛化 | 44.98 | 43.93 | 27.46 | 85.82 | 16.24 | 37.02 | 30.48 | 40.85 | |
| DCA | 自训练 | 45.82 | 49.60 | 51.65 | 80.88 | 16.70 | 42.93 | 36.92 | 46.36 | |
| DAFormer | 自训练 | 50.94 | 56.66 | 62.83 | 89.41 | 11.99 | 45.81 | 25.26 | 48.99 | |
| HRDA | 自训练 | 48.25 | 45.24 | 59.16 | 87.17 | 18.82 | 44.94 | 27.61 | 47.31 | |
| MIC | 自训练 | 50.66 | 49.56 | 49.46 | 87.87 | 19.51 | 43.48 | 34.66 | 49.31 | |
| ST-DASegNet(SegFormer) | 对抗训练+自训练 | 51.01 | 54.23 | 60.52 | 87.31 | 15.18 | 47.23 | 36.26 | 50.28 | |
| ConvNeXt-B(基线模型) | 域泛化 | 46.77 | 46.52 | 37.27 | 82.30 | 17.33 | 42.34 | 36.27 | 44.11 | |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度) | 域泛化 | 48.23 | 52.68 | 42.70 | 77.18 | 16.62 | 40.85 | 48.86 | 46.73 | |
| ConvNeXt-XL(光学)+ConvNeXt-B(相对深度) | 域泛化 | 光学+相对深度 | 50.56 | 51.81 | 42.89 | 85.11 | 16.69 | 42.67 | 44.00 | 47.68 |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度)+DF+BAAL | 域泛化 | 48.95 | 52.80 | 52.26 | 84.29 | 18.18 | 44.27 | 46.49 | 49.61 | |
Tab. 3
Cross-domain RS image semantic segmentation comparison results on SUSAN subset test set"
| 方法 | 类型 | 模态 | IoU | mIoU | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 细粒度类别(道路) | 背景 | 粗粒度类别 | ||||||||||
| 跑道 | 滑行道和保险道 | 其他道路 | 植被 | 水系 | 裸地 | 建筑物 | 冰雪 | |||||
| Swin-Base | 域泛化 | 光学 | 62.08 | 11.50 | 34.65 | 37.19 | 89.64 | 33.11 | 51.69 | 44.68 | 42.38 | 45.21 |
| SegFormer(MiT-B5) | 域泛化 | 47.19 | 10.81 | 33.27 | 37.26 | 90.26 | 43.58 | 50.57 | 43.17 | 46.84 | 44.77 | |
| Mask2Former(Swin-Base) | 域泛化 | 49.08 | 18.38 | 37.06 | 31.60 | 89.65 | 50.82 | 51.05 | 46.15 | 45.06 | 46.54 | |
| DAFormer | 自训练 | 光学 | 46.27 | 14.63 | 36.27 | 36.62 | 86.58 | 46.90 | 53.45 | 42.19 | 33.74 | 44.07 |
| HRDA | 自训练 | 55.36 | 15.90 | 33.47 | 32.91 | 90.22 | 44.58 | 52.82 | 39.15 | 47.57 | 45.59 | |
| MIC | 自训练 | 53.86 | 21.25 | 31.85 | 37.44 | 88.45 | 48.96 | 50.96 | 40.82 | 37.73 | 45.70 | |
| ConvNeXt-B(基线模型) | 域泛化 | 光学+相对深度 | 49.71 | 14.72 | 36.10 | 32.57 | 90.36 | 27.38 | 50.19 | 38.61 | 46.91 | 42.95 |
| ConvNeXt-XL | 域泛化 | 59.76 | 9.91 | 27.87 | 35.10 | 90.80 | 48.82 | 52.64 | 36.26 | 48.76 | 45.55 | |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度) | 域泛化 | 67.56 | 17.47 | 35.19 | 34.02 | 89.09 | 29.72 | 51.01 | 48.24 | 42.07 | 46.04 | |
| ConvNeXt-XL(光学)+ConvNeXt-B(相对深度) | 域泛化 | 66.35 | 14.45 | 35.83 | 37.21 | 89.65 | 39.06 | 52.28 | 46.44 | 42.91 | 47.13 | |
| ConvNeXt-B(光学)+ConvNeXt-S(相对深度)+DF+BAAL | 域泛化 | 61.30 | 20.56 | 37.45 | 39.64 | 89.20 | 52.62 | 46.53 | 47.27 | 35.66 | 47.80 | |
| [1] | Wu Kang, Zhang Yingying, Ru Lixiang, et al. A semantic-enhanced multi-modal remote sensing foundation model for Earth observation[J]. Nature Machine Intelligence, 2025, 7(8): 1235-1249. |
| [2] | Li Yansheng, Shi Te, Zhang Yongjun, et al. Learning deep semantic segmentation network under multiple weakly-supervised constraints for cross-domain remote sensing image semantic segmentation[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2021, 175: 20-33. |
| [3] |
沈秭扬, 倪欢, 管海燕. 遥感图像跨域语义分割的无监督域自适应对齐方法[J]. 测绘学报, 2023, 52(12): 2115-2126. DOI: .
doi: 10.11947/j.AGCS.2023.20220483 |
|
Shen Ziyang, Ni Huan, Guan Haiyan. Unsupervised domain adaptation alignment method for cross-domain semantic segmentation of remote sensing images[J]. Acta Geodaetica et Cartographica Sinica, 2023, 52(12): 2115-2126. DOI: .
doi: 10.11947/j.AGCS.2023.20220483 |
|
| [4] | 张继贤, 顾海燕, 杨懿, 等. 高分辨率遥感影像智能解译研究进展与趋势[J]. 遥感学报, 2021, 25(11): 2198-2210. |
| Zhang Jixian, Gu Haiyan, Yang Yi, et al. Research progress and trend of high-resolution remote sensing imagery intelligent interpretation[J]. Journal of Remote Sensing, 2021, 25(11): 2198-2210. | |
| [5] |
顾海燕, 杨懿, 李海涛, 等. 高分辨率遥感影像样本库动态构建与智能解译应用[J]. 测绘学报, 2024, 53(6): 1165-1179. DOI: .
doi: 10.11947/j.AGCS.2024.20230469 |
|
Gu Haiyan, Yang Yi, Li Haitao, et al. Dynamic construction of high-resolution remote sensing image sample datasets and intelligent interpretation applications[J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(6): 1165-1179. DOI: .
doi: 10.11947/j.AGCS.2024.20230469 |
|
| [6] | Li Da, Yang Yongxin, Song Yizhe, et al. Learning to generalize: Meta-learning for domain generalization[C]//Proceedings of 2018 AAAI Conference on Artificial Intelligence. 2018, 32(1): 11596. |
| [7] | Kim D, Yoo Y, Park S, et al. SelfReg: self-supervised contrastive regularization for domain generalization[C]//Proceedings of 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal: IEEE, 2021: 9599-9608. |
| [8] | Pan Xingang, Luo Ping, Shi Jianping, et al. Two at once: enhancing learning and generalization capacities via IBN-net[C]//Proceedings of 2018 Computer Vision. Cham: Springer International Publishing, 2018: 484-500. |
| [9] | Hoyer L, Dai Dengxin, Van Gool L. DAFormer: improving network architectures and training strategies for domain-adaptive semantic segmentation[C]//Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 9914-9925. |
| [10] | Hoyer L, Dai Dengxin, Wang Haoran, et al. MIC: masked image consistency for context-enhanced domain adaptation[C]//Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 11721-11732. |
| [11] | Hoyer L, Dai Dengxin, Van Gool L. Domain adaptive and generalizable network architectures and training strategies for semantic image segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(1): 220-235. |
| [12] | Zhao Qi, Lyu Shuchang, Zhao Hongbo, et al. Self-training guided disentangled adaptation for cross-domain remote sensing image semantic segmentation[J]. International Journal of Applied Earth Observation and Geoinformation, 2024, 127: 103646. |
| [13] | Chen Jie, Zhu Jingru, He Peien, et al. Unsupervised domain adaptation for building extraction of high-resolution remote sensing imagery based on decoupling style and semantic features[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4406917. |
| [14] | Hu Wenshuai, Li Wei, Li Hengchao, et al. Unsupervised domain adaptation with hierarchical masked dual-adversarial network for end-to-end classification of multisource remote sensing data[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 4409917. |
| [15] | Guo Xin, Lao Jiangwei, Dang Bo, et al. SkySense: a multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery[C]//Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 27662-27673. |
| [16] | Li Xuyang, Li Chenyu, Ghamisi P, et al. FlexiMo: a flexible remote sensing foundation model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 5606516. |
| [17] | Li Xuyang, Li Chenyu, Vivone G, et al. SeaMo: a season-aware multimodal foundation model for remote sensing[J]. Information Fusion, 2026, 125: 103334. |
| [18] | Wang Junjue, Zheng Zhuo, Ma Ailong, et al. LoveDA: a remote sensing land-cover dataset for domain adaptive semantic segmentation[PP/OL]. V6. arXiv (2022-05-31) [2026-04-16]. https://doi.org/10.48550/arXiv.2110.08733. |
| [19] | Yang Lihe, Kang Bingyi, Huang Zilong, et al. Depth anything: unleashing the power of large-scale unlabeled data[C]//Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 10371-10381. |
| [20] | Liu Zhuang, Mao Hanzi, Wu Chaoyuan, et al. A ConvNet for the 2020s[C]//Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 11966-11976. |
| [21] | Liu Ze, Lin Yutong, Cao Yue, et al. Swin transformer: hierarchical vision transformer using shifted windows[C]//Proceedings of 2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 9992-10002. |
| [22] | Berman M, Triki A R, Blaschko M B. The lovasz-softmax loss: a tractable surrogate for the optimization of the intersection-over-union measure in neural networks[C]//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 4413-4421. |
| [23] | Ji Shunping, Wei Shiqing, Lu Meng. Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set[J]. IEEE Transactions on Geoscience and Remote Sensing, 2019, 57(1): 574-586. |
| [24] | Luo Yawei, Zheng Liang, Guan Tao, et al. Taking a closer look at domain shift: category-level adversaries for semantics consistent domain adaptation[C]//Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019: 2502-2511. |
| [25] | Tsai Y H, Hung W C, Schulter S, et al. Learning to adapt structured output space for semantic segmentation[C]//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018: 7472-7481. |
| [26] | Mei Ke, Zhu Chuang, Zou Jiaqi, et al. Instance adaptive self-training for unsupervised domain adaptation[C]//Proceedings of 2020 Computer Vision. Cham: Springer International Publishing, 2020: 415-430. |
| [27] | Wu Linshan, Lu Ming, Fang Leyuan. Deep covariance alignment for domain adaptive remote sensing image segmentation[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5620811. |
| [28] | Xie Enze, Wang Wenhai, Yu Zhiding, et al. segFormer: Simple and efficient design for semantic segmentation with transformers[J]. Advances in Neural Information Processing Systems, 2021, 34: 12077-12090. |
| [29] | Xiao Tete, Liu Yingcheng, Zhou Bolei, et al. Unified perceptual parsing for scene understanding[C]//Proceedings of 2018 Computer Vision-ECCV 2018. Cham: Springer International Publishing, 2018: 432-448. |
| [30] | Yang Lihe, Kang Bingyi, Huang Zilong, et al. Depth anything V2[PP/OL]. V2. arXiv (2024-10-20) [2026-04-16]. https://doi.org/10.48550/arXiv.2406.09414. |
| [31] | Ebrahimi S, Arik S O, Nama T, et al. CROME: cross-modal adapters for efficient multimodal LLM[PP/OL]. arXiv (2024-08-13) [2026-04-16]. https://doi.org/10.48550/arXiv.2408.06610. |
| [32] | Zhang Haodi, Yu Anzhu, Gao Kuiliang, et al. M2Caps: learning multi-modal capsules of optical and SAR images for land cover classification[J]. International Journal of Digital Earth, 2025, 18: 2447347. |
| [1] | Xijiang Chen, Minkun Zeng, Wei Xuan, Jingui Zou, Xianghong Hua. Multi-scale supervoxel feature aggregation network for three-dimensional point cloud semantic segmentation [J]. Acta Geodaetica et Cartographica Sinica, 2026, 55(8): 1414-1424. |
| [2] | Zejiao WANG, Longgang XIANG, Meng WANG, Xingjuan WANG, Qing LIU. Hierarchical feature and diversified attention fusion network for collaborative extraction of road surface and centerline [J]. Acta Geodaetica et Cartographica Sinica, 2026, 55(3): 548-563. |
| [3] | Wenjun HUANG, Qun SUN, Qing XU, Long FAN, Anzhu YU, Fubing ZHANG. A global coastal DEM super-resolution reconstruction method integrating frequency-domain features and topographic priors [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(8): 1518-1531. |
| [4] | Jie WAN, Zhong XIE, Yongyang XU, Liufeng TAO. A U-shaped graph convolution network method for semantic segmentation of vehicle LiDAR point clouds towards urban road scenes [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(7): 1280-1293. |
| [5] | Yiming ZHAO, Kelin HU, Kelong TU, Yaxian QING, Chao YANG, Kunlun QI, Huayi WU. Multi-label scene classification method based on fusion of SAR and optical remote sensing images [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(5): 911-923. |
| [6] | Yungang CAO, Peng YANG, Jiangbo GONG, Gao ZHU, Xingyu SHEN. A road extraction method integrating spatial-relation enhancement and heterogeneous feature fusion [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(12): 2219-2232. |
| [7] | Liangxiong GONG, Xinghua LI, Yuanming CHENG, Xingyou ZHAO, Renping XIE, Honggen WANG. A lightweight remote sensing images change detection network utilizing spatio-temporal difference enhancement and adaptive feature fusion [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(1): 136-153. |
| [8] | Jianbo TANG, Zhiyuan HU, Ju PENG, Heyan XIA, Junjie DING, Yuyu ZHANG, Xiaoming MEI. A road intersection recognition method in crowdsourced trajectory data by fusing visual features and motion features [J]. Acta Geodaetica et Cartographica Sinica, 2025, 54(1): 182-193. |
| [9] | Bo HU, Hanxin CHEN, Song REN, Yinghao QU, Qingyi LIU, Xinyue TU, Datao WANG. A post-processing algorithm for automatic recognition of tunnel crack diseases based on segmentation masks [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(9): 1715-1724. |
| [10] | Fubing ZHANG, Qun SUN, Jingzhen MA, Shijie SUN, Bowei WEN. An intelligent classification method for building shape based on fusion of global and local features [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(9): 1842-1852. |
| [11] | Xin YAN, Li SHEN, Junjie PAN, Yanshuai DAI, Jicheng WANG, Xiaoli ZHENG, Zhi-lin LI. Weakly supervised building change detection integrating multi-scale feature fusion and spatial refinement for high resolution remote sensing images [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(8): 1586-1597. |
| [12] | Tao XU, Yuanwei YANG, Xianjun GAO, Zhiwei WANG, Yue PAN, Shaohua LI, Lei XU, Yanjun WANG, Bo LIU, Jing YU, Fengmin WU, Haoyu SUN. Integrated graph convolution and multi-scale features for the overhead catenary system point cloud semantic segmentation [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(8): 1624-1633. |
| [13] | LIN Yunhao, WANG Yanjun, LI Shaochun, CAI Hengfan. A coupled DeepLab and Transformer approach for fine classification of crop cultivation types in remote sensing [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(2): 353-366. |
| [14] | ZHANG Caili, XIANG Longgang, LI Yali, GAO Songfeng, PAN Chuanjiao. Road section navigation attribute mining [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(2): 367-378. |
| [15] | Genyun SUN, Chao SUN, Aizhu ZHANG. Road extraction networks fusing multiscale and edge features [J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(12): 2233-2243. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||