
测绘学报 ›› 2026, Vol. 55 ›› Issue (8): 1452-1464.doi: 10.11947/j.AGCS.2026.20260063
• 摄影测量学与遥感 • 上一篇
赵翔宇1(
), 张春菊1(
), 裴一帆1, 李晨曦1, 张俊2, 徐薇3,4, 兰春5, 吕临峰1, 梁宏博1
收稿日期:2025-11-24
修回日期:2026-07-23
发布日期:2026-09-09
通讯作者:
张春菊
E-mail:2023110780@mail.hfut.edu.cn;zcjtwz@sina.com
作者简介:赵翔宇(2002—),男,硕士,主要从事遥感影像智能解译研究。E-mail:2023110780@mail.hfut.edu.cn
基金资助:
Xiangyu Zhao1(
), Chunju Zhang1(
), Yifan Pei1, Chenxi Li1, Jun Zhang2, Wei Xu3,4, Chun Lan5, Linfeng Lü1, Hongbo Liang1
Received:2025-11-24
Revised:2026-07-23
Published:2026-09-09
Contact:
Chunju Zhang
E-mail:2023110780@mail.hfut.edu.cn;zcjtwz@sina.com
About author:Zhao Xiangyu (2002—), male, master, majors in intelligent interpretation of remote sensing images. E-mail: 2023110780@mail.hfut.edu.cn
Supported by:摘要:
针对高分辨率遥感语义分割中尺度变化显著、密集小目标与细边界易受纹理干扰、长尾分布导致稀有类别学习不足等问题,本文提出了一种基于分割一切模型(SAM)的端到端框架(SCA-SAM)。该方法在编码器多阶段嵌入尺度上下文注意力(SCA)模块,通过逐层累积多尺度特征信息,强化小目标与复杂边界的特征表达;同时引入低秩适配(LoRA)在注意力映射层面开展参数高效微调,仅更新少量参数即可实现对遥感影像纹理与空间结构差异的高效迁移。设计面向类别不平衡与易混淆的复合损失,进一步提升稀有类别与边界区域的判别能力。在UAVid数据集、ISPRS Vaihingen数据集与ISPRS Potsdam数据集上的试验证明,SCA-SAM在各类别量化指标与整体视觉效果上均取得稳定提升,并在较低参数增量下表现出更优的分割精度与更强泛化稳定性。
中图分类号:
赵翔宇, 张春菊, 裴一帆, 李晨曦, 张俊, 徐薇, 兰春, 吕临峰, 梁宏博. SCA-SAM:基于尺度上下文注意力与SAM高效迁移的遥感影像语义分割方法[J]. 测绘学报, 2026, 55(8): 1452-1464.
Xiangyu Zhao, Chunju Zhang, Yifan Pei, Chenxi Li, Jun Zhang, Wei Xu, Chun Lan, Linfeng Lü, Hongbo Liang. SCA-SAM: semantic segmentation for remote sensing images based on scale context attention and efficient SAM transfer[J]. Acta Geodaetica et Cartographica Sinica, 2026, 55(8): 1452-1464.
表1
UAVid验证集定量对比结果"
| 模型 | IoU | mIoU | F1值 | OA | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 建筑 | 道路 | 树木 | 灌丛 | 移动车辆 | 静止车辆 | 人 | 杂物 | ||||
| UNetFormer | 90.06 | 78.13 | 76.01 | 66.69 | 68.61 | 66.64 | 43.22 | 62.01 | 68.92 | 80.89 | 86.67 |
| SegFormer | 85.20 | 69.55 | 64.76 | 57.02 | 59.64 | 52.80 | 34.66 | 49.36 | 59.12 | 73.33 | 80.31 |
| DCSwin | 90.12 | 78.10 | 76.49 | 66.28 | 75.32 | 69.08 | 45.79 | 62.86 | 70.50 | 82.07 | 86.98 |
| CMTFNet | 91.70 | 80.74 | 78.24 | 69.12 | 75.72 | 69.42 | 43.59 | 65.48 | 71.75 | 82.82 | 88.27 |
| SAM(Frozen) | 88.94 | 74.06 | 71.94 | 56.60 | 71.12 | 62.10 | 40.04 | 56.03 | 65.11 | 78.00 | 83.84 |
| EfficientViTSAM | 90.41 | 76.47 | 77.51 | 68.14 | 69.01 | 65.64 | 37.69 | 62.73 | 68.45 | 80.35 | 87.10 |
| SAM-LST | 88.68 | 75.17 | 75.93 | 66.16 | 70.18 | 66.13 | 45.83 | 59.71 | 68.47 | 80.69 | 85.85 |
| RSAM-Seg | 92.12 | 78.98 | 78.11 | 68.88 | 72.19 | 67.90 | 42.31 | 65.48 | 70.75 | 82.10 | 88.15 |
| SCA-SAM | 92.56 | 79.80 | 79.14 | 71.96 | 76.62 | 72.42 | 49.24 | 65.86 | 73.45 | 84.14 | 88.83 |
表2
ISPRS Vaihingen验证集定量对比结果"
| 模型 | IoU | mIoU | F1值 | OA | ||||
|---|---|---|---|---|---|---|---|---|
| 不透水表面 | 建筑 | 灌丛 | 树木 | 车辆 | ||||
| UNetFormer | 88.43 | 92.97 | 74.29 | 82.06 | 81.38 | 83.83 | 91.07 | 91.86 |
| SegFormer | 88.95 | 92.90 | 74.85 | 82.28 | 81.01 | 84.00 | 91.17 | 92.04 |
| DCSwin | 87.21 | 90.39 | 73.92 | 82.45 | 80.72 | 82.94 | 90.57 | 91.23 |
| CMTFNet | 88.59 | 93.43 | 75.05 | 83.48 | 81.96 | 84.50 | 91.48 | 92.27 |
| SAM(Frozen) | 78.11 | 82.83 | 62.58 | 76.47 | 61.63 | 72.32 | 83.65 | 85.72 |
| EfficientViTSAM | 86.31 | 91.73 | 73.55 | 81.15 | 73.81 | 81.31 | 89.53 | 91.05 |
| SAM-LST | 86.44 | 91.04 | 73.96 | 82.71 | 76.37 | 82.10 | 90.04 | 91.20 |
| RSAM-Seg | 88.37 | 92.27 | 74.62 | 83.01 | 76.22 | 82.90 | 90.50 | 91.82 |
| SCA-SAM | 89.27 | 93.33 | 75.54 | 83.80 | 81.07 | 84.60 | 91.53 | 92.45 |
表3
ISPRS Potsdam验证集定量对比结果"
| 模型 | IoU | mIoU | F1值 | OA | ||||
|---|---|---|---|---|---|---|---|---|
| 不透水表面 | 建筑 | 灌丛 | 树木 | 车辆 | ||||
| UNetFormer | 84.30 | 90.73 | 76.38 | 79.02 | 90.77 | 84.24 | 91.33 | 89.54 |
| SegFormer | 84.54 | 90.57 | 74.47 | 76.51 | 88.07 | 82.83 | 90.48 | 89.02 |
| DCSwin | 85.17 | 91.35 | 76.25 | 77.41 | 91.05 | 84.25 | 91.32 | 89.72 |
| CMTFNet | 86.84 | 93.39 | 77.47 | 79.03 | 90.57 | 85.46 | 92.04 | 90.74 |
| SAM(Frozen) | 71.17 | 76.52 | 61.51 | 62.46 | 79.90 | 70.31 | 82.34 | 80.06 |
| EfficientViTSAM | 84.78 | 91.70 | 75.89 | 79.02 | 87.24 | 83.73 | 91.04 | 89.78 |
| SAM-LST | 82.59 | 87.65 | 74.72 | 76.90 | 90.49 | 82.47 | 90.27 | 88.26 |
| RSAM-Seg | 84.30 | 90.36 | 74.72 | 77.43 | 89.00 | 83.16 | 90.68 | 89.03 |
| SCA-SAM | 86.92 | 93.46 | 77.92 | 79.11 | 92.02 | 85.89 | 92.28 | 90.81 |
表4
模型消融试验分析"
| 试验组 | 训练策略 | 可训参数量 | 吞吐量(张/s) | 总参数量 | mIoU/(%) | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| ViT | SCA模块 | 解码器 | 总可训 | UAVid | ISPRS Vaihingen | ISPRS Potsdam | ||||
| SCA-SAM | LoRA | 0.15×106 | 1.48×106 | 4.34×106 | 5.97×106 | 11.89 | 114.23×106 | 73.45 | 84.60 | 85.89 |
| w/o SCA | LoRA | 0.15×106 | 0×106 | 4.34×106 | 4.49×106 | 13.66 | 94.16×106 | 71.99 | 82.61 | 84.12 |
| Full Train | 全训练 | 28.39×106 | 1.48×106 | 4.34×106 | 34.22×106 | 8.99 | 114.23×106 | 75.05 | 84.63 | 86.34 |
| Full Freeze | 全冻结 | 0.00×106 | 1.48×106 | 4.34×106 | 5.82×106 | 12.01 | 114.23×106 | 29.15 | 82.01 | 84.27 |
表5
损失消融试验分析"
| 损失组合 | CB-FocalCE | Focal-Tversky | TopK-CE | IoU | mIoU | F1值 | OA | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 建筑 | 道路 | 树木 | 灌丛 | 移动车辆 | 静止车辆 | 人 | 杂物 | |||||||
| 1 | √ | √ | √ | 92.56 | 79.80 | 79.14 | 71.96 | 76.62 | 72.42 | 49.24 | 65.86 | 73.45 | 84.14 | 88.83 |
| 2 | √ | √ | 92.14 | 80.98 | 78.57 | 69.58 | 75.47 | 70.25 | 43.77 | 66.82 | 72.20 | 83.13 | 88.61 | |
| 3 | √ | √ | 92.75 | 81.92 | 78.47 | 69.60 | 74.55 | 69.40 | 42.22 | 68.11 | 72.13 | 83.01 | 88.89 | |
| 4 | √ | √ | 92.48 | 80.70 | 77.78 | 69.49 | 73.24 | 68.77 | 43.92 | 66.96 | 71.67 | 82.78 | 88.47 | |
| [1] |
陈超, 梁锦涛, 杨刚, 等. 面向土地覆盖精准分类的遥感特征参数优选方法[J]. 测绘学报, 2024, 53(7): 1401-1416. DOI: .
doi: 10.11947/j.AGCS.2024.20230327 |
|
Chen Chao, Liang Jintao, Yang Gang, et al. Remote sensing parameters optimization for accurate land cover classifi-cation[J]. Acta Geodaetica et Cartographica Sinica, 2024, 53(7): 1401-1416. DOI: .
doi: 10.11947/j.AGCS.2024.20230327 |
|
| [2] | 蒋伟杰, 张春菊, 徐兵, 等. AED-Net:滑坡灾害遥感影像语义分割模型[J]. 地球信息科学学报, 2023, 25(10): 2012-2025. |
| Jiang Weijie, Zhang Chunju, Xu Bing, et al. AED-Net: semantic segmentation model for landslide recognition from remote sensing images[J]. Journal of Geo-Information Science, 2023, 25(10): 2012-2025. | |
| [3] | Wu Kang, Zhang Yingying, Ru Lixiang, et al. A semantic-enhanced multi-modal remote sensing foundation model for Earth observation[J]. Nature Machine Intelligence, 2025, 7(8): 1235-1249. |
| [4] | 杨明旺, 赵丽科, 叶林峰, 等. 基于卷积神经网络的遥感影像建筑物提取方法综述[J]. 地球信息科学学报, 2024, 26(6): 1500-1516. |
| Yang Mingwang, Zhao Like, Ye Linfeng, et al. A review of convolutional neural networks related methods for building extraction from remote sensing images[J]. Journal of Geo-Information Science, 2024, 26(6): 1500-1516. | |
| [5] | Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston: IEEE, 2015: 3431-3440. |
| [6] |
伍广明, 陈奇, Shibasaki R, 等. 基于U型卷积神经网络的航空影像建筑物检测[J]. 测绘学报, 2018, 47(6): 864-872. DOI: .
doi: 10.11947/j.AGCS.2018.20170651 |
|
Wu Guangming, Chen Qi, Shibasaki R, et al. High precision building detection from aerial imagery using a U-Net like convolutional architecture[J]. Acta Geodaetica et Cartographica Sinica, 2018, 47(6): 864-872. DOI: .
doi: 10.11947/j.AGCS.2018.20170651 |
|
| [7] | 许泽宇, 沈占锋, 李杨, 等. 增强型DeepLab算法和自适应损失函数的高分辨率遥感影像分类[J]. 遥感学报, 2022, 26(2): 406-415. |
| Xu Zeyu, Shen Zhanfeng, Li Yang, et al. Classification of high-resolution remote sensing images based on enhanced DeepLab algorithm and adaptive loss function[J]. National Remote Sensing Bulletin, 2022, 26(2): 406-415. | |
| [8] | 柴华彬, 严超, 邹友峰, 等. 利用PSP Net实现湖北省遥感影像土地覆盖分类[J]. 武汉大学学报(信息科学版), 2021, 46(8): 1224-1232. |
| Chai Huabin, Yan Chao, Zou Youfeng, et al. Land cover classification of remote sensing image of Hubei province by using PSP Net[J]. Geomatics and Information Science of Wuhan University, 2021, 46(8): 1224-1232. | |
| [9] | Dosovitskiy A. An image is worth 16x16 words: transformers for image recognition at scale[C]//Proceedings of 2021 International Conference on Learning Representations. San Diego: Open Review.net, 2021. |
| [10] | Chen Wei, Bruzzone L, Dang Bo, et al. REST: holistic learning for end-to-end semantic segmentation of whole-scene remote sensing imagery[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(1): 693-710. |
| [11] | Li Yansheng, Wu Yuning, Cheng Gong, et al. MEET: a million-scale dataset for fine-grained geospatial scene classification with zoom-free remote sensing imagery[J]. IEEE/CAA Journal of Automatica Sinica, 2025, 12(5): 1004-1023. |
| [12] | Kirillov A, Mintun E, Ravi N, et al. Segment anything[C]//Proceedings of 2023 IEEE/CVF International Conference on Computer Vision (ICCV). Paris: IEEE, 2023: 4015-4026. |
| [13] | Wang Di, Zhang Jing, Du Bo, et al. SAMRS: scaling-up remote sensing segmentation dataset with segment anything model[C]//Proceedings of 2023 Advances in Neural Information Processing Systems 36. New Orleans: NeurIPS, 2023: 8815-8827. |
| [14] | Yan Zhiyuan, Li Junxi, Li Xuexue, et al. RingMo-SAM: a foundation model for segment anything in multimodal remote-sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5625716. |
| [15] | Liu Zhuoran, Li Zizhen, Liang Ying, et al. RSPS-SAM: a remote sensing image panoptic segmentation method based on SAM[J]. Remote Sensing, 2024, 16(21): 4002. |
| [16] | Zheng Linghao, Pu Xinyang, Zhang Su, et al. Tuning a SAM-based model with multicognitive visual adapter to remote sensing instance segmentation[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 2737-2748. |
| [17] | 刘思涌, 赵毅力. 微调SAM的遥感图像高效语义分割模型DP-SAM[J]. 中国图象图形学报, 2025, 30(8): 2884-2896. |
| Liu Siyong, Zhao Yili. DP-SAM: efficient semantic segmentation of remote sensing images by fine-tuning SAM[J]. Journal of Image and Graphics, 2025, 30(8): 2884-2896. | |
| [18] | Lu Xiaoyan, Weng Qihao. Multi-LoRA fine-tuned segment anything model for urban man-made object extraction[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5637519. |
| [19] | Shan Zhe, Liu Yang, Zhou Lei, et al. ROS-SAM: high-quality interactive segmentation for remote sensing moving object[C]//Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025: 3625-3635. |
| [20] | Hu E J, Shen Yelong, Wallis P, et al. LoRA: low-rank adaptation of large language models[C]//Proceedings of 2021 International Conference on Learning Representations. San Diego: Open Review.net, 2021. |
| [21] | Lü Ye, Vosselman George, Xia Guisong, et al. UAVid: a semantic segmentation dataset for UAV imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 165: 108-119. |
| [22] | Gerke M, Rottensteiner F, Wegner J D, et al. ISPRS semantic labeling contest[C]//Proceedings of 2014 Photogrammetric Computer Vision. Zurich: [s.n.], 2014. |
| [23] | Li Sirui, Peng Linkai, Zhang Zheyuan, et al. TAGS: 3D tumor-adaptive guidance for SAM[PP/OL]. V2. arXiv (2025-08-27) [2026-01-10]. https://doi.org/10.48550/arXiv.2505.17096. |
| [24] | 罗健伟, 张银胜. DMFPNet:增强多尺度目标感知的双路径高分辨率遥感图像分割算法[J]. 地球信息科学学报, 2025, 27(5): 1195-1213. |
| Luo Jianwei, Zhang Yinsheng. DMFPNet: dual-path high-resolution remote sensing image segmentation algorithm for enhanced multiscale target detection[J]. Journal of Geo-Information Science, 2025, 27(5): 1195-1213. | |
| [25] |
刘帅, 李笑迎, 于梦, 等. 高分辨率遥感图像双解耦语义分割网络模型[J]. 测绘学报, 2023, 52(4): 638-647. DOI: .
doi: 10.11947/j.AGCS.2023.20210455 |
|
Liu Shuai, Li Xiaoying, Yu Meng, et al. Dual decoupling semantic segmentation model for high-resolution remote sensing images[J]. Acta Geodaetica et Cartographica Sinica, 2023, 52(4): 638-647. DOI: .
doi: 10.11947/j.AGCS.2023.20210455 |
|
| [26] | Wang Libo, Li Rui, Zhang Ce, et al. UNetFormer: a UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 190: 196-214. |
| [27] | Xie Enze, Wang Wenhai, Yu Zhiding, et al. SegFormer: simple and efficient design for semantic segmentation with transformers[C]//Proceedings of the 35th Conference on Neural Information Processing Systems. Red Hook: Curran Associates, 2021: 12077-12090. |
| [28] | Wang Libo, Li Rui, Duan Chenxi, et al. A novel transformer based semantic segmentation scheme for fine-resolution remote sensing images[J]. IEEE Geoscience and Remote Sensing Letters, 2022, 19: 6506105. |
| [29] | Wu Honglin, Huang Peng, Zhang Min, et al. CMTFNet: CNN and multiscale transformer fusion network for remote-sensing image semantic segmentation[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 2004612. |
| [30] | Zhang Zhuoyang, Cai Han, Han Song. EfficientViT-SAM: accelerated segment anything model without performance loss[C]//Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Seattle: IEEE, 2024: 7859-7863. |
| [31] | Chai Shurong, Jain R K, Teng Shiyu, et al. Ladder fine-tuning approach for SAM integrating complementary network[J]. Procedia Computer Science, 2024, 246: 4951-4958. |
| [32] | Zhang Jie, Li Yunxin, Yang Xubing, et al. RSAM-SEG: a SAM-based model with prior knowledge integration for remote sensing image semantic segmentation[J]. Remote Sensing, 2025, 17(4): 590. |
| [1] | 龚俊, 罗建军, 柯胜男, 黎施彬, 陈蔚聪, 秦乐, 汤圣君. 一种自适应剪枝的室内场景三维高斯SLAM实时高保真建模方法[J]. 测绘学报, 2026, 55(7): 1266-1277. |
| [2] | 谭清华, 齐恒, 施泓羽, 唐炉亮, 阚子涵, 杨红, 孙乐乐, 刘亚非, 顾正雄. 城市场景多源低空风险量化与航路生成[J]. 测绘学报, 2026, 55(7): 1306-1320. |
| [3] | 麻源源. 南极冰盖/冰架典型区域表面流速时空变化研究[J]. 测绘学报, 2026, 55(7): 1326-1326. |
| [4] | 有泽, 王丽英, 宇翼巍, 秦志伟, 谢春喜, 耿一末, 李鑫奥. 基于RMLS点云的盾构隧道管片接缝多测度交互识别指标[J]. 测绘学报, 2026, 55(6): 1101-1115. |
| [5] | 刘硕, 孙海丽, 钟若飞. 多尺度边缘信息融合的隧道渗漏水检测方法[J]. 测绘学报, 2026, 55(6): 1116-1127. |
| [6] | 符茵. 基于多源遥感影像光流估计的贡嘎山冰川流速场建模与动态演化分析[J]. 测绘学报, 2026, 55(6): 1134-1134. |
| [7] | 张波. 贡嘎山冰川冰湖动态演化SAR遥感监测与分析[J]. 测绘学报, 2026, 55(6): 1135-1135. |
| [8] | 翟若明. 基于点云信息的室内三维重建关键技术研究[J]. 测绘学报, 2026, 55(6): 1136-1136. |
| [9] | 吴楠. 崇明岛滨海湿地的遥感植被分类与生物量时空演变研究[J]. 测绘学报, 2026, 55(6): 1140-1140. |
| [10] | 韦朋成, 蒋贵宇, 沈航毅, 黄海峰, 张溶玲. 高精度激光点云配准驱动的毫米级地表形变检测方法[J]. 测绘学报, 2026, 55(5): 866-880. |
| [11] | 林雨准, 王淑香, 芮杰, 金飞, 姜建芳, 左溪冰, 刘潇, 邹毓杰. 一种稀疏标签优化的异构数据道路提取方法[J]. 测绘学报, 2026, 55(5): 881-893. |
| [12] | 时天东, 赵玲, 赵文豪, 齐霁, 崔浩, 彭程里, 张新长. 时空信息显式引导的高分光学遥感影像可控生成方法[J]. 测绘学报, 2026, 55(5): 894-908. |
| [13] | 王政文, 杨俊涛, 康志忠, 张宇涛, 王旭哲, 张雪. 三维点云语义辅助的多楼层室内空间拓扑模型构建方法[J]. 测绘学报, 2026, 55(5): 909-926. |
| [14] | 尉锐, 李杰, 刘汇慧, 吴美茹, 林镠鹏, 袁强强, 郑莉. 面向洪水灾害的视觉-文本协同表征的异质遥感变化检测方法[J]. 测绘学报, 2026, 55(5): 927-940. |
| [15] | 曲英杰. 基于神经辐射场的多时相卫星影像地表面三维重建研究[J]. 测绘学报, 2026, 55(5): 941-941. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||