测绘学报 ›› 2026, Vol. 55 ›› Issue (8): 1400-1413.doi: 10.11947/j.AGCS.2026.20260730

• 摄影测量学与遥感 • 上一篇    

融合视觉大模型与图神经网络的举证照片土地利用分类方法

王建梅1(), 段瑜1, 张绍明1(), 李炘妍2,3   

  1. 1.同济大学测绘与地理信息学院,上海 200092
    2.广东省国土资源测绘院,广东 广州 510663
    3.自然资源部华南热带亚热带自然资源监测重点实验室,广东 广州 510663
  • 收稿日期:2026-03-19 修回日期:2026-08-10 发布日期:2026-09-09
  • 通讯作者: 张绍明 E-mail:jianmeiw@tongji.edu.cn;zhangshaoming@tongji.edu.cn
  • 作者简介:王建梅(1971—),女,博士,副教授,研究方向为计算机视觉与遥感、空间数据挖掘。E-mail:jianmeiw@tongji.edu.cn
  • 基金资助:
    国家自然科学基金(42271367);广东省科技计划项目(2024A1111120008);广东省自然资源科技项目(GDZRZYKJ2025003)

Land use classification from evidence photos via the integration of vision foundation models and graph neural networks

Jianmei Wang1(), Yu Duan1, Shaoming Zhang1(), Xinyan Li2,3   

  1. 1.College of Surveying and Geo-Informatics, Tongji University, Shanghai 200092, China
    2.Surveying and Mapping Institute of Land and Resources Department of Guangdong Province, Guangzhou 510663, China
    3.Key Laboratory of Natural Resources Monitoring in Tropical and Subtropical Area of South China, Ministry of Natural Resources, Guangzhou 510663, China
  • Received:2026-03-19 Revised:2026-08-10 Published:2026-09-09
  • Contact: Shaoming Zhang E-mail:jianmeiw@tongji.edu.cn;zhangshaoming@tongji.edu.cn
  • About author:Wang Jianmei (1971—), female, PhD, associate professor, majors in computer vision and remote sensing, and spatial data mining. E-mail: jianmeiw@tongji.edu.cn
  • Supported by:
    The National Natural Science Foundation of China(42271367);Guangdong Provincial Science and Technology Plan Project(2024A1111120008);Science and Technology Project of Guangdong Provincial Department of Natural Resources(GDZRZYKJ2025003)

摘要:

国土调查“互联网+举证”模式下,基于举证照片的土地利用类别审核高度依赖人工目视判读,存在效率低、主观性强等问题。土地利用类别由承载功能及空间组织模式决定,不同类别举证照片常呈现相似视觉表征,仅依赖视觉特征的方法难以实现精细化分类。为此,本文提出一种融合视觉大模型与图神经网络的举证照片土地利用分类方法。该方法利用视觉大模型提取地物实例级语义特征,通过多视角照片三维重建将跨视角实例映射至统一空间坐标系,构建融合语义信息与三维空间关系的图结构以实现图斑级信息联合表征,进而利用图神经网络完成土地利用分类。为验证方法有效性,本文构建了源自国土变更核查真实业务的建设用地精细类别举证照片数据集。试验结果表明,相较于性能最优的特征级融合ResNet-50基线模型,本文方法总体精度提高了9.87个百分点,宏平均F1值提升了11.46个百分点,其中零售商业用地F1值由25.84%提升至61.41%。

关键词: 土地利用分类, 举证照片, 视觉大模型, 图神经网络

Abstract:

Under the “Internet plus verification” framework of national land surveys, the review of land use categories based on evidence photos still heavily relies on manual visual interpretation, resulting in low efficiency and strong subjectivity. Land use categories are determined by functional attributes and spatial organization patterns, while evidence photos of different categories often exhibit similar visual characteristics. Therefore, methods relying solely on visual features remain insufficient for fine-grained land use classification. To address this issue, this paper proposes a land use classification method integrating vision foundation models and graph neural networks. The proposed method employs vision foundation models to extract instance-level semantic features of ground objects, and utilizes multi-view photo-based 3D reconstruction to map cross-view instances into a unified spatial coordinate system. A graph-based representation integrating semantic information and three-dimensional spatial relationships is then constructed to jointly characterize parcel-scale information, and a graph neural network is further employed to perform land use classification. To evaluate the effectiveness of the proposed method, a dedicated evidence-photo dataset for fine-grained construction land categories was constructed from real-world land change verification tasks. Experimental results demonstrate that, compared with the best-performing feature-level fusion ResNet-50 baseline, the proposed method improves the overall accuracy by 9.87 percentage points and increases the macro-F1 score by 11.46 percentage points. In particular, the F1 score of retail commercial land is improved from 25.84% to 61.41%.

Key words: land use classification, evidence photos, vision foundation model, graph neural network

中图分类号: