Acta Geodaetica et Cartographica Sinica ›› 2026, Vol. 55 ›› Issue (8): 1400-1413.doi: 10.11947/j.AGCS.2026.20260730

• Photogrammetry and Remote Sensing • Previous Articles    

Land use classification from evidence photos via the integration of vision foundation models and graph neural networks

Jianmei Wang1(), Yu Duan1, Shaoming Zhang1(), Xinyan Li2,3   

  1. 1.College of Surveying and Geo-Informatics, Tongji University, Shanghai 200092, China
    2.Surveying and Mapping Institute of Land and Resources Department of Guangdong Province, Guangzhou 510663, China
    3.Key Laboratory of Natural Resources Monitoring in Tropical and Subtropical Area of South China, Ministry of Natural Resources, Guangzhou 510663, China
  • Received:2026-03-19 Revised:2026-08-10 Published:2026-09-09
  • Contact: Shaoming Zhang E-mail:jianmeiw@tongji.edu.cn;zhangshaoming@tongji.edu.cn
  • About author:Wang Jianmei (1971—), female, PhD, associate professor, majors in computer vision and remote sensing, and spatial data mining. E-mail: jianmeiw@tongji.edu.cn
  • Supported by:
    The National Natural Science Foundation of China(42271367);Guangdong Provincial Science and Technology Plan Project(2024A1111120008);Science and Technology Project of Guangdong Provincial Department of Natural Resources(GDZRZYKJ2025003)

Abstract:

Under the “Internet plus verification” framework of national land surveys, the review of land use categories based on evidence photos still heavily relies on manual visual interpretation, resulting in low efficiency and strong subjectivity. Land use categories are determined by functional attributes and spatial organization patterns, while evidence photos of different categories often exhibit similar visual characteristics. Therefore, methods relying solely on visual features remain insufficient for fine-grained land use classification. To address this issue, this paper proposes a land use classification method integrating vision foundation models and graph neural networks. The proposed method employs vision foundation models to extract instance-level semantic features of ground objects, and utilizes multi-view photo-based 3D reconstruction to map cross-view instances into a unified spatial coordinate system. A graph-based representation integrating semantic information and three-dimensional spatial relationships is then constructed to jointly characterize parcel-scale information, and a graph neural network is further employed to perform land use classification. To evaluate the effectiveness of the proposed method, a dedicated evidence-photo dataset for fine-grained construction land categories was constructed from real-world land change verification tasks. Experimental results demonstrate that, compared with the best-performing feature-level fusion ResNet-50 baseline, the proposed method improves the overall accuracy by 9.87 percentage points and increases the macro-F1 score by 11.46 percentage points. In particular, the F1 score of retail commercial land is improved from 25.84% to 61.41%.

Key words: land use classification, evidence photos, vision foundation model, graph neural network

CLC Number: