Acta Geodaetica et Cartographica Sinica ›› 2026, Vol. 55 ›› Issue (8): 1414-1424.doi: 10.11947/j.AGCS.2026.20260064

• Photogrammetry and Remote Sensing • Previous Articles    

Multi-scale supervoxel feature aggregation network for three-dimensional point cloud semantic segmentation

Xijiang Chen1(), Minkun Zeng1, Wei Xuan2(), Jingui Zou3, Xianghong Hua3   

  1. 1.School of Safety Science and Emergency Management, Wuhan University of Technology, Wuhan 430070, China
    2.School of Civil Engineering and Architecture, Wuhan University of Technology, Wuhan 430070, China
    3.School of Geodesy and Geomatics, Wuhan University, Wuhan 430070, China
  • Received:2025-11-24 Revised:2026-07-23 Published:2026-09-09
  • Contact: Wei Xuan E-mail:chenxijiang@whu.edu.cn;xuanwei1988@whut.edu.cn
  • About author:Chen Xijiang (1984—), male, PhD, associate professor, majors in point cloud data analysis research. E-mail: chenxijiang@whu.edu.cn
  • Supported by:
    The National Natural Science Foundation of China(42171428; 42271447)

Abstract:

Existing point cloud Transformer module equate spatial proximity with semantic similarity. While valid on continuous surfaces, this assumption often fails at physical boundaries where spatially adjacent points belong to distinct objects, leading to the challenge of boundary feature blurring. This deficiency stems from the inherent limitations of standard KNN grouping. To better extract the semantic information of point clouds, we propose multi-scale supervoxel feature aggregation network (MS-SFA-Net), which injects geometric constraints directly into the encoder. This network replaces simple aggregation with a dual-layer masked attention mechanism. Specifically, it strictly confines intra-supervoxel attention within supervoxels generated by VCCS, and subsequently utilizes inter-supervoxel attention to capture global relationships among supervoxels. On ScanNet v2, the module exhibits significant generalization capability: achieving gains of 1.2% over the Point Transformer V3 baseline and 10.8% over PointNet++, validating its robustness across different backbone networks. Moreover, on the test set of the outdoor Semantic KITTI dataset, our network achieves 1.1% performance improvement over the baseline model, demonstrating its generalization efficacy across diverse application scenarios.

Key words: semantic segmentation, point cloud, supervoxel, deep learning, indoor 3D

CLC Number: