测绘学报 ›› 2026, Vol. 55 ›› Issue (8): 1414-1424.doi: 10.11947/j.AGCS.2026.20260064

• 摄影测量学与遥感 • 上一篇    

多尺度超体素特征聚合网络的三维点云语义分割

陈西江1(), 曾敏坤1, 宣伟2(), 邹进贵3, 花向红3   

  1. 1.武汉理工大学安全科学与应急管理学院,湖北 武汉 430070
    2.武汉理工大学土木工程与建筑学院,湖北 武汉 430070
    3.武汉大学测绘学院,湖北 武汉 430070
  • 收稿日期:2025-11-24 修回日期:2026-07-23 发布日期:2026-09-09
  • 通讯作者: 宣伟 E-mail:chenxijiang@whu.edu.cn;xuanwei1988@whut.edu.cn
  • 作者简介:陈西江(1984—),男,博士,副教授,研究方向为点云数据分析。E-mail:chenxijiang@whu.edu.cn
  • 基金资助:
    国家自然科学基金(42171428; 42271447)

Multi-scale supervoxel feature aggregation network for three-dimensional point cloud semantic segmentation

Xijiang Chen1(), Minkun Zeng1, Wei Xuan2(), Jingui Zou3, Xianghong Hua3   

  1. 1.School of Safety Science and Emergency Management, Wuhan University of Technology, Wuhan 430070, China
    2.School of Civil Engineering and Architecture, Wuhan University of Technology, Wuhan 430070, China
    3.School of Geodesy and Geomatics, Wuhan University, Wuhan 430070, China
  • Received:2025-11-24 Revised:2026-07-23 Published:2026-09-09
  • Contact: Wei Xuan E-mail:chenxijiang@whu.edu.cn;xuanwei1988@whut.edu.cn
  • About author:Chen Xijiang (1984—), male, PhD, associate professor, majors in point cloud data analysis research. E-mail: chenxijiang@whu.edu.cn
  • Supported by:
    The National Natural Science Foundation of China(42171428; 42271447)

摘要:

现有的点云Transformer模块将空间邻近性等同于语义相似性。虽然这种假设在连续表面上有效,但在物理边界处,空间相邻的点往往属于不同的对象,导致模型面临边界特征模糊的挑战。这一缺陷源于标准KNN分组的局限性。为了更好地提取点云的语义信息,本文提出了一种多尺度超体素特征聚合网络(MS-SFA-Net),将几何约束直接注入编码器。该网络利用双层掩码注意力机制取代了简单的聚合。通过将超体素内注意力严格限制在VCCS生成的超体素内部,再利用超体素间注意力捕捉超体素之间的全局关系。在ScanNet v2上,该模块展现了显著的泛化能力:在Point Transformer V3基线上提升了1.2%,在PointNet++上提升了10.8%,验证了其在不同主干网络中的稳健性。更进一步地,在室外数据集Semantic KITTI测试集上,该网络相较于基线模型实现了1.1%的性能提升,证明了其应用场景的泛化效能。

关键词: 语义分割, 点云, 超体素, 深度学习, 室内三维

Abstract:

Existing point cloud Transformer module equate spatial proximity with semantic similarity. While valid on continuous surfaces, this assumption often fails at physical boundaries where spatially adjacent points belong to distinct objects, leading to the challenge of boundary feature blurring. This deficiency stems from the inherent limitations of standard KNN grouping. To better extract the semantic information of point clouds, we propose multi-scale supervoxel feature aggregation network (MS-SFA-Net), which injects geometric constraints directly into the encoder. This network replaces simple aggregation with a dual-layer masked attention mechanism. Specifically, it strictly confines intra-supervoxel attention within supervoxels generated by VCCS, and subsequently utilizes inter-supervoxel attention to capture global relationships among supervoxels. On ScanNet v2, the module exhibits significant generalization capability: achieving gains of 1.2% over the Point Transformer V3 baseline and 10.8% over PointNet++, validating its robustness across different backbone networks. Moreover, on the test set of the outdoor Semantic KITTI dataset, our network achieves 1.1% performance improvement over the baseline model, demonstrating its generalization efficacy across diverse application scenarios.

Key words: semantic segmentation, point cloud, supervoxel, deep learning, indoor 3D

中图分类号: