基于机器视觉的数字摄影测量的新理论新方法

摄影测量与深度学习

  • 龚健雅 ,
  • 季顺平
展开
  • 武汉大学遥感信息工程学院, 湖北 武汉 430079
龚健雅(1957-),男,博士,教授,中国科学院院士,长期从事地理信息理论和几何遥感基础研究。E-mail:gongjy@whu.edu.cn

收稿日期: 2017-11-30

  修回日期: 2018-03-28

  网络出版日期: 2018-06-21

基金资助

国家自然科学基金(41471288)

Photogrammetry and Deep Learning

  • GONG Jianya ,
  • JI Shunping
Expand
  • School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430079, China

Received date: 2017-11-30

  Revised date: 2018-03-28

  Online published: 2018-06-21

Supported by

The National Natural Science Foundation of China (No.41471288)

摘要

深度学习正逐渐占领与“学习”相关的诸多研究领域,也对摄影测量这门学科造成冲击和促进。根据摄影测量学的定义:“利用光学像片研究被摄物体的形状、位置、大小、特性及相互位置关系”,其研究对象包括几何与语义。本文从这两个方面回顾和探讨深度学习目前的应用现状,并对其影响下的摄影测量的发展进行展望。在几何上,基于卷积神经元网络的学习架构已经广泛用于图像匹配、SLAM及三维重建,取得了较好的效果,但仍需进一步改进。在语义上,由于传统的手工设计方法未能将语义信息以工程化的形式确定并生成类似4D产品的各类语义“专题图”,语义部分长期受到忽视。深度学习强大的泛化能力、对任意函数的拟合能力及极高的稳定性,正使得专题图的自动制作成为可能。笔者通过道路网、建筑物、作物分类等应用实例,回顾已经取得的研究成果,并预计:利用光学像片生成高精度的语义专题图,在不远的未来即将实现;并可能成为摄影测量的一类标准产品。最后,针对几何和语义,分别介绍了笔者的两个相关研究:基于深度学习的航空图像匹配以及基于3D卷积神经元网络的精细农作物分类专题图自动提取。

本文引用格式

龚健雅 , 季顺平 . 摄影测量与深度学习[J]. 测绘学报, 2018 , 47(6) : 693 -704 . DOI: 10.11947/j.AGCS.2018.20170640

Abstract

Deep learning has become popular and the mainstream in types of researches related to learning,and has shown its impact on photogrammetry.According to the definition of photogrammetry,a subject that researches shapes,locations,sizes,characteristics and inter-relationships of real objects from optical images,photogrammetry considers two aspects,geometry and semantics.From the two aspects,we review the history of deep learning and discuss its current applications on photogrammetry,and forecast the future development of photogrammetry.In geometry,the deep convolutional neural network (CNN) has been widely applied in stereo matching,SLAM and 3D reconstruction,and has made some effect but needs more improvement.In semantics,conventional empirical and handcrafted methods have failed to extract the semantic information accurately and failed to produce types of “semantic thematic map” as 4D productions (DEM,DOM,DLG,DRG) of photogrammetry,which causes the semantic part of photogrammetry be ignored for a long time.The powerful generalization capacity,ability to fit any functions and stability under types of situations of deep leaning is making the automated production of thematic maps possible.We review the achievements that have been obtained in road network extraction,building detection and crop classification,etc.,and forecast that producing high-accuracy semantic thematic maps directly from optical images will become reality and these maps will become a type of standard products of photogrammetry.At last,we introduce two current researches related to geometry and semantics respectively.One is stereo matching of aerial images based on deep learning and transfer learning; the other is fine crop classification from satellite special-temporal images based on 3D CNN.

参考文献

[1] 龚健雅,季顺平.从摄影测量到计算机视觉[J].武汉大学学报(信息科学版),2017,42(11):1518-1522. GONG Jianya,JI Shunping.From Photogrammetry to Computer Vision[J].Geomatics and Information Science of Wuhan University,2017,42(11):1518-1522.
[2] BOYLE W S,SMITH G E.Charge Coupled Semiconductor Devices[J].The Bell System Technical Journal,1970,49(4):587-593.
[3] ASHBY W R.An Introduction to Cybernetics[M].London:Chapman & Hall Ltd,1961.
[4] FODOR J A,PYLYSHYN Z W.Connectionism and Cognitive Architecture:A Critical Analysis[J].Cognition,1988,28(1-2):3-71.
[5] HINTON G E,OSINDERO S,TEH Y W.A Fast Learning Algorithm for Deep Belief Nets[J].Neural Computation,2006,18(7):1527-1554.
[6] SUYKENS J A K,VANDERWALLE J.Least Squares Support Vector Machine Classifiers[J].Neural Processing Letters,1999,9(3):293-300.
[7] KOLLER D,FRIEDMAN N.Probabilistic Graphical Models:Principles and Techniques[M].Cambridge:MIT Press,2009.
[8] BENGIO Y,LAMBLIN P,POPOVICI D,et al.Greedy Layer-Wise Training of Deep Networks[C]//Proceedings of the 19th International Conference on Neural Information Processing Systems.Canada:ACM,2006:153-160.
[9] KRIZHEVSKY A,SUTSKEVER I,HINTON G E.Imagenet Classification with Deep Convolutional Neural Networks[C]//Proceedings of the 25th International Conference on Neural Information Processing Systems.Lake Tahoe,Nevada:ACM,2012:1097-1105.
[10] MEHTA P,SCHWAB D J.An Exact Mapping between the Variational Renormalization Group and Deep Learning[J].arXiv Preprint arXiv:1410.3831,2014.
[11] TISHBY N,PEREIRA F C,BIALEK W.The Information Bottleneck Method[J].arXiv Preprint arXiv:physics/0004057,2000.
[12] HINTON G,DENG Li,YU Dong,et al.Deep Neural Networks for Acoustic Modeling in Speech Recognition:The Shared Views of Four Research Groups[J].IEEE Signal Processing Magazine,2012,29(6):82-97.
[13] LECUN Y,BOSER B,DENKER J S,et al.Backpropagation Applied to Handwritten Zip Code Recognition[J].Neural Computation,1989,1(4):541-551.
[14] KENDALL A,GRIMES M,CIPOLLA R.Posenet:A Convolutional Network for Real-time 6-dof Camera Relocalization[C]//Proceedings of 2015 IEEE International Conference on Computer Vision.Santiago,Chile:IEEE,2015:2938-2946.
[15] KITTI.The KITTI Vision Benchmark Suite[DB/OL].[2018-03-01].http://www.cvlibs.net/datasets/kitti.
[16] BENGIO Y,COURVILLE A,VINCENT P.Representation Learning:A Review and New Perspectives[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2013,35(8):1798-1828.
[17] NG A.Sparse Autoencoder[R].CS294A Lecture Notes,2011,72(2011):1-19.
[18] SANGER T D.Optimal Unsupervised Learning in A Single-layer Linear Feedforward Neural Network[J].Neural Networks,1989,2(6):459-473.
[19] RUCK D W,ROGERS S K,KABRISKY M,et al.The Multilayer Perceptron as an Approximation to ABayes Optimal Discriminant Function[J].IEEE Transactions on Neural Networks,1990,1(4):296-298.
[20] MIKOLOV T,KARAFIÁT M,BURGET L,et al.Recurrent Neural Network Based Language Model[C]//Proceedings of the 11th Annual Conference of the International Speech Communication Association.Makuhari,Chiba,Japan:International Speech Communication Association,2010,2:3.
[21] MINSKY M L,PAPERT S A.Perceptrons[M].Cambridge:MIT Press,1969.
[22] NAIR V,HINTON G E.Rectified Linear Units Improve Restricted Boltzmann Machines[C]//Proceedings of the 27th International Conference on Machine Learning.Haifa,Israel:ACM,2010:807-814.
[23] SHORE J,JOHNSON R.Axiomatic Derivation of the Principle of Maximum Entropy and the Principle of Minimum Cross-entropy[J].IEEE Transactions on Information Theory,1980,26(1):26-37.
[24] MORÉ J J.The Levenberg-Marquardt Algorithm:Implementation and Theory[M]//WATSON G A.Numerical Analysis.Berlin,Heidelberg:Springer,1978:105-116.
[25] LE CUN Y,BOSER B E,DENKER J S,et al.Handwritten Digit Recognition with a Back-propagation Network[M]//TOURETZKY D S.Advances in Neural Information Processing Systems.San Francisco,CA:Morgan Kaufmann Publishers Inc.,1990:396-404.
[26] GOODFELLOW I,BENGIO Y,COURVILLE A.Deep Learning[M].Cambridge,Massachusetts:MIT Press,2016.
[27] HORN B.Robot Vision[M].Cambridge:MIT Press,1986.
[28] GRAHAM B.Fractional Max-pooling[J].arXiv Preprint arXiv:1412.6071,2014.
[29] ZEILER M D,FERGUS R.Visualizing and Understanding Convolutional Networks[C]//European Conference on Computer Vision.Zurich,Switzerland:Springer,2014:818-833.
[30] SZEGEDY C,LIU W,JIA Y,et al.Going Deeper with Convolutions[J].arXiv Preprint arXiv:1409.4842,2014.
[31] SIMONYAN K,ZISSERMAN A.Very Deep Convolutional Networks for Large-scale Image Recognition[J].arXiv Preprint arXiv:1409.1556,2014.
[32] HE Kaiming,ZHANG Xianyu,REN Shaoqing,et al.Deep Residual Learning for Image Recognition[C]//Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition.Las Vegas,NV:IEEE,2016:770-778.
[33] KENDALL A,CIPOLLA R.Modelling Uncertainty in Deep Learning for Camera Relocalization[C]//Proceedings of 2016 IEEE International Conference on Robotics and Automation.Stockholm,Sweden:IEEE,2016:4762-4769.
[34] ŽBONTAR J,LECUN Y.Computing the Stereo Matching Cost with a Convolutional Neural Network[C]//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition.Boston,MA:IEEE,2015:1592-1599.
[35] KENDALL A,MARTIROSYAN H,DASGUPTA S,et al.End-to-end Learning of Geometry and Context for Deep Stereo Regression[C]//Proceedings of the IEEE Conference on Computer Vision.Venice,Italy:IEEE,2017:66-75.
[36] SEKI A,POLLEFEYS M.SGM-Nets:Semi-global Matching with Neural Networks[C]//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops.Honolulu,HI:IEEE,2017:21-26.
[37] MAYER N,ILG E,HÄUSSER P,et al.A Large Dataset to Train Convolutional Networks for Disparity,Optical Flow,and Scene Flow Estimation[C]//Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition.Las Vegas,NV:IEEE,2016:4040-4048.
[38] LUO Wenjie,SCHWING A G,URTASUN R.Efficient Deep Learning for Stereo Matching[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.Las Vegas,NV:IEEE,2016:5695-5703.
[39] MARR D.Vision:A Computational Investigation into the Human Representation and Processing of Visual Information[M].San Francisco:W.H.Freeman and Company,1982.
[40] CHENG Guangliang,WANG Ying,XU Shibiao,et al.Automatic Road Detection and Centerline Extraction via Cascaded End-to-end Convolutional Neural Network[J].IEEE Transactions on Geoscience and Remote Sensing,2017,55(6):3322-3337.
[41] LI Peikang,ZANG Yu,WANG Cheng,et al.Road Network Extraction via Deep Learning and Line Integral Convolution[C]//Proceedings of the IEEE Conference on Geoscience and Remote Sensing Symposium (IGARSS).Beijing,China:IEEE,2016:1599-1602.
[42] MNIH V,HINTON G E.Learning to Detect Roads in High-resolution Aerial Images[C]//Proceedings of the 11th European Conference on Computer Vision.Heraklion,Crete,Greece:Springer,2010:210-223.
[43] WANG Jun,SONG Jingwei,CHEN Mingquan,et al.Road Network Extraction:A Neural-dynamic Framework Based on Deep Learning and a Finite State Machine[J].International Journal of Remote Sensing,2015,36(12):3144-3169.
[44] PANBOONYUEN T,JITKAJORNWANICH K,LAWAWIROJWONG S,et al.Road Segmentation of Remotely-sensed Images Using Deep Convolutional Neural Networks with Landscape Metrics and Conditional Random Fields[J].Remote Sensing,2017,9(7):680.
[45] VAKALOPOULOU M,KARANTZALOS K,KOMODAKIS N,et al.Building Detection in Very High Resolution Multispectral Data with Deep Learning Features[C]//Proceedings of the IEEE Conference on Geoscience and Remote Sensing Symposium (IGARSS).Milan,Italy:IEEE,2015:1873-1876.
[46] KUSSUL N,LAVRENIUK M,SKAKUN S,et al.Deep Learning Classification of Land Cover and Crop Types Using Remote Sensing Data[J].IEEE Geoscience and Remote Sensing Letters,2017,14(5):778-782.
[47] CASTELLUCCIO M,POGGI G,SANSONE C,et al.Land Use Classification in Remote Sensing Images by Convolutional Neural Networks[J].arXiv Preprint arXiv:1508.00092,2015.
[48] ZHANG Liangpei,ZHANG Lefei,DU Bo.Deep Learning for Remote Sensing Data:A Technical Tutorial on the State of the Art[J].IEEE Geoscience and Remote Sensing Magazine,2016,4(2):22-40.
文章导航

/