Multi-modal Semantic Labeling

Overview

Pixel-wise semantic labeling that fuses RGB imagery with depth, using encoder–decoder convolutional architectures. Depth carries geometric structure that colour alone does not, and fusing the two modalities improves labeling of structural elements in buildings.

Encoder-decoder CNN for multi-modal semantic labeling

Related publications

  • Multimodal Semantic Segmentation: Fusion of RGB and Depth Data in Convolutional Neural Networks (book chapter), 2019
  • Semantic Labeling of Structural Elements in Buildings by Fusing RGB and Depth Images in an Encoder-Decoder CNN Framework, ISPRS 2018