Panoptic-DeepLab

Bowen Cheng, Maxwell D. Collins, Yukun Zhu, Ting Liu, Thomas S. Huang, Hartwig Adam, Liang-Chieh Chen

Introduction

Our bottom-up Panoptic-DeepLab is conceptually simple and delivers state-of-the-art panoptic segmentation results kirillov2018panoptic. We adopt dual-ASPP and dual-decoder modules, specific to semantic segmentation and instance segmentation, respectively. The semantic segmentation branch follows the typical design of any semantic segmentation model (e.g., DeepLab deeplabv3plus2018), while the instance segmentation prediction involves a simple instance center regression ballard1981generalizing; kendall2018multi, where the model learns to predict instance centers as well as the offset from each pixel to its corresponding center.

We perform experiments on Cityscapes cordts2016cityscapes and Mapillary Vistas neuhold2017mapillary datasets. On Cityscapes test set, a single Panoptic-DeepLab model achieves state-of-the-art performance of 65.5% PQ, 39.0% AP, and 84.2% mIoU, ranking first at all three Cityscapes tasks when comparing with published works. On Mapillary Vistas validation set, our best single model attains 40.6% PQ, while employing an ensemble of 6 models reaches a performance of 42.2% PQ.

To summarize, our contributions are as follows.

Panoptic-DeepLab is the first bottom-up approach that demonstrates state-of-the-art results for panoptic segmentation on Cityscapes and Mapillary Vistas.

Panoptic-DeepLab is the first single model (without fine-tuning on different tasks) that simultaneously ranks first at all three Cityscapes benchmarks.

Panoptic-DeepLab is simple in design, requiring only three loss functions during training, and introducing extra marginal parameters as well as additional slight computation overhead when building on top of a modern semantic segmentation model.

Methods

As illustrated in Fig. 1, our proposed Panoptic-DeepLab is deployed in a bottom-up single-shot manner for panoptic segmentation kirillov2018panoptic. Panoptic-DeepLab consists of four components: (1) an encoder backbone shared for both semantic segmentation and instance segmentation, (2) decoupled ASPP modules and (3) decoupled decoder modules specific to each task, and (4) task-specific prediction heads.

Architecture: The encoder backbone is adapted from an ImageNet-pretrained neural network paired with atrous convolution for extracting denser feature maps in its last block. The ASPP modules and decoder modules are separate for semantic segmentation and instance segmentation. Our light-weight decoder module gradually recovers the spatial resolution by a factor of 2; in each upsampling stage we apply only a single convolution.

Semantic segmentation: We employ the typical softmax cross entropy loss for semantic segmentation.

Class-agnostic instance segmentation: Motivated by Hough Voting ballard1981generalizing; kendall2018multi, we represent each object instance by its center of mass, encoded by a 2-D Gaussian with standard deviation of 8 pixels. For every foreground pixel (i.e., pixel whose class is a ‘thing’), we further predict the offset to its corresponding mass center. In particular, we adopt the Mean Squared Error (MSE) loss to minimize the distance between predicted heatmaps and 2D Gaussian-encoded groundtruth heatmaps. We use L1L_{1} loss for the offset prediction, which is only activated at pixels belonging to object instances. During inference, we group predicted foreground pixels by their closest predicted mass center, forming our class-agnostic instance segmentation results.

Panoptic segmentation: Given the predicted semantic segmentation and class-agnostic instance segmentation results, we adopt a fast and parallelizable method to merge the results, following the “majority vote” principle proposed in DeeperLab yang2019deeperlab. In particular, the semantic label of a predicted instance mask is inferred by the majority vote of the corresponding predicted semantic labels.

Experiments

Cityscapes cordts2016cityscapes: The dataset consists of 2975, 500, and 1525 traffic-related images for training, validation, and testing, respectively. It contains 8 ‘thing’ and 11 ‘stuff’ classes.

Mapillary Vistas neuhold2017mapillary: A large-scale traffic-related dataset, containing 18K, 2K, and 5K images for training, validation and testing, respectively. It contains 37 ‘thing’ classes and 28 ‘stuff’ classes in a variety of image resolutions, ranging from 1024×7681024\times 768 to more than 4000×60004000\times 6000

Experimental setup: We report mean IoU, average precision (AP), and panoptic quality (PQ) to evaluate the semantic, instance, and panoptic segmentation results.

All our models are trained using TensorFlow on 32 TPUs. We adopt a similar training protocol as in deeplabv3plus2018. In particular, we use the ‘poly’ learning rate policy with an initial learning rate of 0.0010.001, fine-tune the batch normalization parameters, perform random scale data augmentation during training, and optimize with Adam. On Cityscapes, our best setting is obtained by training with whole image (i.e., crop size equal to 1025×20491025\times 2049) with batch size 32. On Mapillary Vistas, we resize the images to 2177 pixels at the longest side to handle the large input variations, and randomly crop 1025×10251025\times 1025 patches during training with batch size 64. We set training iterations to 60K and 150K for Cityscapes and Mapillary Vistas, respectively. During evaluation, due to the sensitivity of PQ xiong2019upsnet; li2018learning; porzi2019seamless, we re-assign to ‘VOID’ label all ‘stuff’ segments whose areas are smaller than a threshold. The thresholds on Cityscapes and Mapillary Vistas are 2048 and 4096, respectively. Additionally, we adopt multi-scale inference (scales equal to {0.5,0.75,1,1.25,1.5,1.75,2}\{0.5,0.75,1,1.25,1.5,1.75,2\}) and left-right flipped inputs, to further improve the performance. For all the reported results, unless specified, Xception-71 deeplabv3plus2018 is employed as the backbone in Panoptic-DeepLab.

We conduct ablation studies on the Cityscapes validation set, as summarized in Tab. 1. Replacing the SGD momentum optimizer with the Adam optimizer yields 0.6% PQ improvement. Instead of using the sigmoid cross entropy loss for training the heatmap (i.e., instance center prediction), it brings 1.1% PQ improvement by applying the Mean Squared Error (MSE) loss to minimize the distance between the predicted heatmap and the 2D Gaussian-encoded groundtruth heatmap. It is more effective to adopt both dual-decoder and dual-ASPP, which gives us 0.7% PQ improvement while maintaining similar AP and mIoU. Employing a large crop size 1025×20491025\times 2049 (instead of 513×1025513\times 1025) during training further improves the AP and mIoU by 0.6% and 0.9% respectively. Finally, increasing the feature channels from 128 to 256 in the semantic segmentation branch achieves our best result of 63.0% PQ, 35.3% AP, and 80.5% mIoU. For reference, we train a Semantic-DeepLab under the same setting as the best Panoptic-DeepLab, showing that multi-task learning does not bring extra gain to mIoU. Note that Panoptic-DeepLab adds marginal parameters and small computation overhead over Semantic-DeepLab.

2 Cityscapes Results

Val set: In Tab. 2, we report our Cityscapes validation set results. When using only Cityscapes fine annotations, our best Panoptic-DeepLab, with mutli-scale inputs and left-right flips, outperforms the best bottom-up approach, SSAP, by 3.0% PQ and 1.2% AP, and is better than the best proposal-based approach, AdaptIS, by by 2.1% PQ, 2.2% AP, and 2.3% mIoU. When using extra data, our best Panoptic-DeepLab outperforms UPSNet by 5.2% PQ, 3.5% AP, and 3.9% mIoU, and Seamless by 2.0% PQ and 2.4% mIoU. Note that we do not exploit any other data, such as COCO, Cityscapes coarse annotations, depth, or video.

Test set: For the test set results, we additionally employ the trick proposed in deeplabv3plus2018 where we apply atrous convolution in the last two blocks within the backbone, with rate 2 and 4 respectively, during inference. We found this brings an extra 0.4% AP and 0.2% mIoU on val set but no improvement over PQ. We do not use this trick for the Mapillary Vistas Challenge. As shown in Tab. 3, our single unified Panoptic-DeepLab achieves state-of-the-art results, ranking first at all three Cityscapes tasks, when comparing with published works. Our model ranks second in the instance segmentation track when also taking into account unpublished entries.

3 Mapillary Vistas Challenge

Val set: In Tab. 4, we report Mapillary Vistas val set results. Our best single Panoptic-DeepLab model, with multi-scale inputs and left-right flips, outperforms the bottom-up approach, DeeperLab, by 8.3% PQ, and the top-down approach, Seamless, by 2.6% PQ. In Tab. 5, we report our results with three families of network backbones. We observe that naïve HRNet-W48 slightly under-performs Xception-71. Due to the diverse image resolutions in Mapillary Vistas, we found it important to enrich the context information as well as to keep high-resolution features. Therefore, we propose a simple modification for HRNet wang2019deep and Auto-DeepLab liu2019auto. For modified HRNet, called HRNet+, we keep its ImageNet-pretrained head and further attach dual-ASPP and dual-decoder modules. For modified Auto-DeepLab, called Auto-DeepLab+, we remove the stride in the original 1/32 branch (which improves PQ by 1%). To summarize, using Xception-71 strikes the best accuracy and speed trade-off, while HRNet-W48+ achieves the best PQ of 40.6%. Finally, our ensemble of 6 models attains a 42.2% PQ, 18.2% AP, and 58.7% mIoU.

References