DeepLab2: A TensorFlow Library for Deep Labeling

Mark Weber, Huiyu Wang, Siyuan Qiao, Jun Xie, Maxwell D. Collins, Yukun Zhu, Liangzhe Yuan, Dahun Kim, Qihang Yu, Daniel Cremers, Laura Leal-Taixe, Alan L. Yuille, Florian Schroff, Hartwig Adam, Liang-Chieh Chen

Introduction

Deep labeling refers to solving certain computer vision problems by assigning a predicted value for each pixel (i.e., label each pixel) in an image or video with a deep neural network . Typical dense prediction problems include, but not limited to, semantic segmentation , instance segmentation , panoptic segmentation , depth estimation , video panoptic segmentation , and depth-aware video panoptic segmentation .

Going beyond our previous open source libraryhttps://github.com/tensorflow/models/tree/master/research/deeplab in 2018 (which could only tackle image semantic segmentation with the first few DeepLab model variants ), we introduce DeepLab2, a modern TensorFlow library for deep labeling, aiming to provide a unified and easy-to-use TensorFlow codebase for general dense pixel labeling tasks. Re-implemented in TensorFlow2, this release includes all our recently developed DeepLab model variants , model training and evaluation code, and several pretrained checkpoints, allowing the community to reproduce and further improve upon the state-of-art systems. We hope that the open source DeepLab2 would facilitate future research on dense pixel labeling tasks, and anticipate novel breakthroughs and new applications that adopt this technology.

In the following sections, we detail a few popular dense prediction tasks as well as the provided state-of-the-art models in the DeepLab2 library.

Dense Pixel Labeling Tasks

Several computer vision problems could be formulated as dense pixel labeling. In this section, we briefly introduce some typical examples of dense pixel labeling tasks.

Image Semantic Segmentation , one step further than image-level classification for scene understanding, recognizes objects within an image with pixel-level accuracy, requiring precise outline of objects. It is usually formulated as pixel-wise classification , where each pixel is labeled by a predicted value encoding its semantic class.

Image Instance Segmentation recognizes and localizes object instances with pixel-level accuracy in an image. Existing models are mostly based on the top-down approach (i.e., bounding box detection followed by segmentation) and formulate the problem as mask detection (one step further than bounding box detection for instance-level understanding). On the contrary, our system tackles instance segmentation from the bottom-up perspective, detecting (or more precisely, grouping) instances on top of segmentation prediction. Therefore, our system labels each ‘thing’ pixel by a predicted value encoding both semantic class and instance identity (and the ‘stuff’ pixels are ignored). Note that our model generates non-overlapping instance masks, unlike other proposal-based models.

Image Panoptic Segmentation unifies semantic segmentation and instance segmentation. The task disallows overlapping instance masks and requires labeling each pixel (including ‘thing’ and ‘stuff’ pixels) with a predicted value encoding both semantic class and instance identity. We would like to highlight that our whole system, including the training and evaluation pipelines, uses the panoptic label format (i.e., panoptic_label=semantic_label×label_divisor+instance_idpanoptic\_label=semantic\_label\times label\_divisor+instance\_id), and thus do not take (during training mode) or generate (during inference mode) any overlapping masks . This is very different from most existing modern panoptic segmentation models , which are trained with overlapping instance masks.

Monocular Depth Estimation attempts to understand the 3D geometry of a scene by labeling each pixel with an estimated depth value.

Video Panoptic Segmentation extends the image panoptic segmentation to the video domain, where a temporally consistent instance identity is enforced across the video sequence.

Depth-aware Video Panoptic Segmentation provides in-depth scene understanding by solving a joint task of depth estimation, panoptic segmentation, and pixel-level tracking. Each pixel in a video is labeled with semantic class, temporally consistent instance identity, and estimated depth value.

Model Garden

Herein, we briefly introduce our developed DeepLab model variants that are included in the DeepLab2 library.

DeepLab , where atrous convolution (also known as convolution with holes or dilated convolutionhttps://www.tensorflow.org/api_docs/python/tf/nn/atrous_conv2d) is intensively exploited for semantic segmentation (see several previous works that effectively use atrous convolution in different ways). Specifically, DeepLabv1 employs atrous convolution to explicitly control the feature resolution computed by convolutional neural network backbones . Additionally, atrous convolution allows us to effectively enlarge the model’s field of view without increasing the number of parameters. As a result, the proposed ASPP module (Atrous Spatial Pyramid Pooling) in the follow-up DeepLab models effectively aggregates multi-scale information by employing parallel atrous conovlutions with multiple sampling rates.

Panoptic-DeepLab , a simple, fast, and strong bottom-up (i.e., proposal-free) baseline for panoptic segmentation. Panoptic-DeepLab adopts the dual-ASPP and dual-decoder structures specific to semantic and instance segmentation, respectively. The semantic segmentation branch is the same as DeepLab, while the instance segmentation branch is class-agnostic, involving a simple instance center regression . Even though simple, Panoptic-DeepLab yields state-of-the-art performance on multiple panoptic segmentation benchmarks.

Axial-DeepLab , building on top of the proposed Axial-ResNet backbones that efficiently capture long-range context with precise position information. Axial-ResNet enables a huge or even global receptive field in all the layers of a backbone by replacing spatial convolutions with axial-attention layers sequentially applied to the height- and width-axis. Additionally, a position-sensitive self-attention formulation is proposed to preserve context position in the huge receptive field. As a result, Axial-DeepLab, employing the proposed Axial-ResNet as backbone in the Panoptic-DeepLab framework , outperforms state-of-the-art convolutional counterparts on multiple panoptic segmentation benchmarks.

MaX-DeepLab , the first fully end-to-end system for panoptic segmentation. MaX-DeepLab directly predicts a set of segmentation masks and their corresponding semantic classes with a mask transformer, removing the needs for previous hand-designed modules (e.g., box anchors , thing-stuff merging heuristic or module ). The mask transformer is trained with a proposed PQ-style loss function and employs a dual-path architecture that enables an Axial-ResNet to read and write a global memory, allowing efficient communication (feature information exchange) between any Axial-ResNet layer and the transformer. MaX-DeepLab achieves the state-of-the-art panoptic segmentation results on COCO , outperforming both proposal-based and proposal-free approaches.

Motion-DeepLab , a unified model for the task of video panoptic segmentation, which requires to segment and track every pixel. It is built on top of Panoptic-DeepLab and uses an additional branch to regress each pixel to its center location in the previous frame. Instead of using a single RGB image as input, the network input contains two consecutive frames, i.e., the current and previous frame, as well as the center heatmap from the previous frame . The output is used to assign consistent track IDs to all instances throughout a video sequence.

ViP-DeepLab , a unified model that jointly tackles monocular depth estimation and video panoptic segmentation. It extends Panoptic-DeepLab by adding a depth prediction head to perform monocular depth estimation and a next-frame instance branch to generate panoptic predictions with temporally consistent instance IDs for videos. ViP-DeepLab achieves state-of-the-art performance on multiple benchmarks, including KITTI monocular depth estimation , KITTI multi-object tracking and segmentation , and Cityscapes video panoptic segmentation .

Supported Network Backbones

In this section, we briefly introduce the network backbones that are supported by the DeepLab2 library.

MobileNetv3 , a light-weight backbone designed for mobile devices , which thus could be used as a baseline for fast on-device model comparison.

ResNet , along with some modern modifications (e.g., Inception stem , stochastic drop path ). This backbone could be used as a general baseline.

SWideRNet , which scales Wide ResNets in both width (number of channels) and depth (number of layers). This could be used for strong model comparison.

Axial-ResNet and Axial-SWideRNet , which replaces ResNet and SWideRNet residual blocks with axial-attention blocks in the last few stages. This hybrid CNN-Transformer (more precisely, only self-attention modules as the Transformer encoder) architecture could be used as a baseline for self-attention based model comparison. Besides the default axial-attention blocks, we also support global or non-local attention blocks.

MaX-DeepLab , which incorporates transformer blocks to Axial-ResNets in a dual-path manner, allowing efficient communication between any Axial-ResNet layers and the transformers. When turning off the transformer path, this backbone could be also adopted in the Panoptic-DeepLab framework .

To briefly highlight the effectiveness of DeepLab2, our Panoptic-DeepLab employing the Axial-SWideRNet as network backbone achieves 68.0% PQ or 83.5% mIoU on Cityscapes validation set, with single-scale inference and ImageNet-1K pretrained checkpoints. For more detailed results, we refer the readers to the provided model zoo https://github.com/google-research/deeplab2/tree/main/g3doc/projects.

Finally, we would like to emphasize that we have implemented the Axial-Block in a general way that encompasses not only transformer-based blocks (i.e., axial-attention , global-attention , dual-path transformer ) but also convolutional residual blocks (i.e., basic block, bottleneck block , wide residual block , each w/ or w/o Switchable Atrous Convolution ). Our design allows users to easily develop novel neural networks that efficiently combine convolution , attention , and transformer (i.e., self- and cross- attention) operations.

Data Augmentation During Training

In addition to the typical data augmentation (i.e. random scaling, left-right flipping, and random cropping) used for dense prediction tasks, we also support:

A random color jittering found by AutoAugment . In , we apply this augmentation policy with magnitude 1.01.0, and 0.20.2 on COCO and Cityscapes datasets, respectively.

Conclusion

We open-source DeepLab2, containing all our recent research results, in the hope that it would facilitate future research on dense prediction tasks. The codebase is still under active development, and any contribution to the codebase from the community is very welcome. Finally, to reiterate, the code along with model zoo could be found at https://github.com/google-research/deeplab2.

We would like to thank Michalis Raptis for the feedback on the paper, Jiquan Ngiam and Amil Merchant for Hungarian Matching implementation, and the support from Google Mobile Vision.

References